Self-Supervised No-Reference Image Quality Assessment Method and System

Through the self-supervised learning method, image restoration tasks are used for pre-training, and parameters are transferred through the knowledge transfer channel, the problem of insufficient training data in the reference-free image quality evaluation method is solved, efficient image quality evaluation is achieved, and the model's performance and perception ability of distorted information are improved.

CN114358204BActive Publication Date: 2025-07-01INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210028835.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-07-01
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

The existing reference-free image quality evaluation method results in poor model performance due to insufficient training data, and the models pre-trained on the classification task are difficult to fully extract the distorted information of the image.

Method used

A self-supervised learning method is adopted to construct a self-supervised reference-free image quality evaluation model through image restoration tasks as agent tasks. The model includes a prior knowledge learning subnet of the shared encoder and an image quality evaluation subnet. The parameters of the pre-trained prior knowledge learning subnet are transferred to the image quality evaluation subnet through the knowledge transfer channel for fine-tuning training.

Benefits of technology

It effectively breaks through the limitations of insufficient training data, improves the performance of the image quality evaluation model, improves the accuracy and accuracy of image quality recognition, and enhances the model's perception of global and local distortion information of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358204B_ABST
    Figure CN114358204B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image quality assessment, and particularly relates to a self-supervised no-reference image quality assessment method and system, aiming to solve the problem of poor performance of the image quality assessment model due to insufficient training data in the prior art. The present invention includes: constructing a self-supervised no-reference image quality assessment model composed of a prior knowledge learning sub-network and an image quality assessment sub-network with a shared encoder; using the image restoration task as a proxy task for the image quality assessment task to pre-train the prior knowledge learning sub-network; performing knowledge transfer through a knowledge transfer channel established between the prior knowledge learning sub-network and the decoder of the image quality assessment sub-network; performing model fine-tuning training on the image quality assessment task; and performing quality assessment of no-reference images through the trained model. The model of the present invention can obtain good performance by training only on less data, has high training efficiency, and high accuracy in image quality assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Image quality assessment (IQA) is widely used in scenarios such as audio-visual entertainment, medical imaging, and aerial remote sensing. According to the availability of reference images, IQA methods can be divided into full-reference, semi-reference, and no-reference methods. In many application scenarios, it is difficult to obtain reference images. Therefore, the no-reference IQA (NR-IQA) method, which evaluates the quality of distorted images without reference images, is more practical and has received more extensive attention. The key to the NR-IQA method is to extract an effective representation of the distorted image and then map it to a quality score. Traditional NR-IQA methods usually use manually designed features, such as wavelet transform coefficients, discrete transform coefficients, and luminance coefficients of images. In recent years, convolutional neural networks (CNNs) have made a series of progress in the field of computer vision, and many scholars have also used the powerful feature extraction ability of CNNs to improve the performance of NR-IQA methods. However, training an excellent CNN model requires a large amount of manually labeled data, that is, the quality scores of images. This score is generally obtained by having subjects strictly follow the scoring rules to give the quality scores of distorted images under laboratory conditions and then taking the average of the quality scores of each image. Therefore, the production cost of the IQA dataset is high and time-consuming. Difficulty in obtaining a large-scale labeled sample has become one of the main factors restricting the performance of NR-IQA models. To alleviate the problem of insufficient training data, many NR-IQA methods adopt a pre-training mechanism. First, they are pre-trained on other tasks and then fine-tuned on the NR-IQA task. Representative models include DBCNN and MetaIQA. These methods are generally pre-trained on classification tasks, but classification tasks mainly focus on high-level semantic information and pay less attention to low-level distortion information. Therefore, they are not conducive to the quality assessment of distorted images.

[0003] Some literature has proposed a method for no-reference image quality assessment through a deep bilinear convolutional neural network (DBCNN scheme) [1]. First, the VGG-16 module is a model pre-trained for object classification tasks on the ImageNet dataset. Then, an S-CNN module with the same structure as VGG-16 is set up and pre-trained on the classification tasks of distortion types and distortion degrees. VGG-16 is responsible for extracting the semantic features of the input image, and S-CNN is responsible for extracting the distortion features of the input image. Finally, the semantic and distortion features are fused using the method of bilinear pooling and input into the FC layer to predict the quality score of the image. Some other literature has proposed a method for no-reference image quality assessment through meta-learning (MetaIQA scheme) [2]. Using the idea of meta-learning, first, some distortion types are mastered through pre-training, and then the learned knowledge is transferred to unknown distortion types to explore the correlation between known and unknown distortions, so as to perform better learning. The model training process mainly includes: (1) Selecting distorted images of some distortion types and their corresponding quality scores to pre-train the model; (2) Using the distorted images of all distortion types in the target IQA dataset and their corresponding quality scores to fine-tune the model pre-trained in (1), and finally obtaining the NR-IQA model.

[0004] However, the DBCNN scheme uses two pre-trained networks to extract semantic and distortion information, but the S-CNN for extracting distortion information is also trained on classification tasks. This results in poor pre-training effects when there are many distortion types and degrees, and the classification task belongs to supervised learning, which will also reduce the generalization ability of the model to a certain extent. The MetaIQA scheme uses the idea of meta-learning and inputs distorted images and their corresponding scores into the scheme network for pre-training to let the network perceive the score levels of different distortion information in advance. However, this pre-training process still depends on the IQA dataset, and the production processes and scoring rules of different IQA datasets are different, which will lead to the problem that the model is not adaptable when fine-tuning on other datasets. And because only the image quality scores are constrained in the pre-training process, this will cause the scheme model to be unable to accurately extract the semantic and distortion information of the image, resulting in the problem of weak model interpretability.

[0005] Generally speaking, in the NR-IQA task, predicting the quality score of an image often requires considering both high-level semantics and low-level distortion information at the same time. However, the above two schemes both use models pre-trained on classification tasks (such as VGG, ResNet) to extract semantic information and rarely consider the distortion information of the image. Therefore, they cannot fully adapt to the NR-IQA task.

[0006] The following literature is the technical background information related to the present invention:

[0007] [1] Zhang W, Ma K, Yan J, et al. Blind Image Quality Assessment Using A Deep Bilinear Convolutional Neural Network[J]. IEEE Transactions on Circuits & Systems for Video Technology, 2019: 1-1.

[0008] [2] Zhu H, Li L, Wu J, et al. MetaIQA: Deep Meta-learning for No-Reference Image Quality Assessment[J]. 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. Summary of the Invention

[0009] To solve the above problems in the prior art, that is, the problem of poor performance of the image quality assessment model due to insufficient training data in the prior art, the present invention provides a self-supervised no-reference image quality assessment method, which includes:

[0010] Step S10: Construct a self-supervised no-reference image quality assessment model, which includes a prior knowledge learning sub-network and an image quality assessment sub-network sharing an encoder, and a knowledge transfer channel is set between the decoder of the prior knowledge learning sub-network and the decoder of the image quality assessment sub-network;

[0011] Step S20: Use the image restoration task as a proxy task for the image quality assessment task, construct a first training set for the image restoration task, and pre-train the prior knowledge learning sub-network through the first training set;

[0012] Step S30: Transfer the parameters of the decoder of the pre-trained prior knowledge learning sub-network to the decoder of the image quality assessment sub-network through the knowledge transfer channel;

[0013] Step S40: Obtain a small second training set for the image quality assessment task, and perform fine-tuning training on the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer through the second training set;

[0014] Step S50: Use the fine-tuned prior knowledge learning sub-network and image quality assessment sub-network as the trained self-supervised no-reference image quality assessment model, and perform no-reference image quality assessment through the trained self-supervised no-reference image quality assessment model.

[0015] In some preferred embodiments, for the prior knowledge learning sub-network and the image quality assessment sub-network, their decoders are distortion information perception modules;

[0016] The distortion information perception module includes a first splicing layer, an Inception layer, a second splicing layer, and three 3×3 convolutional layers connected in sequence;

[0017] The Inception layer includes a first branch, a second branch, a third branch, and a fourth branch. The first branch is a 1×1 convolutional layer, the second branch is a 1×1 convolutional layer and a 5×5 convolutional layer connected in sequence, the third branch is a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in sequence, and the fourth branch is a max pooling layer and a 3×3 convolutional layer connected in sequence.

[0018] In some preferred embodiments, the knowledge transfer channel includes a knowledge collection module and a knowledge distribution module;

[0019] The knowledge collection module includes m branch inputs and a third splicing layer. Each branch input includes a 1×1 convolutional layer and a pooling layer connected in sequence. The pooling layers of the m branches are connected to the first splicing layer together, where m is the number of scales of the multi-scale context features of the images output by the decoder of the prior knowledge learning sub-network;

[0020] The knowledge distribution module includes n branch outputs connected to the third splicing layer. Each branch output includes a fourth splicing layer and a 1×1 convolutional layer connected in sequence, where n is the number of scales of the multi-scale context features of the images output by the decoder of the image quality assessment sub-network.

[0021] In some preferred embodiments, for the prior knowledge learning sub-network, the first loss function pre-trained based on the first training set for the image restoration task is:

[0022]

[0023] where I respectively represents the reference images in the first training set and the corresponding restored images I, represents the reference image and the first loss function between the corresponding restored images I, represents the reference image the L1 loss function between the reference image and the corresponding restored image I represents the reference image the SSIM loss function between the reference image and the corresponding restored image I represents the reference image the perceptual loss function between the reference image and the corresponding restored image I, where ε, ρ, μ are preset hyperparameters for adjusting the proportions of the L1 loss function, SSIM loss function, and perceptual loss function in the first loss function, and N is the number of image pairs of the pre-trained reference images in the first training set and the corresponding restored image I

[0024] In some preferred embodiments, the preset hyperparameters for adjusting the proportions of the L1 loss function, SSIM loss function, and perceptual loss function in the first loss function are respectively set to ε = 1, ρ = 0.08, and μ = 1

[0025] In some preferred embodiments, the reference image and the L1 loss function between the corresponding restored image I which is expressed as

[0026]

[0027] where ‖‖1 represents the L1 norm

[0028] In some preferred embodiments, the reference image and the SSIM loss function between the corresponding restored image I which is expressed as

[0029]

[0030] where represents calculating the structural similarity between the reference image and the corresponding restored image I

[0031] In some preferred embodiments, the reference image and the perceptual loss function between the corresponding restored image I which is expressed as

[0032]

[0033] where represents the feature map of the t-th layer of the prior knowledge learning sub-network with the reference image as the input, and ψ t (I) represents the reference image The corresponding restored image I is the feature map of the t-th layer of the input prior knowledge learning sub-network, T represents the number of layers of the pre-trained prior knowledge learning sub-network, and Θ t represents the total number of feature points on the feature map of the t-th layer, and ‖‖1 represents the L1 norm.

[0034] In some preferred embodiments, for the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer, the second loss function for fine-tuning training based on the second training set of the image quality assessment task is:

[0035]

[0036] where represents the second loss function, g and p respectively represent the labeled quality scores of the images in the second training set and the predicted quality scores of the corresponding image quality assessment sub-network, and || represents taking the absolute value.

[0037] On the other hand, the present invention proposes a self-supervised no-reference image quality assessment system, which includes the following modules:

[0038] A model construction module configured to construct a self-supervised no-reference image quality assessment model, the model includes a prior knowledge learning sub-network and an image quality assessment sub-network sharing an encoder, and a knowledge transfer module is provided between the decoder of the prior knowledge learning sub-network and the decoder of the image quality assessment sub-network;

[0039] A pre-training module configured to use the image restoration task as a proxy task for the image quality assessment task, construct a first training set for the image restoration task, and pre-train the prior knowledge learning sub-network through the first training set;

[0040] A knowledge transfer module configured to transfer the parameters of the decoder of the pre-trained prior knowledge learning sub-network to the decoder of the image quality assessment sub-network through a knowledge transfer channel;

[0041] A fine-tuning training module configured to obtain a small amount of second training set for the image quality assessment task, and perform fine-tuning training on the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer through the second training set;

[0042] An image quality assessment module configured to use the fine-tuned prior knowledge learning sub-network and image quality assessment sub-network as a trained self-supervised no-reference image quality assessment model, and perform no-reference image quality assessment through the trained self-supervised no-reference image quality assessment model.

[0043] Advantages of the present invention:

[0044] (1) The self-supervised reference-free image quality assessment method of the present invention uses self-supervised learning to pre-train the prior knowledge learning sub-network of the self-supervised reference-free image quality assessment model on an image restoration task, and then transfers it to the image quality assessment task for fine-tuning of the image quality assessment sub-network. It can break through the limitation of insufficient training data for image quality assessment, effectively improve the performance of the image quality assessment model, and further improve the accuracy and precision of image quality recognition.

[0045] (2) The self-supervised reference-free image quality assessment method of the present invention establishes a knowledge transfer channel between the prior knowledge learning sub-network and the image quality assessment sub-network of the self-supervised reference-free image quality assessment model, which can achieve the aggregation and distribution of knowledge and promote the effective transfer of knowledge.

[0046] (3) The self-supervised reference-free image quality assessment method of the present invention has a distortion information perception module as the decoder of the sub-network, which can effectively perceive the global and local distortion information of the image and improve the accuracy of image restoration and image quality assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Other features, objectives, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0048] Figure 1 is a schematic flowchart of the self-supervised reference-free image quality assessment method of the present invention;

[0049] Figure 2 is a schematic structural diagram of the distortion perception module of an embodiment of the self-supervised reference-free image quality assessment method of the present invention;

[0050] Figure 3 is a schematic structural diagram of the knowledge transfer channel of an embodiment of the self-supervised reference-free image quality assessment method of the present invention;

[0051] Figure 4 is a schematic diagram of the quality assessment of a reference-free image of an embodiment of the self-supervised reference-free image quality assessment method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0053] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will detail the present application with reference to the drawings and in combination with the embodiments.

[0054] The present invention provides a self-supervised no-reference image quality assessment method, which adopts a new NR-IQA (no-reference image quality assessment) training method. Using the idea of self-supervised learning, first, an agent task is determined. This agent task should have a strong correlation with the target NR-IQA task, the training data is easy to obtain, sufficient in quantity, and does not require a human-annotated supervision signal, but can directly obtain the supervision signal from the input data itself. The image restoration task meets the above conditions, so the present invention selects the image restoration task as the agent task. A dataset for the image restoration task is constructed, and the self-supervised no-reference image quality assessment model is pre-trained on this dataset. Through pre-training, the prior knowledge learning sub-network of the self-supervised no-reference image quality assessment model can learn effective feature representations of distorted images in the image restoration task, including semantic features and distortion features, and regard them as a kind of prior knowledge. Then, the pre-trained prior knowledge learning sub-network is transferred to the image quality assessment task for fine-tuning of the image quality assessment sub-network. Since prior knowledge about semantic and distortion features has been learned in the pre-training stage, the image quality assessment sub-network only needs less IQA training data to quickly converge to a better performance.

[0055] The IQA dataset contains several distorted images with different distortion types and degrees. The distortion of the images may be caused by natural reasons such as shooting jitter, device noise, overexposure, etc., or is artificially added distortion (such as Gaussian blur, Gaussian noise, overexposure, etc.) to high-definition images. Each distorted image has an artificially annotated image quality score.

[0056] A self-supervised no-reference image quality assessment method of the present invention, the no-reference image quality assessment method includes:

[0057] Step S10, constructing a self-supervised no-reference image quality assessment model, the model includes a prior knowledge learning sub-network with a shared encoder and an image quality assessment sub-network, and a knowledge transfer channel is provided between the decoder of the prior knowledge learning sub-network and the image quality assessment sub-network;

[0058] Step S20, taking the image restoration task as the agent task of the image quality assessment task, constructing a first training set for the image restoration task, and pre-training the prior knowledge learning sub-network through the first training set;

[0059] Step S30: Transfer the parameters of the decoder of the pre-trained prior knowledge learning sub-network to the decoder of the image quality assessment sub-network through the knowledge transfer channel;

[0060] Step S40: Obtain a small second training set for the image quality assessment task, and perform fine-tuning training on the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network with transferred parameters through the second training set;

[0061] Step S50: Use the fine-tuned prior knowledge learning sub-network and image quality assessment sub-network as the trained self-supervised no-reference image quality assessment model, and perform quality assessment of no-reference images through the trained self-supervised no-reference image quality assessment model.

[0062] To more clearly illustrate the self-supervised no-reference image quality assessment method of the present invention, the following combines Figure 1 Expand and detail each step in the embodiments of the present invention.

[0063] The self-supervised no-reference image quality assessment method according to the first embodiment of the present invention includes steps S10 - S50, and each step is described in detail as follows:

[0064] Step S10: Construct a self-supervised no-reference image quality assessment model, the model includes a prior knowledge learning sub-network and an image quality assessment sub-network sharing an encoder, and a knowledge transfer channel is provided between the decoders of the prior knowledge learning sub-network and the image quality assessment sub-network.

[0065] In view of the fact that both sub-networks need to fully perceive the distortion information of the input image, the prior knowledge learning sub-network can restore a clean image, and the image quality assessment sub-network can correctly predict the quality score of the image. However, since the decoder takes the high-level semantic features output by the encoder as input, the low-level distortion information is lost to a certain extent. In response to this, the present invention designs a special network module for the decoders of the two sub-networks, called the distortion perception module. By transferring the distortion features in the encoder to the decoder, the distortion information perception ability of the decoder is enhanced, and the correct prediction result is output. This module takes the feature map output by the previous layer of the decoder and the feature map output by the middle layer of the encoder as input. First, it passes through a splicing layer to splice the two inputs together along the channel layer, and then passes through an inception layer. The inception layer contains four branches inside, and each branch consists of multiple convolutional kernels or pooling layers of different sizes. The inception layer is used in this module to mine distortion information of different scales. The feature maps output by the four branches pass through a splicing layer and are spliced together along the channel. Finally, it passes through 3 convolutional layers to output the feature map with enhanced distortion information.

[0066] AsFigure 2 As shown in the figure, it is a schematic structural diagram of a distortion perception module of an embodiment of the self-supervised reference-free image quality assessment method of the present invention. The distortion information perception module includes a first splicing layer, an Inception layer, a second splicing layer, and three 3×3 convolutional layers connected in sequence:

[0067] The Inception layer includes a first branch, a second branch, a third branch, and a fourth branch. The first branch is a 1×1 convolutional layer. The second branch is a 1×1 convolutional layer and a 5×5 convolutional layer connected in sequence. The third branch is a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in sequence. The fourth branch is a max pooling layer and a 3×3 convolutional layer connected in sequence.

[0068] The present invention designs a knowledge transfer channel for connecting two sub-networks. This channel includes two parts: knowledge collection and knowledge distribution. The knowledge collection part is used to aggregate the multi-scale context features of the prior knowledge sub-network. The features of each scale map are adjusted to the same number of channels through a 1×1 convolutional layer respectively, then adjusted to the same size through a pooling layer respectively, and finally spliced together to form an aggregated feature. In the knowledge distribution part, the aggregated feature is distributed to the feature maps of different scales of the image quality assessment sub-network. Each scale of feature map retains the useful features selected from the aggregated feature and discards the useless features. In each distribution process, first, the aggregated feature is spliced with the feature of a certain scale in the image quality assessment sub-network through a splicing layer, and then the two are fused through a 1×1 convolutional layer to realize feature screening. The fused feature continues to be transmitted backward along the image quality assessment sub-network.

[0069] As Figure 3 shown in the figure, it is a schematic structural diagram of a knowledge transfer channel of an embodiment of the self-supervised reference-free image quality assessment method of the present invention. The knowledge transfer channel includes a knowledge collection module and a knowledge distribution module:

[0070] The knowledge collection module includes m branch inputs and a third splicing layer. Each branch input includes a 1×1 convolutional layer and a pooling layer connected in sequence. The pooling layers of the m branches are connected to the first splicing layer together, where m is the number of scales of the multi-scale context features of the image output by the decoder of the prior knowledge learning sub-network;

[0071] The knowledge distribution module includes n branch outputs connected to the third splicing layer. Each branch output includes a fourth splicing layer and a 1×1 convolutional layer connected in sequence, where n is the number of scales of the multi-scale context features of the image output by the decoder of the image quality assessment sub-network.

[0072] In one embodiment of the present invention, the encoder of the prior knowledge learning sub-network and the image quality assessment sub-network adopts the first 5 layers of ResNet-50. The decoder of the prior knowledge learning sub-network contains 5 distortion perception modules. The decoder of the image quality assessment sub-network contains 6 basic units and 4 convolutional layers. Each basic unit contains a convolutional layer, a ReLU activation layer, and an average pooling layer. The outputs of the first 4 distortion perception modules in the decoder of the prior knowledge learning sub-network are connected to the output features of the 2nd - 5th basic units of the image quality assessment through the knowledge transfer channel. The output of the last basic unit of the image quality assessment sub-network passes through 4 convolutional layers and then outputs the final quality score.

[0073] Step S20: Use the image restoration task as a proxy task for the image quality assessment task, construct a first training set for the image restoration task, and pre-train the prior knowledge learning sub-network through the first training set.

[0074] In one embodiment of the present invention, a large-scale high-definition distortion-free image is collected. For example, the Waterloo dataset contains 4,744 high-definition distortion-free images. The collected images are used as reference images. A random type of distortion, such as noise, blur, compression distortion, overexposure, etc., is randomly imposed on each reference image, and the degree of distortion is random, thereby generating a large-scale distorted image. Each distorted image and its corresponding reference image form a training sample, and a first training set for training the prior knowledge learning sub-network is obtained.

[0075] Randomly select a training sample from the first training set, input the distorted image in the training sample into the prior knowledge learning sub-network, output a restored image, and calculate the first loss in the pre-training. The first loss function is shown in Equation (1):

[0076]

[0077] Where: I respectively represents the reference image in the first training set and the corresponding restored image I, represents the reference image and the first loss function between the corresponding restored image I, represents the reference image and the L1 loss function between the corresponding restored image I, represents the reference image and the SSIM loss function between the corresponding restored image I, represents the reference image The perceptual loss function between the reference image and the corresponding restored image I, where ε, ρ, and μ are preset hyperparameters for adjusting the proportions of the L1 loss function, SSIM loss function, and perceptual loss function in the first loss function, and N is the number of image pairs of the pre-trained reference images in the first training set. and the corresponding restored image I.

[0078] The preset hyperparameters for adjusting the proportions of the L1 loss function, SSIM loss function, and perceptual loss function in the first loss function are respectively set as ε = 1, ρ = 0.08, and μ = 1.

[0079] Reference image The L1 loss function between the reference image is expressed as shown in Equation (2):

[0080]

[0081] where || ||1 represents the L1 norm.

[0082] Reference image The SSIM loss function between the reference image is expressed as shown in Equation (3):

[0083]

[0084] where represents the structural similarity between the reference image and the corresponding restored image I.

[0085] Reference image The perceptual loss function between the reference image is expressed as shown in Equation (4):

[0086]

[0087] where represents the feature map of the t-th layer of the prior knowledge learning sub-network with the reference image as the input, μ t (I) represents the feature map of the t-th layer of the prior knowledge learning sub-network with the corresponding restored image I of the reference image as the input, T represents the number of layers of the pre-trained prior knowledge learning sub-network, Θ t represents the total number of feature points on the feature map of the t-th layer, and || ||1 represents the L1 norm.

[0088] The parameters in the prior knowledge learning sub-network are adjusted by the gradient backpropagation algorithm. The prior knowledge learning sub-network is iteratively trained with the training samples in the first training set continuously until the set loop end condition is met, and a pre-trained prior knowledge learning sub-network is obtained. In an embodiment of the present invention, the loop end condition is set to loop 5000 times.

[0089] Step S30, transfer the parameters of the decoder of the pre-trained prior knowledge learning sub-network to the decoder of the image quality assessment sub-network through the knowledge transfer channel.

[0090] Step S40, obtain a small second training set of image quality assessment tasks, and perform fine-tuning training on the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer through the second training set.

[0091] In an embodiment of the present invention, an IQA training data set is selected. For example, existing IQA data sets include TID2013, LIVE, CSIQ data sets, etc. Taking the TID2013 data set as an example, it contains 3000 distorted images, among which there are 24 types of distortions and 5 levels of distortion degrees. Each distorted image has an artificially labeled image quality score, and a distorted image and its corresponding artificially labeled image quality score form a sample, obtaining the second training set for fine-tuning training of the image quality assessment sub-network.

[0092] The image quality assessment sub-network and the prior knowledge learning sub-network are trained on the image quality assessment task. Randomly select a training sample from the second training set, input the distorted images in the training sample into the prior knowledge learning sub-network and the image quality assessment sub-network respectively, splice the outputs of the first 4 distortion perception modules of the decoder of the prior knowledge learning sub-network together through the splicing layer to obtain the feature f1, and transfer f1 to the 2nd - 5th basic units of the image quality assessment sub-network through the knowledge transfer channel. The image quality assessment sub-network outputs a predicted image quality score. For the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer, the second loss function for fine-tuning training based on the second training set of the image quality assessment task is shown in Equation (5):

[0093]

[0094] Among them, represents the second loss function, g and p respectively represent the labeled quality score of the image in the second training set and the predicted quality score of the corresponding image quality assessment sub-network, and || represents taking the absolute value.

[0095] Adjust the parameters in the image quality assessment sub-network and fine-tune the parameters in the prior knowledge learning sub-network through the gradient backpropagation algorithm.

[0096] Continuously iterate for model training until the number of loops meets a set value, such as 5000 times, at which point the training ends and the parameters of both sub-networks are all fixed.

[0097] Step S50: Use the prior knowledge learning sub-network and the image quality assessment sub-network trained by fine-tuning as the trained self-supervised no-reference image quality assessment model, and perform the quality assessment of no-reference images through the trained self-supervised no-reference image quality assessment model.

[0098] After being trained through the above steps, the self-supervised no-reference image quality assessment model of the present invention can be directly used for the quality assessment of distorted images. As Figure 4 shown, it is a schematic diagram of the quality assessment of a no-reference image in an embodiment of the self-supervised no-reference image quality assessment method of the present invention. In actual use, the distorted image is input into this device, and the distorted image enters the prior knowledge learning sub-network and the image quality assessment sub-network respectively. The outputs of the first 4 distortion perception modules of the decoder of the prior knowledge learning sub-network are spliced together through a splicing layer to obtain feature f1, and f1 is transmitted to the 2nd - 5th basic units of the image quality assessment sub-network through the knowledge transfer channel. The image quality assessment sub-network outputs the quality score predicted for the distorted image.

[0099] Although each step is described in the above order in the above embodiments, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.

[0100] The self-supervised no-reference image quality assessment system according to the second embodiment of the present invention, the no-reference image quality assessment system includes the following modules:

[0101] A model construction module, configured to construct a self-supervised no-reference image quality assessment model, the model includes a prior knowledge learning sub-network and an image quality assessment sub-network sharing an encoder, and a knowledge transfer module is provided between the decoders of the prior knowledge learning sub-network and the image quality assessment sub-network;

[0102] A pre-training module, configured to use the image restoration task as a proxy task for the image quality assessment task, construct a first training set for the image restoration task, and perform pre-training of the prior knowledge learning sub-network through the first training set;

[0103] A knowledge transfer module, configured to transfer the parameters of the decoder of the pre-trained prior knowledge learning sub-network to the decoder of the image quality assessment sub-network through a knowledge transfer channel;

[0104] The fine-tuning training module is configured to obtain a second training set of a small number of image quality assessment tasks, and perform fine-tuning training on the prior knowledge learning sub-network pre-trained by the second training set and the image quality assessment sub-network after parameter transfer;

[0105] The image quality assessment module is configured to use the fine-tuned prior knowledge learning sub-network and the image quality assessment sub-network as a trained self-supervised no-reference image quality assessment model, and perform quality assessment of no-reference images through the trained self-supervised no-reference image quality assessment model.

[0106] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and related descriptions of the above-described system can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0107] It should be noted that the above-described self-supervised no-reference image quality assessment system provided by the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. For the names of the modules and steps involved in the embodiments of the present invention, they are only used to distinguish each module or step, and are not regarded as an improper limitation of the present invention.

[0108] An electronic device according to a third embodiment of the present invention includes:

[0109] At least one processor; and

[0110] A memory communicatively connected to at least one of the processors; wherein,

[0111] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned self-supervised no-reference image quality assessment method.

[0112] A computer-readable storage medium according to a fourth embodiment of the present invention stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned self-supervised no-reference image quality assessment method.

[0113] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and related descriptions of the above-described storage device and processing device can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0114] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0115] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or represent a specific order or sequence.

[0116] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, so that a process, method, article, or device / equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent in these processes, methods, articles, or devices / equipment.

[0117] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A self-supervised no-reference image quality assessment method, characterized in that The no-reference image quality assessment method includes: Step S10, constructing a self-supervised no-reference image quality assessment model, the model includes a prior knowledge learning sub-network and an image quality assessment sub-network sharing an encoder, and a knowledge transfer channel is arranged between the decoders of the prior knowledge learning sub-network and the image quality assessment sub-network; the knowledge transfer channel includes two parts: knowledge collection and knowledge distribution; for the prior knowledge learning sub-network and the image quality assessment sub-network, their decoders are distortion information perception modules; Step S20, using the image restoration task as a proxy task for the image quality assessment task, constructing a first training set for the image restoration task, and pre-training the prior knowledge learning sub-network through the first training set; Step S30, transferring the parameters of the decoder of the pre-trained prior knowledge learning sub-network to the decoder of the image quality assessment sub-network through the knowledge transfer channel; Step S40, obtaining a small second training set for the image quality assessment task, and performing fine-tuning training on the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer through the second training set; Step S50, using the fine-tuned prior knowledge learning sub-network and image quality assessment sub-network as the trained self-supervised no-reference image quality assessment model, and performing no-reference image quality assessment through the trained self-supervised no-reference image quality assessment model.

2. The no-reference image quality assessment method based on self-supervision according to claim 1, wherein The distortion information perception module includes a first splicing layer, an Inception layer, a second splicing layer, and three 3×3 convolutional layers connected in sequence; The Inception layer includes a first branch, a second branch, a third branch, and a fourth branch. The first branch is a 1×1 convolutional layer, the second branch is a 1×1 convolutional layer and a 5×5 convolutional layer connected in sequence, the third branch is a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in sequence, and the fourth branch is a max pooling layer and a 3×3 convolutional layer connected in sequence.

3. The no-reference image quality assessment method based on self-supervision according to claim 1, characterized in that The knowledge transfer channel includes a knowledge collection module and a knowledge distribution module; The knowledge collection module includes m branch inputs and a third splicing layer. Each branch input includes a 1×1 convolutional layer and a pooling layer connected in sequence. The pooling layers of the m branches are connected to the first splicing layer together, where m is the number of scales of the multi-scale context features of the image output by the decoder of the prior knowledge learning sub-network; The knowledge distribution module includes n branch outputs connected to the third splicing layer. Each branch output includes a fourth splicing layer and a 1×1 convolutional layer connected in sequence, where n is the number of scales of the multi-scale context features of the image output by the decoder of the image quality assessment sub-network.

4. The no-reference image quality assessment method based on self-supervision according to claim 3, wherein For the prior knowledge learning sub-network, the first loss function for pre-training based on the first training set of the image restoration task is: Among them, respectively represent the reference images in the first training set and the corresponding restored image I, represents the first loss function between the reference image and the corresponding restored image I, represents the L1 loss function between the reference image and the corresponding restored image I, represents the SSIM loss function between the reference image and the corresponding restored image I, represents the perceptual loss function between the reference image and the corresponding restored image I. ε, ρ, μ are preset hyperparameters for adjusting the proportions of the L1 loss function, SSIM loss function, and perceptual loss function in the first loss function. N is the number of image pairs of the reference images and the corresponding restored images I in the pre-trained first training set and the corresponding restored image I.

5. The method for self-supervised reference-free image quality assessment according to claim 4, wherein The preset hyperparameters for adjusting the proportions of the L1 loss function, SSIM loss function, and perceptual loss function in the first loss function are respectively set as ε = 1, ρ = 0.08, μ = 1.

6. The no-reference image quality assessment method based on self-supervision according to claim 4, characterized in that The reference image and the L1 loss function between the corresponding restored image I which is expressed as: Among them, ‖‖1 represents the L1 norm.

7. The no-reference image quality assessment method based on self-supervision according to claim 4, wherein The reference image and the SSIM loss function between the corresponding restored image I which is expressed as: Among them, represents the structural similarity between the reference image and the corresponding restored image I.

8. The no-reference image quality assessment method based on self-supervision according to claim 4, characterized in that The reference image and the perceptual loss function between the corresponding restored image I which is expressed as: Among them, represents the feature map of the t-th layer of the prior knowledge learning sub-network with the reference image as the input, ψ t (I) represents the feature map of the t-th layer of the prior knowledge learning sub-network with the restored image I corresponding to the reference image as the input, T represents the number of layers of the pre-trained prior knowledge learning sub-network, Θ t represents the total number of feature points on the feature map of the t-th layer, and ‖‖1 represents the L1 norm.

9. The no-reference image quality assessment method based on self-supervision according to claim 1, wherein For the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer, the second loss function for fine-tuning training based on the second training set of the image quality assessment task is: Among them, represents the second loss function, where g and p represent the label quality score of the images in the second training set and the predicted quality score of the corresponding image quality assessment sub-network respectively, and || represents taking the absolute value.

10. A self-supervised no-reference image quality assessment system, characterized in that, The no-reference image quality assessment system includes the following modules: A model construction module configured to construct a self-supervised no-reference image quality assessment model, the model including a prior knowledge learning sub-network and an image quality assessment sub-network sharing an encoder, and a knowledge transfer module is provided between the decoder of the prior knowledge learning sub-network and the decoder of the image quality assessment sub-network; A pre-training module configured to use the image restoration task as a proxy task for the image quality assessment task, construct a first training set for the image restoration task, and perform pre-training of the prior knowledge learning sub-network through the first training set; A knowledge transfer module configured to transfer the parameters of the decoder of the pre-trained prior knowledge learning sub-network to the decoder of the image quality assessment sub-network through a knowledge transfer channel; A fine-tuning training module configured to obtain a small amount of the second training set of the image quality assessment task, and perform fine-tuning training of the pre-trained prior knowledge learning sub-network and the image quality assessment sub-network after parameter transfer through the second training set; An image quality assessment module configured to use the fine-tuned prior knowledge learning sub-network and image quality assessment sub-network as a trained self-supervised no-reference image quality assessment model, and perform quality assessment of no-reference images through the trained self-supervised no-reference image quality assessment model.

Citation Information

Patent Citations

  • Meta-learning-based no-reference image quality data processing method and intelligent terminal

    CN110728656A

  • Non-reference image quality evaluation method based on deep feature transfer learning

    CN113421237A