Image Quality Assessment Method Based on Deep Mutual Learning and Dual-Scale Feature Fusion

Through deep mutual learning and dual-scale feature fusion methods, the local and non-local features of images are constrained by the Resnet50 and Vision Transformer networks to achieve consistency constraints on the local and non-local features of the image, solving the problem of low accuracy in image quality evaluation in the prior art, and achieving higher accuracy image quality evaluation.

CN115375663BActive Publication Date: 2025-07-11GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211038963.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-11
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

The existing reference-free image quality evaluation method extracted by a single CNN network cannot well reflect the global quality characteristics in the original image that play an important role in human eye perception, resulting in low evaluation accuracy.

Method used

Using deep mutual learning and dual-scale feature fusion methods, Resnet50 and Vision Transformer networks are used to constrain the local and non-local features of the image consistently. By building an initial quality evaluation model, feature fusion is combined with deep learning technology to improve the evaluation accuracy.

Benefits of technology

It enhances the feature extraction capability of the network, reduces the prediction deviation caused by the image after horizontal flip, can more accurately reflect the global quality characteristics in the original image that play an important role in human eye perception, and improves the evaluation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375663B_ABST
    Figure CN115375663B_ABST
Patent Text Reader

Abstract

A no-reference image quality assessment method based on deep mutual learning and dual-scale feature fusion provided by an embodiment of the present application. The method includes determining a target distorted image to be subject to no-reference image quality assessment; performing horizontal flipping on the target distorted image to obtain a target mirror image; constructing an initial quality assessment model, where the initial quality assessment model includes first and second Resnet50 networks for extracting local features from the image, and first and second VisionTransformer networks for extracting non-local features from the image; inputting the target distorted image into the first Resnet50 network and the first VisionTransformer network, and inputting the target mirror image into the second Resnet50 network and the second Vision Transformer network for model training. During the training process, consistency constraints are imposed on the local features and non-local features between images through the method of deep mutual learning, and the model output result is determined by fusing the local and non-local features of the image; when the model training is completed, a target quality assessment model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and more particularly, to an image quality evaluation method based on deep mutual learning and dual-scale feature fusion. Background Art

[0002] Images play an important role in people's daily work, entertainment, social applications, etc. Accurately evaluating image quality can not only provide necessary assistance for other tasks in computer vision, but also bring a comfortable visual experience to people in the Internet era and promote the prosperity of the Internet economy.

[0003] Currently, image quality assessment is divided into subjective and objective evaluations. Among them, the objective evaluation method of no-reference image quality assessment that does not require a reference image has received more attention from researchers due to the wide range of its application scope and application prospects.

[0004] Existing no-reference objective evaluation methods evaluate image quality by combining convolutional neural networks (CNNs), and this technology has also shown good performance currently. However, the way of predicting the image quality score through the regression task based on the image quality perception features extracted by a single CNN network makes the extracted image quality perception features unable to well reflect the global quality features that play an important role in human eye perception in the original image, resulting in the problem of low evaluation accuracy. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide an image quality evaluation method based on deep mutual learning and dual-scale feature fusion, which can improve the evaluation accuracy of image quality.

[0006] The embodiments of the present application also provide a no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion, including the following steps:

[0007] Determine a target distorted image to be subject to no-reference image quality assessment;

[0008] Horizontally flip the target distorted image to obtain a corresponding target mirror image;

[0009] Construct an initial quality evaluation model, where the initial quality evaluation model includes first and second Resnet50 networks for extracting local features from an image, and first and second VisionTransformer networks for extracting non-local features from the image;

[0010] Input the target distorted image into the first Resnet50 network and the first Vision Transformer network, and input the target mirror image into the second Resnet50 network and the second Vision Transformer network for model training. During the training process, perform consistency constraints on the local features and non-local features between images through the method of deep mutual learning to improve the evaluation accuracy, and determine the model output result by fusing the local and non-local features of the images;

[0011] When the model training is completed, obtain the target quality evaluation model, and input the image to be evaluated into the target quality evaluation model to obtain the predicted quality evaluation score of the image to be evaluated.

[0012] In a second aspect, an embodiment of the present application further provides a reference-free image quality evaluation system based on deep mutual learning and dual-scale feature fusion. The system includes an image acquisition module, a mirror image processing module, a model construction module, a model training module, and a quality evaluation module, where:

[0013] The image acquisition module is used to determine the target distorted image to be subjected to reference-free image quality evaluation;

[0014] The mirror image processing module is used to horizontally flip the target distorted image to obtain the corresponding target mirror image;

[0015] The model construction module is used to construct an initial quality evaluation model, and the initial quality evaluation model includes the first and second Resnet50 networks for extracting local features from images, and the first and second Vision Transformer networks for extracting non-local features from images;

[0016] The model training module is used to input the target distorted image into the first Resnet50 network and the first Vision Transformer network, and input the target mirror image into the second Resnet50 network and the second Vision Transformer network for model training. During the training process, perform consistency constraints on the local features and non-local features between images through the method of deep mutual learning to improve the evaluation accuracy, and determine the model output result by fusing the local and non-local features of the images;

[0017] The quality evaluation module is used to obtain the target quality evaluation model when the model training is completed, and input the image to be evaluated into the target quality evaluation model to obtain the predicted quality evaluation score of the image to be evaluated.

[0018] In a third aspect, an embodiment of the present application further provides a readable storage medium, which includes a program for a no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion. When the program for the no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion is executed by a processor, the steps of a no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion as described in any one of the above are implemented.

[0019] As can be seen from the above, a no-reference image quality evaluation method, system, and readable storage medium provided by an embodiment of the present application, on the one hand, fully consider the self-consistency of the image, that is, the original image and the mirror image obtained after its horizontal flipping should be the same for the human visual system, and these two versions of the image should have the same evaluation score. By means of deep mutual learning, consistency constraints are imposed on the local features and non-local features between the original image and the mirror image. The self-consistency of the no-reference image is used to make up for the lack of a reference image when using the no-reference method, enhance the feature extraction ability of the network, and reduce the prediction deviation caused by horizontal flipping of the image. On the other hand, using the powerful deep learning technology with fitting ability, dual-scale feature fusion of the local features and non-local features of the image is carried out, so that the extracted image quality perception features can well reflect the global quality features that play an important role in human eye perception in the original image, further ensuring the evaluation accuracy.

[0020] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will be obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of a no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion provided by an embodiment of the present application;

[0023] Figure 2 It is a schematic structural diagram of a Resnet50 network;

[0024] Figure 3It is a schematic structural diagram of the Vision Transformer network;

[0025] Figure 4 It is a schematic structural diagram of the Transformer Encoder network;

[0026] Figure 5 It is a schematic structural diagram of a no-reference image quality evaluation system based on deep mutual learning and dual-scale feature fusion provided by an embodiment of the present application. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0028] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first" and "second" are only used for differential description and cannot be understood as indicating or implying relative importance.

[0029] Please refer to Figure 1 , Figure 1 is a flowchart of a no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion in some embodiments of the present application. The method includes the following steps:

[0030] Step S100, determining a target distorted image to be evaluated for no-reference image quality.

[0031] Specifically, before model training, an image dataset for no-reference image quality evaluation will be obtained first and divided into a training set and a test set. Among them, the training set is used for model training, and the test set is used to evaluate the performance of the model. In the current embodiment, the division ratio of the training set and the test set is not limited and can be flexibly adjusted in different embodiments.

[0032] It should be noted that the image dataset selected in the current embodiment is a large-scale natural distortion dataset containing real distortions. It contains a total of 10,073 images taken by more than a thousand different models of cameras. Among them, there are two types of image resolutions, namely 1024×768 and 512×384. In the current embodiment, the image dataset with a resolution of 1024×768 will be selected, and 80% of the images will be randomly selected to form a training set, and the remaining 20% of the images will be selected to form a test set.

[0033] Step S200: Horizontally flip the target distorted image to obtain the corresponding target mirror image.

[0034] Specifically, an image processing tool such as PS can be used to horizontally flip the target distorted image. Since horizontally flipping the image is not the core innovation point of this solution, the current embodiment will not elaborate on this too much.

[0035] Step S300: Construct an initial quality evaluation model. The initial quality evaluation model includes the first and second Resnet50 networks for extracting local features from the image, and the first and second VisionTransformer networks for extracting non-local features from the image.

[0036] Specifically, the initial quality evaluation model includes 2 Resnet50 networks and 2 Vision Transformer networks. Among them, the parameter initialization methods between the same type of networks are different, and specific details can be referred to in the following content.

[0037] Step S400: Input the target distorted image into the first Resnet50 network and the first VisionTransformer network, and input the target mirror image into the second Resnet50 network and the second VisionTransformer network for model training. During the training process, the local features and non-local features between the images are constrained for consistency through the method of deep mutual learning to improve the evaluation accuracy, and the model output result is determined by fusing the local and non-local features of the image.

[0038] Specifically, the first Resnet50 network and the second Resnet50 network respectively take the original image (i.e., the target distorted image) and the mirror image obtained after its horizontal flipping (i.e., the target mirror image) as network inputs, and the outputs obtained are the local feature representations of the original image and the local feature representations of the mirror image. The first Vision Transformer network and the second Vision Transformer network respectively take the original image and the mirror image obtained after its horizontal flipping as network inputs, and the outputs obtained are the non-local feature representations of the original image and the non-local feature representations of the mirror image.

[0039] In the current embodiment, fully considering the self-consistency of the image, deep mutual learning is performed on the obtained various local features and non-local features, and consistency constraints are imposed during the training process of the model, reducing the prediction deviation caused by horizontal flipping of the image and improving the accuracy of image quality assessment.

[0040] Finally, when performing image quality assessment, average pooling operations will be first performed on the local features and non-local features of the input image respectively, and then the obtained pooling results will be concatenated together and output via a preset fully connected layer.

[0041] In one of the embodiments, the above image quality assessment process can be represented by the following formula:

[0042]

[0043] Among them, Score represents the predicted image quality evaluation score, represents the local features of the input image, represents the non-local features of the input image, AvgPool(*) represents the average pooling operation, Concat(*) represents the concatenation operation, and FC represents the preset fully connected layer.

[0044] Step S500, when ending the model training, obtain the target quality evaluation model, input the image to be evaluated into the target quality evaluation model, and obtain the predicted quality evaluation score of the image to be evaluated.

[0045] Specifically, the Pearson linear correlation coefficient (PLCC) and the Spearman rank correlation coefficient (SROCC) can be used to evaluate the performance of the network model. Among them, the Pearson linear correlation coefficient (PLCC) evaluates the performance from the perspective of the accuracy of the objective quality evaluation method prediction. The Spearman rank correlation coefficient (SROCC) focuses on measuring the monotonic consistency of the objective quality prediction scores of the images, and the value of it only relates to the sorting result of the images in the prediction result, reducing the consideration of the relative distance between the predicted value and the true value.

[0046] It should be noted that in the image quality assessment task, the closer the values of the Pearson linear correlation coefficient (PLCC) and the Spearman rank correlation coefficient (SROCC) are to 1, the better the effect of the network model. In the current embodiment, the calculation formulas of the Pearson linear correlation coefficient (PLCC) and the Spearman rank correlation coefficient (SROCC) are not limited.

[0047] As can be seen from the above, on the one hand, when fully considering the self-consistency of the image, that is, the original image and its mirror image obtained by horizontal flipping should be the same for the human visual system, and these two versions of the image should have the same evaluation score. By means of deep mutual learning, consistency constraints are imposed on the local features and non-local features between the original image and the mirror image. The self-consistency of the reference-free image is used to make up for the lack of the reference image when using the reference-free method, enhance the feature extraction ability of the network, and reduce the prediction deviation caused by horizontal flipping of the image. On the other hand, by using the powerful deep learning technology with strong fitting ability, dual-scale feature fusion of local features and non-local features of the image is carried out, so that the extracted image quality perception features can well reflect the global quality features that play an important role in human eye perception in the original image, further ensuring the evaluation accuracy.

[0048] In one embodiment, in step S100, determining the target processing object to be subjected to reference-free image quality assessment includes:

[0049] Step S1000, obtaining the initial processing object to be subjected to reference-free image quality assessment, and preprocessing the initial processing object according to a preset data processing method to obtain the corresponding target processing object, where the data processing method includes at least one of an image denoising method for smoothing and removing random noise in the image, a resolution unification method for randomly shearing the image to adjust the resolution of the image to a preset size, an image restoration method for filling in the missing parts of the image, and an image enhancement method for strengthening or suppressing key information in the image to improve the visual effect of the image.

[0050] Specifically:

[0051] (1) Common image denoising methods include Gaussian filtering (scanning each pixel in the image with a template (or convolution, mask), and replacing the value of the pixel at the center of the template with the weighted average gray value of the pixels in the neighborhood determined by the template), median filtering (sorting all the pixels in the neighborhood and then taking the median value as the pixel at the center of the neighborhood), etc. The embodiments of the present application do not limit them.

[0052] (2) When performing resolution unification processing, randomly shear the processed image to adjust the image resolution to 224×224, so that the processed image can meet the input specifications of the Resnet50 network and the Vision Transformer network, and retain more information useful for quality evaluation in the original image, improving the evaluation accuracy.

[0053] (3) Image restoration, that is, using the prior knowledge of the degradation process to restore the original appearance of the degraded image. Among them, the basic idea of image restoration includes first establishing an image restoration model, and then fitting the degraded image according to this model.

[0054] In one embodiment, the image restoration model can be processed by continuous mathematics and discrete mathematics, and the implementation of the processing term can be in the spatial domain convolution or in the frequency domain multiplication.

[0055] (4) Image enhancement is to make the original unclear image clear, or emphasize some interesting features, suppress the uninteresting features, so as to improve the image quality, enrich the information volume, and strengthen the image interpretation and recognition effect.

[0056] It should be noted that the method of image enhancement is to attach some information or transform data to the original image by certain means, selectively highlight the interesting features in the image or suppress (cover) some unnecessary features in the image, so that the image matches the visual response characteristics.

[0057] Through the above embodiments, by preprocessing the training images, noise in the training images can be avoided and the visual effect of the images can be improved, effectively improving the image quality.

[0058] In one embodiment, please refer to Figure 2 , the Resnet50 network includes an initial convolutional layer, a max pooling layer, and a residual network composed of 4 residual block layers connected in sequence, where: the convolutional kernel size of the initial convolutional layer is 7×7×64, and the stride is 2, which is used to perform convolutional operations on the input image to convert it into a 2D target feature vector; the max pooling layer is used to perform feature dimensionality reduction processing on the basis of the target feature vector while ensuring the features remain unchanged to retain the significant features of the image; the residual network is used to increase the depth considerably to improve the accuracy of feature extraction.

[0059] Through the above embodiments, the advantages of CNN in perceiving local features of images are fully utilized, the learning of local features of images is strengthened, and the performance of the network in image quality evaluation is improved.

[0060] In one embodiment, please refer to Figure 2, the first residual block layer contains a first convolutional kernel group composed of 3 sequentially arranged convolutional kernels. The size of the first convolutional kernel group is [1×1×64, 3×3×64, 1×1×256]. Among them, the number of channels of the first feature vector output via the first residual block layer is 256; the second residual block layer contains a second convolutional kernel group composed of 4 sequentially arranged convolutional kernels. The size of the second convolutional kernel group is [1×1×128, 3×3×128, 1×1×512]. Among them, the number of channels of the second feature vector output via the second residual block layer is 512; the third residual block layer contains a third convolutional kernel group composed of 6 sequentially arranged convolutional kernels. The size of the third convolutional kernel group is [1×1×256, 3×3×256, 1×1×1024]. Among them, the number of channels of the third feature vector output via the third residual block layer is 1024; the fourth residual block layer contains a fourth convolutional kernel group composed of 3 sequentially arranged convolutional kernels. The size of the fourth convolutional kernel group is [1×1×512, 3×3×512, 1×1×2048]. Among them, the number of channels of the fourth feature vector output via the fourth residual block layer is 2048.

[0061] In the above embodiments, combined with the deep residual network, shallow and deep features can be effectively extracted from the image. By increasing the depth considerably, the accuracy of feature extraction can be improved.

[0062] In one of the embodiments, please refer to Figure 3 , the Vision Transformer network includes a Patch Embedding network and a Transformer Encoder network, where: the convolutional kernel size of the Patch Embedding network is 8×8, and the convolutional stride is 8, which is used to perform a convolutional operation on the input image to convert it into a 2D feature vector where N = HW / P 2 represents the final number of blocks, which also serves as the effective input sequence length of the Transformer Encoder network. (H, W) represents the resolution of the input image, and (P, P) represents the resolution of each image patch in the input image.

[0063] Specifically, the total number of feature maps of the Patch Embedding network is 768. In the actual application process, the Patch Embedding network will convert the input image into a 2D feature vector through convolutional operation, and map the total number of its features to the size constantly used in the Transformer Encoder network, that is, 768 dimensions, so as to meet the input requirements of the Transformer Encoder network.

[0064] Please refer to Figure 4, the Transformer Encoder network includes a Layer Norm layer, a Multi-Head Attention layer, and an MLP layer connected in sequence, where: the calculation formula of the Multi-Head Attention layer includes:

[0065]

[0066]

[0067] Among them, Q represents the query matrix corresponding to the input vector, K represents the preset key matrix, V represents the preset value matrix, d k represents the dimension of the input vector, T represents the transpose of the matrix; softmax(*) represents the activation function; head i represents the i-th head, W1 represents a learnable weight matrix, and Concat(*) represents the concatenation operation.

[0068] It should be noted that multi-head attention is to establish different projection information in multiple different projection spaces. It will project the input matrix in different directions, and after obtaining the corresponding output matrices, splice them together.

[0069] This process is similar to integration. Among them, the difference between multi-head and single-head is that multiple single-heads are replicated, but the weight coefficients involved will be different. It can be analogized to a neural network model and multiple identical neural network models, but due to different initializations, the weights will be different. head i represents the i-th head, which can be set to 12.

[0070] Specifically, the Layer Norm layer is used to perform layer normalization operations on the feature vectors. The calculation formula of the attention module in the Multi-HeadAttention (MSA) layer is the above-mentioned Attention(Q, K, V), and its specific calculation form can refer to the above formula. The calculation formula of the activation function softmax is Among them, e j represents the exponential value obtained by the j-th component, and e i represents the exponential value obtained by the i-th component.

[0071] In one of the embodiments, the Transformer Encoder network is represented by the following formula:

[0072]

[0073] Among them, z0 represents the feature vector obtained by processing through the Patch Embedding network plus the class encoding xclass and the result obtained after position encoding E pos the obtained result; represents the first feature slice obtained after the input image is processed by the Patch Embedding network; LN(*) represents the Layer Norm layer, MSA(*) represents the Multi-Head Attention layer, and MLP(*) represents the MLP layer; z t-1 represents the output feature output by the (t - 1)-th layer in the Transformer Encoder network, z′ t represents the intermediate output feature output by the t-th layer in the Transformer Encoder network, z t represents the final output feature output by the t-th layer in the Transformer Encoder network; L represents the depth of the Transformer Encoder network, represents the output obtained after processing the class encoding as a feature through the Transformer Encoder network.

[0074] In the above embodiments, combined with the advantage of the Vision Transformer network in perceiving the non-local features of the image, the learning of the non-local features of the image is strengthened, and the performance of the network in image quality assessment is improved.

[0075] In one of the embodiments, during the training process, the method further includes: performing model constraint based on a pre-constructed consistency loss function and a mean square error loss function, where the consistency loss function and the mean square error loss function are represented by the following formulas:

[0076]

[0077]

[0078]

[0079] In the above formula, L1 represents the first loss function used by the entire network composed of the first Resnet50 network and the first Vision Transformer network, L2 represents the second loss function used by the entire network composed of the second Resnet50 network and the second Vision Transformer network; s represents the model output result, g represents the benchmark result, and B represents the size of a batch during training; L mse represents the mean square error loss function, L con represents the consistency loss function; represents the first local feature output by the first Resnet50 network, represents the second local feature output via the second Resnet50 network; represents the first non-local feature output via the first Vision Transformer network, represents the second non-local feature output via the second Vision Transformer network; represents the two-norm, where f1 and f2 represent the same type but different local or non-local features in the loss function L con in the same kind but different local or non-local features.

[0080] Please refer to Figure 5 , which is a reference-free image quality assessment system based on deep mutual learning and dual-scale feature fusion. The system 500 includes an image acquisition module 501, a mirror image processing module 502, a model construction module 503, a model training module 504, and a quality assessment module 505, where:

[0081] The image acquisition module 501 is used to determine the target distorted image to be subjected to reference-free image quality assessment.

[0082] The mirror image processing module 502 is used to horizontally flip the target distorted image to obtain the corresponding target mirror image.

[0083] The model construction module 503 is used to construct an initial quality assessment model. The initial quality assessment model includes the first and second Resnet50 networks for extracting local features from images, and the first and second Vision Transformer networks for extracting non-local features from images.

[0084] The model training module 504 is used to input the target distorted image into the first Resnet50 network and the first Vision Transformer network, and input the target mirror image into the second Resnet50 network and the second Vision Transformer network for model training. During the training process, the local features and non-local features between images are constrained for consistency through deep mutual learning to improve the evaluation accuracy, and the local and non-local features of the image are fused to determine the model output result.

[0085] The quality assessment module 505 is used to obtain the target quality assessment model when the model training ends, and input the image to be evaluated into the target quality assessment model to obtain the predicted quality assessment score of the image to be evaluated.

[0086] In one embodiment, the above modules are further used to implement the methods in any optional implementation manner of the above embodiment, and the embodiments of the present application do not make any limitations in this regard.

[0087] As can be seen from the above, a no-reference image quality evaluation system based on deep mutual learning and dual-scale feature fusion disclosed in the present application, on the one hand, fully considers the self-consistency of the image, that is, the original image and the mirror image obtained after its horizontal flipping should be the same for the human visual system, and these two versions of the image should have the same evaluation score. Through the method of deep mutual learning, consistency constraints are imposed on the local features and non-local features between the original image and the mirror image, and the lack of a reference image when using the no-reference method is compensated by the self-consistency of the no-reference image, enhancing the feature extraction ability of the network and reducing the prediction deviation caused by horizontal flipping of the image. On the other hand, by using the powerful deep learning technology with strong fitting ability, dual-scale feature fusion of local features and non-local features of the image is carried out, so that the extracted image quality perception features can well reflect the global quality features that play an important role in human eye perception in the original image, further ensuring the evaluation accuracy.

[0088] The embodiment of the present application provides a readable storage medium. When the computer program is executed by a processor, it executes the method in any optional implementation manner of the above embodiment. Among them, the readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, a magnetic disk or an optical disc.

[0089] On the one hand, when fully considering the self-consistency of an image, that is, the original image and its mirror image obtained by horizontal flipping should be the same for the human visual system, and these two versions of the image should have the same evaluation score. Through the method of deep mutual learning, consistency constraints are imposed on the local features and non-local features between the original image and the mirror image. Utilizing the self-consistency of the reference-free image makes up for the lack of a reference image when using the reference-free method, enhances the feature extraction ability of the network, and reduces the prediction deviation caused by horizontal flipping of the image. On the other hand, by using the powerful deep learning technology with strong fitting ability, dual-scale feature fusion of the local features and non-local features of the image is carried out, so that the extracted image quality perception features can well reflect the global quality features that play an important role in human eye perception in the original image, further ensuring the evaluation accuracy.

[0090] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0091] In addition, the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0092] Furthermore, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0093] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0094] The above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A no-reference image quality assessment method based on deep mutual learning and dual-scale feature fusion, characterized in that, Including the following steps: Determine a target distorted image to be subject to reference-free image quality assessment; Horizontally flip the target distorted image to obtain a corresponding target mirror image; Construct an initial quality evaluation model, where the initial quality evaluation model includes first and second Resnet50 networks for extracting local features from an image, and first and second Vision Transformer networks for extracting non-local features from an image; Input the target distorted image into the first Resnet50 network and the first Vision Transformer network, and input the target mirror image into the second Resnet50 network and the second Vision Transformer network for model training. During the training process, consistency constraints are imposed on the local features and non-local features between images through the method of deep mutual learning to improve the evaluation accuracy, and the model output result is determined by fusing the local and non-local features of the image; When the model training is completed, obtain a target quality evaluation model, input the image to be evaluated into the target quality evaluation model, and obtain the predicted quality evaluation score of the image to be evaluated; During the training process, the method further includes: Conduct model constraints based on a pre-constructed consistency loss function and a mean square error loss function, where the consistency loss function and the mean square error loss function are represented by the following formulas: ; ; ; In the above formula, L 1 represents the first loss function used by the entire network composed of the first Resnet50 network and the first Vision Transformer network, L 2 represents the second loss function used by the entire network composed of the second Resnet50 network and the second Vision Transformer network; s represents the model output result, g represents the benchmark result, B represents the size of a batch during training; L mse represents the mean squared error loss function, L con represents the consistency loss function; represents the first local feature output by the first Resnet50 network, represents the second local feature output by the second Resnet50 network; represents the first non-local feature output by the first Vision Transformer network, represents the second non-local feature output by the second Vision Transformer network; represents the two-norm, f 1 and f 2 in the loss function L con represent the same type but different local or non-local features.

2. The method according to claim 1, wherein The determination of the target processing object to be subject to reference-free image quality assessment includes: Obtain an initial processing object to be subject to reference-free image quality assessment, and preprocess the initial processing object according to a preset data processing method to obtain a corresponding target processing object, where the data processing method includes at least one of an image denoising method for smoothing and removing random noise in the image, a resolution unification method for randomly shearing the image to adjust the resolution of the image to a preset size, an image restoration method for complementing missing parts in the image, and an image enhancement method for strengthening or suppressing key information in the image to improve the visual effect of the image.

3. The method according to claim 1, wherein The Resnet50 network includes an initial convolutional layer, a max pooling layer, and a residual network composed of 4 residual block layers connected in sequence, where: The convolutional kernel size of the initial convolutional layer is 7×7×64, and the stride is 2, which is used to perform a convolutional operation on the input image to convert it into a 2D target feature vector; The max pooling layer is used to perform feature dimensionality reduction processing on the basis of the target feature vector while ensuring the features remain unchanged to retain the significant features of the image; The residual network is used to increase the depth accordingly to improve the accuracy of feature extraction.

4. The method according to claim 3, wherein The first residual block layer contains a first convolutional kernel group composed of 3 convolutional kernels arranged in sequence, and the size of the first convolutional kernel group is [1×1×64, 3×3×64, 1×1×256], where the number of channels of the first feature vector output by the first residual block layer is 256; The second residual block layer contains a second convolutional kernel group composed of 4 sequentially arranged convolutional kernels. The size of the second convolutional kernel group is [1×1×128, 3×3×128, 1×1×512]. Among them, the number of channels of the second feature vector output via the second residual block layer is 512; The third residual block layer contains a third convolutional kernel group composed of 6 sequentially arranged convolutional kernels. The size of the third convolutional kernel group is [1×1×256, 3×3×256, 1×1×1024]. Among them, the number of channels of the third feature vector output via the third residual block layer is 1024; The fourth residual block layer contains a fourth convolutional kernel group composed of 3 sequentially arranged convolutional kernels. The size of the fourth convolutional kernel group is [1×1×512, 3×3×512, 1×1×2048]. Among them, the number of channels of the fourth feature vector output via the fourth residual block layer is 2048.

5. The method according to claim 1, characterized in that The Vision Transformer network includes a PatchEmbedding network and a Transformer Encoder network, where: The convolution kernel size of the Patch Embedding network is 8×8, and the convolution stride is 8, which is used to perform convolution operations on the input image to convert it into a 2D feature vector , where represents the final number of blocks, which is also used as the effective input sequence length of the Transformer Encoder network, ( H , W ) represents the resolution of the input image, ( P , P ) represents the resolution of each image patch in the input image; The Transformer Encoder network includes a Layer Norm layer, a Multi-HeadAttention layer, and an MLP layer connected in sequence, where: The calculation formula of the Multi-Head Attention layer includes: ; ; Among them, Q represents the query matrix corresponding to the input vector, K represents the preset key matrix, V represents the preset value matrix, d k represents the dimension of the input vector, T represents the transpose of the matrix; softmax (*) represents the activation function; head i represents the i th head, W 1 represents a learnable weight matrix, Concat (*) represents the connection operation.

6. The method according to claim 5, wherein The Transformer Encoder network is represented by the following formula: ; Among them, z 0 represents the result obtained after adding the feature vector processed by the Patch Embedding network plus class encoding x class , and after position encoding E pos ; represents the first feature slice obtained after processing the input image through the Patch Embedding network; LN (*) represents the Layer Norm layer, MSA (*) represents the Multi-Head Attention layer, MLP (*) represents the MLP layer; z t-1 represents the output feature output by the t -1 layer in the Transformer Encoder network, represents the intermediate output feature output by the t layer in the Transformer Encoder network, z t represents the final output feature output by the t layer in the Transformer Encoder network; L represents the depth of the Transformer Encoder network, represents class the output obtained after processing the encoding as a feature through the Transformer Encoder network.

7. A no-reference image quality assessment system based on deep mutual learning and dual-scale feature fusion, characterized in that, The system includes an image acquisition module, a mirror image processing module, a model construction module, a model training module, and a quality evaluation module, where: The image acquisition module is used to determine a target distorted image to be evaluated for no-reference image quality; The mirror image processing module is used to horizontally flip the target distorted image to obtain a corresponding target mirror image; The model construction module is used to construct an initial quality evaluation model. The initial quality evaluation model includes the first and second Resnet50 networks for extracting local features from images, and the first and second Vision Transformer networks for extracting non-local features from images; The model training module is used to input the target distorted image into the first Resnet50 network and the first Vision Transformer network, and input the target mirror image into the second Resnet50 network and the second Vision Transformer network for model training. During the training process, the local features and non-local features between images are constrained for consistency through the method of deep mutual learning to improve the evaluation accuracy, and the model output result is determined by fusing the local and non-local features of the images; The quality evaluation module is used to obtain a target quality evaluation model when the model training is completed, input the image to be evaluated into the target quality evaluation model, and obtain the predicted quality evaluation score of the image to be evaluated; During the training process, it also includes: Model constraint is performed based on a pre-constructed consistency loss function and a mean square error loss function, where the consistency loss function and the mean square error loss function are represented by the following formulas: ; ; ; In the above formula, L 1 represents the first loss function used by the entire network composed of the first Resnet50 network and the first Vision Transformer network, L 2 represents the second loss function used by the entire network composed of the second Resnet50 network and the second Vision Transformer network; s represents the model output result, g represents the reference result, B represents the size of a batch during training; L mse represents the mean squared error loss function, L con represents the consistency loss function; represents the first local feature output by the first Resnet50 network, represents the second local feature output by the second Resnet50 network; represents the first non-local feature output by the first Vision Transformer network, represents the second non-local feature output by the second Vision Transformer network; represents the second norm, f 1 and f 2 in the loss function L con represent the same type but different local or non-local features.

8. A readable storage medium, characterized in that, The readable storage medium includes a program for a no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion. When the program for the no-reference image quality evaluation method based on deep mutual learning and dual-scale feature fusion is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Optical remote sensing spatial spectrum fusion method, device and equipment without reference image and medium

    CN114581347A

  • Indoor scene monocular image depth estimation method based on deep learning

    CN114638870A