Image quality assessment method, device, electronic device and machine-readable storage medium
Through the feature extraction and score regression layer of the image quality assessment model, the problem of accurate quantitative assessment of various types of image distortion is solved, and the reference value and efficiency of image quality assessment are improved.
Patent Information
- Application Number
- CN202111574337.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2041-12-21
AI Technical Summary
Existing technologies have difficulty accurately judging and quantifying the various types of distortion in images, resulting in insufficient reference value for image quality assessment and a complex assessment process.
An image quality assessment model is adopted, including a feature extraction layer and a score regression layer. Image features are extracted and the evaluation scores of various distortion types are quantified in an end-to-end manner. The feature enhancement layer is used to improve the evaluation accuracy and simplify the evaluation process.
It achieves accurate judgment and quantitative evaluation of various types of image distortion, improves the reference value and efficiency of image quality assessment, and simplifies the processing flow.
Smart Images

Figure CN114299358B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image quality assessment method, device, electronic device, and machine-readable storage medium. Background Art
[0002] In real-world scenarios, images can exhibit significant variations in image quality due to various external factors (such as occlusion, interference, and lighting variations) as well as image acquisition constraints (such as camera resolution and camera angle). Low-quality images not only consume significant storage resources but can also introduce erroneous information, resulting in unnecessary data loss. Therefore, image quality assessment has become an increasingly important research area. Summary of the Invention
[0003] In view of this, the present application provides an image quality assessment method, device, electronic device and machine-readable storage medium to improve the reference value of image quality assessment.
[0004] Specifically, this application is implemented through the following technical solutions:
[0005] According to a first aspect of an embodiment of the present application, there is provided an image quality assessment method, comprising:
[0006] Inputting the image to be evaluated into the image quality assessment model to obtain evaluation scores of N different distortion types of the image to be evaluated;
[0007] The image quality assessment model includes a feature extraction layer and a score regression layer. The feature extraction layer is used to extract image features of the input image; the score regression layer is used to obtain evaluation scores of N different distortion types of the input image based on the image features; N≥2, and N is an integer.
[0008] According to a second aspect of an embodiment of the present application, there is provided an image quality assessment apparatus, comprising:
[0009] an acquisition unit, configured to acquire an image to be evaluated;
[0010] An evaluation unit is configured to input the image to be evaluated into an image quality assessment model to obtain evaluation scores for N different distortion types of the image to be evaluated; wherein the image quality assessment model includes a feature extraction layer and a score regression layer, the feature extraction layer is configured to extract image features of the input image; the score regression layer is configured to obtain evaluation scores for N different distortion types of the input image based on the image features; N ≥ 2, and N is an integer.
[0011] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the image quality assessment method provided in the first aspect.
[0012] According to a fourth aspect of an embodiment of the present application, a machine-readable storage medium is provided, wherein the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the image quality assessment method provided in the first aspect is implemented.
[0013] The technical solution provided by this application can at least bring the following beneficial effects:
[0014] By inputting the image to be evaluated into the image quality assessment model, the feature extraction layer of the image quality assessment model extracts the image features of the image, and the score regression layer of the image quality assessment model obtains the evaluation scores of N different distortion types of the input image based on the image features, not only can the distortion type contained in the image be accurately judged, but each distortion type can also be quantitatively evaluated separately, thereby improving the reference value of image quality assessment; in addition, after the image is input into the image instruction assessment model, the evaluation scores of N different distortion types of the image are obtained in an end-to-end manner, without the need for other post-processing, thereby simplifying the processing flow of image quality assessment and improving the efficiency of image quality assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flowchart of an image quality assessment method shown in an exemplary embodiment of the present application;
[0016] Figure 2 is a module diagram of an image quality assessment model shown in an exemplary embodiment of the present application;
[0017] Figure 3 is a schematic diagram of an image quality assessment process shown in an exemplary embodiment of the present application;
[0018] Figure 4 is a structural diagram of an image quality assessment device shown in an exemplary embodiment of the present application;
[0019] Figure 5 is a structural diagram of another image quality assessment device shown in an exemplary embodiment of the present application;
[0020] Figure 6 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0021] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0022] The terms used in this application are for the purpose of describing particular embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0023] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.
[0024] See Figure 1 , is a flow chart of an image quality assessment method provided in an embodiment of the present application, such as Figure 1 As shown, the image quality assessment method may include the following steps:
[0025] Step S100: Input the image to be evaluated into the image quality assessment model to obtain evaluation scores of N different distortion types of the image to be evaluated.
[0026] Among them, the image quality assessment model includes a feature extraction layer and a score regression layer. The feature extraction layer is used to extract image features of the input image; the score regression layer is used to obtain evaluation scores of N different distortion types of the input image based on the image features; N ≥ 2, and N is an integer.
[0027] In the embodiments of the present application, it is taken into consideration that there may be many different types of distortion in actual images, such as overexposure of the image due to an abnormality in the image acquisition process, blurred images due to excessive compression during image transmission and storage, etc., and that different types of distortion may have different effects on the image processing results in different usage scenarios.
[0028] For example, in the license plate recognition scenario, "occlusion" has a greater impact on the recognition results than "overexposure" and "too dark".
[0029] Therefore, in order to improve the reference value of the image quality assessment results, it is no longer limited to evaluating the overall quality of the image, nor is it limited to detecting the distortion type of the image. Instead, the distortion type included in the image can be determined, and each distortion type can be quantitatively evaluated separately to obtain an assessment score for each distortion type.
[0030] Accordingly, an image to be evaluated can be obtained and input into an image quality assessment model, the feature extraction layer of the image quality assessment model extracts the image features of the image, and the score regression layer of the image quality assessment model (which can be called an N-branch score regression layer) obtains the evaluation scores of N different distortion types of the input image (i.e., the image to be evaluated) based on the image features.
[0031] For example, for any distortion type, a lower evaluation score of the distortion type indicates that the distortion of the input image is more serious.
[0032] In one example, the fractional regression model may include N fully connected layers, each fully connected layer corresponding to one distortion type.
[0033] In another example, the fractional regression model may include a fully connected layer including N outputs, each output corresponding to a distortion type.
[0034] It should be noted that score regression can also be implemented using convolutional layer + global average pooling.
[0035] It can be seen that in Figure 1 In the method flow shown, the image to be evaluated is input into the image quality assessment model, the feature extraction layer of the image quality assessment model extracts the image features of the image, and the score regression layer of the image quality assessment model obtains the evaluation scores of N different distortion types of the input image based on the image features. This method can not only accurately determine the distortion type contained in the image, but also quantitatively evaluate each distortion type separately, thereby improving the reference value of image quality assessment. In addition, after the image is input into the image instruction assessment model, the evaluation scores of N different distortion types of the image are obtained in an end-to-end manner, without the need for other post-processing, thereby simplifying the processing flow of image quality assessment and improving the efficiency of image quality assessment.
[0036] In some embodiments, the above-mentioned image instruction evaluation model may also include a feature enhancement layer, which is used to perform feature enhancement processing on the image features extracted by the feature extraction layer to obtain enhanced features, and the score regression layer obtains evaluation scores of N different distortion types of the input image based on the enhanced feature image features.
[0037] To improve the accuracy of image quality assessment, the image quality assessment model can also include a feature enhancement layer. After the image to be assessed is input into the image quality assessment model, the feature extraction layer can extract features to obtain image features of the image to be assessed. These image features can then be enhanced by the feature enhancement layer to obtain enhanced features.
[0038] The enhanced features output by the feature enhancement layer can be input into the score regression layer, and the score regression layer obtains the evaluation scores of N different distortion types of the image to be evaluated based on the enhanced features.
[0039] Exemplarily, feature enhancement processing may include but is not limited to feature fusion processing and / or attention mechanism processing.
[0040] In one example, the feature enhancement process may include a feature fusion process;
[0041] The above feature enhancement layer implements feature enhancement processing in the following ways:
[0042] Process the output features of each layer of the feature extraction layer into the same number of channels, and obtain the output features of each layer after the first processing;
[0043] Upsampling the output features of each layer after the first processing is performed to obtain the output features of each layer after the second processing; wherein the feature maps of the output features of each layer after the second processing have the same size;
[0044] The output features after the second processing of each layer are subjected to feature fusion processing to obtain fused features.
[0045] For example, in order to improve the accuracy of image quality assessment, when image features are extracted from an input image through a feature extraction layer, feature fusion processing can be performed on the output features of each layer of the feature extraction layer to obtain fused features to achieve feature enhancement.
[0046] For example, after the input image is input into the image quality assessment model, the number of channels and feature map size of the output features of each layer of the feature extraction layer of the image quality assessment model may be different.
[0047] Accordingly, before performing feature fusion on the output features of each layer of the feature extraction layer, the output features of each layer may be processed to have the same number of channels and the same size of feature maps.
[0048] Exemplarily, after the input image is input into the image quality assessment model, the output features of each layer of the feature extraction layer can be obtained respectively, and the output features of each layer of the feature extraction layer can be processed to have the same number of channels to obtain the output features of each layer with the same number of channels (referred to as the output features after the first processing in this article).
[0049] For the output features of each layer after the first processing, upsampling can be performed to make the feature maps of each output feature the same size, thereby obtaining output features with the same feature map size and the same number of channels in each layer (referred to as the output features after the second processing in this article).
[0050] Exemplarily, the specific implementation of processing the output features of each layer of the feature extraction layer into output features with the same number of channels and the same feature map size can be explained below with reference to specific examples.
[0051] Furthermore, feature fusion processing can be performed on the output features of each layer after the second processing to obtain fused features.
[0052] In one example, the feature enhancement processing includes attention mechanism processing; the attention mechanism processing includes spatial attention mechanism processing and / or channel attention mechanism processing;
[0053] The above feature enhancement layer can achieve feature enhancement processing in the following ways:
[0054] Processing image features based on the attention mechanism to determine feature dependencies of image features in spatial dimensions and / or channel dimensions;
[0055] According to the feature dependencies of the image features in the spatial dimension and / or channel dimension, the image features are enhanced to obtain enhanced features.
[0056] For example, it is considered that the extracted image features of the input image usually have feature dependencies in the spatial dimension and the channel dimension.
[0057] Accordingly, in order to achieve feature enhancement, the image features may be enhanced based on the feature dependencies of the image features in the spatial dimension and / or channel dimension.
[0058] Exemplarily, the acquired image features can be processed according to the attention mechanism to determine the feature dependencies of the image features in the spatial dimension and / or channel dimension, and based on the feature dependencies of the image features in the spatial dimension and / or channel dimension, the image features can be enhanced to obtain enhanced features.
[0059] Taking the feature dependency in the spatial dimension as an example, the weighted weight of the feature can be determined based on the similarity of the features at different positions, and the features at all positions can be aggregated and updated by weighted summation to obtain enhanced features.
[0060] For example, the specific implementation of enhancing image features based on the feature dependencies of image features in the spatial dimension and / or channel dimension can be described below in conjunction with specific examples, and the embodiments of the present application will not be repeated here.
[0061] In some embodiments, before inputting the image to be evaluated into the image quality assessment network in step S100, the process further includes:
[0062] The model loss is determined based on the evaluation scores of N different distortion types output by the image quality assessment model after the training samples are input, and the calibration scores of the N different distortion types of the training samples. Based on the model loss, the model parameters of each part of the image quality assessment model are optimized and adjusted through back propagation.
[0063] Exemplarily, during the training process of the image quality assessment model, the model loss can be calculated based on the evaluation scores of N different distortion types output after the training samples are input into the image quality assessment model, and the calibration scores of the N different distortion types of the training samples. Based on the model loss, the model parameters of each part of the image quality assessment model can be optimized and adjusted through back propagation.
[0064] It should be noted that after completing the training of the image quality assessment model, before using the trained image quality assessment model to perform the image quality assessment task, the trained image quality assessment model can also be tested according to a preset test set, and when the trained image quality assessment model meets the test requirements, it can be used to perform the image quality assessment task; otherwise, the trained image quality assessment model can continue to be trained.
[0065] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, the technical solutions provided by the embodiments of the present application are described below in conjunction with specific embodiments.
[0066] In this embodiment, a labeled data set (i.e., a training sample set) can be used to train a deep neural network (which can be called an image multi-attribute quality assessment network, i.e., the above-mentioned image quality assessment model). The deep neural network can evaluate the quality of the image based on the distortion type contained in the image and output evaluation scores for multiple different distortion types.
[0067] In this embodiment, if Figure 2 As shown in Figure 1, the image quality assessment model mainly includes three functional modules: a feature extraction module (such as the above-mentioned feature extraction layer), a feature processing module (for performing feature enhancement processing on image features, i.e. the above-mentioned feature enhancement layer), and a score regression module (such as the above-mentioned N-branch score regression layer). Among them:
[0068] Feature extraction module: This module typically uses a deep convolutional neural network, such as VGG, ResNet, or DensNet. After image preprocessing, the image is first input into the feature extraction module, where image features are extracted through multiple layers of convolution.
[0069] Feature processing module: The image features extracted by the feature extraction module can be input into the feature processing module for feature enhancement processing, such as feature fusion processing, attention mechanism processing, etc., and then input into the score regression module.
[0070] Score regression module: The features output by the feature processing module are passed through the score regression module to obtain N score scalars, which represent the evaluation scores of N distortion types.
[0071] The following combination Figure 3 Describe each functional module.
[0072] 1. Feature Extraction Module
[0073] For example, taking the use of a residual network (such as ResNet18) as a feature extraction module as an example, the output features of each of its four sub-modules Layer_1, Layer_2, Layer_3, and Layer_4 can be used as the input of subsequent modules.
[0074] Since high-level features have stronger semantic information and less noise, the output features of the last layer are usually used for classification or detection in traditional classification and detection tasks.
[0075] However, in the embodiment of the present application, considering that multi-attribute image quality assessment (i.e., quality assessment of multiple different distortion types) focuses more on local details of the image, the requirements for semantics are not very high.
[0076] In addition, considering that the low-level features (output features of Layer_1 / Layer_2 / Layer_3) are less semantic and have more noise, they have higher resolution and contain more location and detail information.
[0077] Therefore, in the embodiment of the present application, the output features of each layer of the ResNet18 network can be extracted, and feature enhancement can be achieved through feature fusion processing.
[0078] It should be noted that feature extraction through the ResNet18 network is only a specific example of a feature extraction module, and is not a limitation on the scope of protection of this application. In the embodiments of this application, other ResNet series networks can also be used, or deep convolutional networks such as VGG and DenseNet can be used for feature extraction.
[0079] 2. Feature Processing Module
[0080] 2.1 Feature Fusion Processing Submodule
[0081] After completing the extraction of features of each layer, the output features of each layer can be fused.
[0082] For example, assuming that the vector dimension of the input image is 3x256x256, the feature map dimensions of the output features of each layer of the ResNet18 network are: the output of Layer_1 is 64x64x64, the output of Layer_2 is 128x32x32, the output of Layer_3 is 256x16x16, and the output of Layer_4 is 512x8x8.
[0083] It can be seen that the output features of each layer are different, so it is necessary to unify the feature map size and number of channels.
[0084] For example, the output features of Layer_2, Layer_3, and Layer_4 can each use 1x1 convolution to change the number of their channels to 64 (the number of channels of Layer_1 feature is 64 and does not need to be changed).
[0085] Then, starting from the output features of Layer_4, upsampling is performed layer by layer from top to bottom, and the feature map of the previous layer is added.
[0086] For example, each upsampling multiple is 2, and the nearest neighbor interpolation method is used; finally, a 3×3 convolution is used to reduce the impact of the upsampling process to obtain the final feature map.
[0087] After the above operations, the feature map dimensions corresponding to the output features of each layer are: Layer_1 (64x64x64), Layer_2 (64x32x32), Layer_3 (64x16x16), Layer_4 (64x8x8).
[0088] For example, since the feature maps of each layer have different sizes, they need to be unified.
[0089] For example, for high-level feature maps, the upsampling method is continued to be used, where the feature map corresponding to Layer_4 uses 3 times of 2-fold upsampling, the feature map corresponding to Layer_3 uses 2 times of upsampling, and the feature map corresponding to Layer_2 uses 1 time of 2-fold upsampling. At this time, the feature map dimension size corresponding to the output features of each layer is 64x64x64.
[0090] Finally, the features of each layer are added together, and then a 3x3 convolution is used to alleviate the upsampling effect to complete the feature fusion process.
[0091] 2.2 Attention Mechanism Processing Submodule
[0092] For example, a dual attention network can be used to capture feature dependencies in the spatial dimension and channel dimension respectively based on the self-attention mechanism.
[0093] Exemplarily, the attention mechanism processing submodule may include a spatial attention module (PositionAttention Module, referred to as PAM) and a channel attention module (ChannelAttention Module, referred to as CAM).
[0094] For example, PAM uses the self-attention mechanism to capture the spatial dependency between any two positions and aggregates and updates the features of all positions through weighted summation, where the weight is determined by the similarity between the corresponding two positions; CAM uses the self-attention mechanism to capture the dependency between any two channels and updates it through weighted summation; finally, the outputs of the two modules are added and fused to further enhance the feature representation.
[0095] For example, it is assumed that after the input image passes through the feature extraction module and the feature fusion submodule, a CxHxW feature map is obtained, and the feature map enters the dual attention network as input.
[0096] In PAM, we first need to get the spatial attention matrix M PAM . M PAM The dimension size is KxK (K=HxW), which is used to represent the dependency relationship between any two positions. Its matrix element M ij , M PAM ∈R N×N Represents the degree of association between the i-th position and the j-th position.
[0097] For example, the original feature map dimension can be adjusted from CxHxW to CxK (K=HxW), and the matrix B(CxK) is obtained, and then the matrix C(KxC) is transposed, and finally M is calculated. PAM =CB to obtain the spatial attention matrix.
[0098] After obtaining the attention matrix, it can be multiplied with the original feature map. Before multiplication, the feature map dimension needs to be adjusted to CxK. After multiplication, the weighted feature map (CxK) is obtained, which is then adjusted to CxHxW and added to the original feature map to obtain the final output.
[0099] For example, the calculation process of the channel self-attention module is basically similar, the main difference is that the channel attention matrix M CAM The dimension size is CxC, so we need to pay attention to the corresponding adjustment dimension during calculation.
[0100] 3. Score Regression Module
[0101] For example, let's take the implementation of score regression through a fully connected layer as an example.
[0102] It should be noted that score regression is not limited to the use of fully connected layers, but can also be implemented using convolutional layers + global average pooling.
[0103] Exemplarily, the score regression module may include N (N = number of distortion types) fully connected layers, each fully connected layer corresponding to a distortion type, or the score regression module may include one fully connected layer, the fully connected layer including N outputs, each output corresponding to a distortion type.
[0104] The input processed feature map will pass through each fully connected layer and output N quality scores S i , where i = 0, 1,…, N-1.
[0105] Among them, S i The lower the value, the more severe the distortion.
[0106] It should be noted that the image multi-attribute quality assessment network provided in the embodiment of the present application may include two stages before being used to perform image quality assessment tasks: a training stage and a testing stage; wherein:
[0107] 4.1 Training Phase
[0108] The image data is extracted by the feature extraction module to obtain multi-scale abstract features. The feature fusion processing submodule fuses features of different scales. The attention mechanism processing submodule calculates the dependency between features at different positions and channels and uses the weighted values of the encoded high-level abstract features to obtain the final fused features. Finally, score regression is performed, and the loss is calculated by comparing with the calibration scores of different distortion types provided by the training samples. Backpropagation is used to adjust the model parameters of the above parts.
[0109] 4.2 Testing Phase
[0110] It should be noted that the testing phase may include batch testing on the test set and evaluating the test results; it may also include performing actual image scoring tasks.
[0111] For images that need to be evaluated for quality, the same input method as the training images is used to extract multi-scale features and fuse them, calculate attention weights and weight features, and finally obtain quality scores for different distortion types.
[0112] It can be seen that in the embodiment of the present application, various types of mixed distorted images can be processed, their distortion types can be detected, and each distortion type can be quantitatively evaluated separately.
[0113] Secondly, the training and testing process for image quality assessment is an end-to-end network. After inputting an image, the scores for each distortion type are directly output, eliminating the need to perform feature extraction and fusion, distortion type classification, and quality assessment separately.
[0114] Furthermore, the technology extracts deep features of different scales and performs feature fusion on the extracted image features, while ensuring the semantics and detail information of the features;
[0115] Finally, the self-attention mechanism is used to model the dependencies between different feature positions and channels, so as to better maintain the consistency between the same distortion types and the differences between different distortion types.
[0116] The above describes the method provided by this application. The following describes the device provided by this application:
[0117] See Figure 4 , is a structural diagram of an image quality assessment device provided in an embodiment of the present application, such as Figure 4 As shown, the image quality assessment device may include:
[0118] An acquisition unit 410 is configured to acquire an image to be evaluated;
[0119] An evaluation unit 420 is configured to input the image to be evaluated into an image quality assessment model to obtain evaluation scores for N different distortion types of the image to be evaluated; wherein the image quality assessment model includes a feature extraction layer and a score regression layer, the feature extraction layer is configured to extract image features of the input image; the score regression layer is configured to obtain evaluation scores for N different distortion types of the input image based on the image features; N ≥ 2, and N is an integer.
[0120] In some embodiments, the image quality assessment model also includes a feature enhancement layer, which is used to perform feature enhancement processing on the image features extracted by the feature extraction layer to obtain enhanced features, and the score regression layer obtains evaluation scores of N different distortion types of the input image based on the enhanced feature image features.
[0121] In some embodiments, the feature enhancement processing includes feature fusion processing;
[0122] The feature enhancement layer implements feature enhancement processing in the following ways:
[0123] Processing the output features of each layer of the feature extraction layer to have the same number of channels, and obtaining the output features of each layer after the first processing;
[0124] Performing upsampling processing on the output features of each layer after the first processing to obtain the output features of each layer after the second processing; wherein the feature maps of the output features of each layer after the second processing have the same size;
[0125] The output features of each layer after the second processing are subjected to feature fusion processing to obtain fused features.
[0126] In some embodiments, the feature enhancement processing includes attention mechanism processing; the attention mechanism processing includes spatial attention mechanism processing and / or channel attention mechanism processing;
[0127] The feature enhancement layer implements feature enhancement processing in the following ways:
[0128] Processing the image features according to an attention mechanism to determine feature dependencies of the image features in a spatial dimension and / or a channel dimension;
[0129] According to the feature dependency relationship of the image features in the spatial dimension and / or the channel dimension, the image features are enhanced to obtain enhanced features.
[0130] In some embodiments, as Figure 5 As shown, the device also includes:
[0131] The training unit 430 is used to determine the model loss based on the evaluation scores of N different distortion types output after the training samples are input into the image quality assessment model and the calibration scores of the N different distortion types of the training samples, and optimize and adjust the model parameters of each part of the image quality assessment model through back propagation based on the model loss.
[0132] An embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the image quality assessment method described above.
[0133] See Figure 6 , is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing machine-executable instructions. The processor 601 and the memory 602 may communicate via a system bus 603. Furthermore, by reading and executing the machine-executable instructions corresponding to the image quality assessment logic in the memory 602, the processor 601 may perform the image quality assessment method described above.
[0134] The memory 602 mentioned herein can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as a hard disk drive), solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.
[0135] In some embodiments, a machine-readable storage medium is also provided. Figure 6 The memory 602 in the machine-readable storage medium stores machine-executable instructions. When executed by the processor, the machine-executable instructions implement the image quality assessment method described above. For example, the storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0136] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0137] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for image quality assessment, characterized in that: include: Inputting the image to be evaluated into the image quality assessment model to obtain evaluation scores of N different distortion types of the image to be evaluated; The image quality assessment model includes a feature extraction layer and a score regression layer. The feature extraction layer is used to extract image features of the input image; the score regression layer is used to obtain assessment scores of N different distortion types of the input image based on the image features; the score regression layer may include N fully connected layers, each fully connected layer corresponding to a distortion type; N ≥ 2, and N is an integer; The image quality assessment model further includes a feature enhancement layer for performing feature enhancement processing on the image features extracted by the feature extraction layer to obtain enhanced features, and the score regression layer obtains assessment scores of N different distortion types of the input image based on the enhanced feature image features; The feature enhancement processing includes feature fusion processing; The feature enhancement layer implements feature enhancement processing in the following ways: Processing the output features of each layer of the feature extraction layer to have the same number of channels, and obtaining the output features of each layer after the first processing; Performing upsampling processing on the output features of each layer after the first processing to obtain the output features of each layer after the second processing; wherein the feature maps of the output features of each layer after the second processing have the same size; Performing feature fusion processing on the output features of each layer after the second processing to obtain fused features; The feature enhancement processing also includes attention mechanism processing; the attention mechanism processing includes spatial attention mechanism processing and / or channel attention mechanism processing; The feature enhancement layer also implements feature enhancement processing in the following ways: Processing the fused features according to an attention mechanism to determine feature dependencies of the fused features in a spatial dimension and / or a channel dimension; According to the feature dependency of the fused features in the spatial dimension and / or channel dimension, the fused features are enhanced to obtain enhanced features.
2. The method according to claim 1, characterized in that Before inputting the image to be evaluated into the image quality assessment network, the method further includes: The model loss is determined based on the evaluation scores of N different distortion types output after the training samples are input into the image quality assessment model, and the calibration scores of the N different distortion types of the training samples. Based on the model loss, the model parameters of each part of the image quality assessment model are optimized and adjusted through back propagation.
3. An image quality assessment device, characterized in that: include: an acquisition unit, configured to acquire an image to be evaluated; An evaluation unit is configured to input the image to be evaluated into an image quality evaluation model to obtain evaluation scores of N different distortion types of the image to be evaluated; wherein the image quality evaluation model includes a feature extraction layer and a score regression layer, the feature extraction layer is configured to extract image features of the input image; the score regression layer is configured to obtain evaluation scores of the N different distortion types of the input image based on the image features; the score regression layer may include N fully connected layers, each fully connected layer corresponding to one distortion type; N ≥ 2, and N is an integer; The image quality assessment model further includes a feature enhancement layer for performing feature enhancement processing on the image features extracted by the feature extraction layer to obtain enhanced features, and the score regression layer obtains assessment scores of N different distortion types of the input image based on the enhanced feature image features; The feature enhancement processing includes feature fusion processing; The feature enhancement layer implements feature enhancement processing in the following ways: Processing the output features of each layer of the feature extraction layer to have the same number of channels, and obtaining the output features of each layer after the first processing; Performing upsampling processing on the output features of each layer after the first processing to obtain the output features of each layer after the second processing; wherein the feature maps of the output features of each layer after the second processing have the same size; Performing feature fusion processing on the output features of each layer after the second processing to obtain fused features; The feature enhancement processing also includes attention mechanism processing; the attention mechanism processing includes spatial attention mechanism processing and / or channel attention mechanism processing; The feature enhancement layer also implements feature enhancement processing in the following ways: Processing the image features according to an attention mechanism to determine feature dependencies of the image features in a spatial dimension and / or a channel dimension; According to the feature dependency relationship of the image features in the spatial dimension and / or the channel dimension, the image features are enhanced to obtain enhanced features.
4. The device according to claim 3, characterized in that The device further comprises: A training unit is configured to determine a model loss based on the evaluation scores of N different distortion types output after the training samples are input into the image quality assessment model and the calibration scores of the N different distortion types of the training samples, and optimize and adjust the model parameters of each part of the image quality assessment model through back propagation based on the model loss.
5. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the method according to claim 1 or 2.
6. A machine-readable storage medium, characterized in that The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method according to claim 1 or 2 is implemented.
Citation Information
Patent Citations
Image blurring and noise evaluation method based on multi-task convolutional neural network
CN107133948A
Attention module-based image quality evaluation method
CN112634238A