Image quality determination method, apparatus, device, and storage medium

By extracting semantic and saliency information from distorted images and calculating pixel weights, combined with objective methods, the problem of discrepancies between image quality assessment results and human visual perception is solved, achieving more accurate image quality assessment.

CN114596287BActive Publication Date: 2025-11-28BIGO TECH PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210238500.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-11-28
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

In existing technologies, image quality evaluation results cannot accurately reflect human subjective perception. Subjective assessment is complex and easily influenced by human factors, while objective assessment does not match human perception.

Method used

By acquiring distorted and original images, semantic and saliency information is extracted, the weight of each pixel is calculated, and image quality information is calculated by combining objective image quality evaluation methods.

Benefits of technology

This improves the accuracy of image quality assessment, making the results more consistent with human subjective perception while meeting objectivity requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596287B_ABST
    Figure CN114596287B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a kind of image quality determination method, device, equipment and storage medium, the method comprises: obtaining distorted image and corresponding original image, information extraction is carried out to the original image to obtain semantic information and saliency information, according to the semantic information and the saliency information determine the pixel weight of each pixel point, based on the pixel weight and the pixel value of the corresponding pixel point of the original image and the distorted image is calculated to obtain image quality information.The image quality evaluation result finally determined in this scheme meets the prerequisite of the requirement of objectivity, more in line with human eye subjective experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of image processing, and in particular, to an image quality determination method, device, equipment and storage medium. BACKGROUND

[0002] Image quality evaluation refers to quantitatively describing the distortion degree between two images with similar contents by subjective and objective methods. It plays a very important role in algorithm analysis and comparison, system performance evaluation, etc. in the field of image / video processing.

[0003] Image quality evaluation can be divided into subjective evaluation and objective evaluation in terms of methods. Subjective evaluation refers to evaluating image quality through human subjective feeling, that is, giving an original reference image and a distorted image, and letting an observer evaluate the distorted image, and generally using mean opinion score or mean opinion score difference to describe. Subjective evaluation needs a large amount of manpower and material resources, and the evaluation result is easily affected by the subjective factors of testers and external conditions. The complexity of the evaluation process seriously affects its accuracy and universality, so it is extremely difficult to apply it to an actual video processing system. In contrast, objective evaluation directly gives a quantitative value of distortion using a mathematical model, and is simple to operate and is widely used in various fields. For example, a series of objective evaluation indexes commonly used in the industry are used to objectively evaluate image quality. Although objective evaluation indexes are simple to operate and easy to implement, different objective indexes do not conform to the subjective feeling of the human eye to the same extent, which leads to unsatisfactory image quality evaluation results. SUMMARY

[0004] Embodiments of the present application provide an image quality determination method, device, equipment and storage medium, which solve the problem that the image quality evaluation result in the related art cannot well reflect the subjective image quality feeling, so that the finally given image quality evaluation result is more consistent with the subjective feeling of the human eye.

[0005] In a first aspect, embodiments of the present application provide an image quality determination method, which comprises:

[0006] obtaining a distorted image and a corresponding original image, and performing information extraction on the original image to obtain semantic information and saliency information;

[0007] determining a pixel weight of each pixel point according to the semantic information and the saliency information;

[0008] calculating image quality information based on the pixel weight and pixel values of corresponding pixel points of the original image and the distorted image.

[0009] In a second aspect, embodiments of the present application further provide an image quality determination device, which comprises:

[0010] an image acquisition module configured to acquire a distorted image and a corresponding original image;

[0011] an information extraction module configured to perform information extraction on the original image to obtain semantic information and saliency information;

[0012] a weight calculation module configured to determine a pixel weight of each pixel point according to the semantic information and the saliency information;

[0013] an image quality calculation module configured to calculate image quality information based on the pixel weight and pixel values of corresponding pixel points of the original image and the distorted image.

[0014] In a third aspect, an image quality determination device is provided, and the device includes:

[0015] one or more processors;

[0016] a storage device configured to store one or more programs,

[0017] when the one or more programs are executed by the one or more processors, the one or more processors implement the image quality determination method provided in the embodiments of the present application.

[0018] In a fourth aspect, a storage medium storing computer-executable instructions is provided, and the computer-executable instructions, when executed by a computer processor, are configured to perform the image quality determination method provided in the embodiments of the present application.

[0019] In a fifth aspect, a computer program product is provided, and the computer program product includes a computer program stored in a computer-readable storage medium, and at least one processor of a device reads and executes the computer program from the computer-readable storage medium, so that the device performs the image quality determination method provided in the embodiments of the present application.

[0020] In the embodiments of the present application, the distorted image and the corresponding original image are acquired, the semantic information and the saliency information are obtained by performing information extraction on the original image, the pixel weight of each pixel point is determined according to the semantic information and the saliency information, and the image quality information is calculated based on the pixel weight and the pixel values of corresponding pixel points of the original image and the distorted image, thereby improving the accuracy of image quality evaluation, and making the finally determined image quality evaluation result more consistent with the subjective feeling of the human eye under the premise of meeting the objective requirement. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 a flowchart of an image quality determination method provided in the embodiments of the present application;

[0022] Figure 2 A flow chart of a method for determining pixel weights of pixel points according to an embodiment of the present application is provided;

[0023] Figure 3 A flow chart of another method for determining image quality according to an embodiment of the present application is provided;

[0024] Figure 4 A diagrammatic view of information extraction to obtain semantic information according to an embodiment of the present application is provided;

[0025] Figure 5 A diagrammatic view of information extraction to obtain saliency information according to an embodiment of the present application is provided;

[0026] Figure 6 A flow chart of another method for determining image quality according to an embodiment of the present application is provided;

[0027] Figure 7 A flow chart of another method for determining image quality according to an embodiment of the present application is provided;

[0028] Figure 8 A structure block diagram of an image quality determination apparatus according to an embodiment of the present application is provided;

[0029] Figure 9 A structure diagram of an image quality determination device according to an embodiment of the present application is provided. DETAILED DESCRIPTION

[0030] The embodiments of the present application will be further described below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, but not to limit the embodiments of the present application. In addition, it should be noted that, for the convenience of description, only the parts related to the embodiments of the present application are shown in the drawings, but not all the structures.

[0031] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class, not limited to the number of objects, for example, the first object can be one or more. In addition, the specification and claims "and / or" means at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in a "or" relationship.

[0032] Figure 1A flowchart of an image quality determination method provided by an embodiment of the present application can be used to evaluate image and video quality. The method can be executed by a computing device such as a server, a smart terminal, a notebook computer, a tablet computer, or the like, and specifically includes the following steps:

[0033] In step S101, a distorted image and a corresponding original image are obtained, and semantic information and saliency information are extracted from the original image.

[0034] In one embodiment, a comparison calculation is performed between the distorted image and the original image to obtain a quantitative value of the distorted image and further obtain a corresponding image quality evaluation. The original image can be a clear and undistorted image, i.e., a reference image of the distorted image. The distorted image is an image that has different distortion conditions relative to the original image, i.e., a noise image. Based on an objective image quality evaluation method, an attention mechanism is introduced in the embodiment of the present application to achieve a more reasonable evaluation of the final image quality.

[0035] In one embodiment, the obtained original image can be an image or a video frame sequence composed of multiple image frames, such as continuous multiple frame live picture images in a live video. Semantic information and saliency information are extracted from the original image. Optionally, the information extraction can be performed by a trained neural network model on the original image, or an image feature extraction algorithm can be used for information extraction. The semantic information represents the category of each pixel point in the image. Specifically, the category of each pixel point can be one of a plurality of object categories that are pre-set, and the object categories can be exemplarily a face, a microphone, a chair, hair, a keyboard, and the like. The saliency information represents the gray scale of each pixel point in the image, which indicates the degree of attention of each pixel to human eye vision, and can be a normalized gray scale image. Optionally, the semantic information and the saliency information are presented in the form of an image mask.

[0036] In step S102, a pixel weight of each pixel point is determined according to the semantic information and the saliency information.

[0037] In one embodiment, after the semantic information and the saliency information of each pixel point are obtained through information extraction, the pixel weight of each pixel point is calculated based on the semantic information and the saliency information. Optionally, as shown in Figure 2 Figure 2 A flowchart of a method for determining a pixel weight of a pixel point provided by an embodiment of the present application, specifically includes:

[0038] In step S1021, a semantic weight of each pixel point is determined according to the semantic information and a corresponding basic weight.

[0039] ​In one embodiment, the semantic information includes the semantic index value of each pixel obtained from information extraction, where different semantic index values ​​correspond to different object categories. For example, taking a live streaming scenario, there are 13 categories, corresponding to index values ​​0 to 12 respectively; that is, semantic index value 0 corresponds to category 1, semantic index value 1 corresponds to category 2, semantic index value 2 corresponds to category 3, and so on. Different object categories are assigned corresponding basic weights, such as category 1 corresponding to basic weight 1, category 2 corresponding to basic weight 2, and so on. Accordingly, the semantic weight of each pixel can be determined based on the semantic information and the corresponding basic weights by: determining the semantic index value of each pixel, and determining the basic weight corresponding to the semantic index value of each pixel as the semantic weight. For example, the basic weights are denoted as ω0, ω1, ..., ω N These correspond to the basic weight values ​​from semantic index 0 to semantic index N, respectively. The semantic weight of the pixel at position (i,j) is denoted as W. se (i,j), then the semantic weight W se (i,j)=ωpse(i,j), where p se (i,j) represents the semantic index value of the image mask at position (i,j). The specific value of this basic weight can be set subjectively or measured through statistical methods.

[0040] Step S1022: Calculate the saliency weight of each pixel based on the saliency information and the corresponding predefined weights.

[0041] In one embodiment, saliency information includes the grayscale value of each pixel obtained from information extraction. Optionally, the grayscale value can be a normalized numerical value, ranging from 0 to 1. Optionally, a value of 0 represents "most easily ignored," and a value of 1 represents "most interesting." The predefined weights include a predefined minimum weight value and a predefined maximum weight value. An exemplary predefined minimum weight value is denoted as ω. min The predefined maximum weight value is denoted as ω. max Optionally, the saliency weight of each pixel is calculated based on the saliency information and the corresponding predefined weights, including: calculating the linear difference between the grayscale values ​​of each pixel based on the predefined minimum weight value and the predefined maximum weight value to obtain the saliency weight of each pixel. For example, the saliency weight is denoted as W. sa (i,j), then W sa (i,j)=(ω max -ω min )p sa (i,j)+ω min , where p sa(i,j) represents the gray value of the image mask at the (i,j) position. The specific size of the predefined maximum weight value and the predefined minimum weight value can be set by subjective experience or measured by statistical methods.

[0042] Step S1023, the pixel weight of each pixel point is calculated according to the semantic weight and the corresponding saliency weight of each pixel point.

[0043] In one embodiment, after determining the semantic weight and the corresponding saliency weight of each pixel point, the final pixel weight can be calculated by multiplying the semantic weight and the corresponding saliency weight of each pixel point. For example, the semantic weight is denoted as W se (i,j), the saliency weight is denoted as W sa (i,j), and the pixel weight of the pixel point at the (i,j) position, i.e., the total weight, is denoted as W(i,j), then W(i,j) = W se (i,j) * W sa (i,j).

[0044] Step S103, the image quality information is calculated based on the pixel weight, and the pixel values of the corresponding pixel points of the original image and the distorted image.

[0045] In one embodiment, after determining the pixel weight based on the semantic information and the saliency information, the pixel weight is combined with an objective image quality evaluation method to calculate the final image quality information. Optionally, the objective image quality evaluation method can be a PSNR (Peak Signal-to-Noise Ratio) algorithm, an SSIM (Structural SIMilarity) algorithm, or a VMAF (Video Multime Assessment Fusion) algorithm.

[0046] In one embodiment, the image quality information is calculated based on the pixel weight and the pixel values of the corresponding pixel points of the original image and the distorted image, i.e., in the quantitative calculation of the distortion degree of the distorted image relative to the original image, the pixel weight is combined to calculate the image quality information, so that the final calculation result meets the objective requirement and is more consistent with the subjective feeling of the human eye.

[0047] It can be known from the above scheme that the semantic information and the saliency information are obtained by acquiring the distorted image and the corresponding original image and performing information extraction on the original image, the pixel weight of each pixel point is determined according to the semantic information and the saliency information, and the image quality information is calculated based on the pixel weight and the pixel values of the corresponding pixel points of the original image and the distorted image, thereby improving the accuracy of image quality evaluation. When image quality evaluation is performed, the attention mechanism is introduced, so that the finally determined image quality evaluation result is more consistent with the subjective feeling of the human eye under the premise of meeting the objective requirement.

[0048] Figure 3 The flowchart of another image quality determination method provided by the embodiment of the present application gives a specific method for obtaining semantic information and saliency information by performing information extraction on an original image, as shown in Figure 3 , which includes:

[0049] In step S201, a distorted image and a corresponding original image are acquired, and semantic information and saliency information are obtained by performing information extraction on the original image through a trained convolutional neural network with a double-branch structure.

[0050] In one embodiment, the semantic information and the saliency information are obtained by performing information extraction on an original image through a neural network model. Specifically, a convolutional neural network with a double-branch structure is used, wherein semantic information is extracted through one branch, and saliency information is extracted through the other branch. As shown in Figure 4 , a schematic diagram of information extraction to obtain semantic information is given, Figure 4 a schematic diagram of information extraction to obtain saliency information is given as shown in Figure 5 , and Figure 5 a schematic diagram of information extraction to obtain saliency information is given.

[0051] In step S202, the pixel weight of each pixel point is determined according to the semantic information and the saliency information.

[0052] In step S203, image quality information is calculated based on the pixel weight and the pixel values of the corresponding pixel points of the original image and the distorted image.

[0053] As known from the above, when image quality is determined, the semantic information and the saliency information are obtained by performing information extraction on an original image through a trained convolutional neural network with a double-branch structure, so that the accuracy and rationality of the obtained semantic information and saliency information are higher, and the finally obtained image quality information can obviously be more consistent with the subjective feeling of the human eye under the premise of meeting the objective requirement.

[0054] Figure 6 The flowchart of another image quality determination method provided by the embodiments of the present application gives a specific double-branch network architecture to implement the process of extracting semantic information and saliency information, as shown in Figure 6 The flowchart of another image quality determination method provided by the embodiments of the present application gives a specific double-branch network architecture to implement the process of extracting semantic information and saliency information, as shown in

[0055] Step S301, obtaining a distorted image and a corresponding original image.

[0056] Step S302, extracting semantic information from the original image by a semantic branch of a convolutional neural network to obtain the semantic information, wherein the semantic branch uses global pooling as a feature weight factor, and the number of channels of each network layer is less than a preset number of channels.

[0057] In one embodiment, the semantic branch has a small number of feature channels and a deep number of layers, which is conducive to extracting semantic context information of the image. Since the semantic branch only needs a large receptive field to capture the semantic context features of the image, the semantic branch adopts a lightweight structure design, with a small number of channels at each layer, and the feature map is quickly down-sampled as the layer level deepens. In addition, in order to further expand the receptive field, the global pooling is used as the feature weight factor. The preset number of channels can be 1 or 2.

[0058] Step S303, extracting saliency information from the original image by a detail branch of a convolutional neural network to obtain the saliency information, wherein the number of network levels of the detail branch is less than a preset number of levels.

[0059] The detail branch has a large number of feature channels and a shallow number of layers to capture spatial detail information of the image, such as texture, edge, etc. The detail branch is mainly used to extract spatial detail information of the image, and adopts a shallow number of levels, a rich number of features, and a large resolution. Specifically, it includes three stages, each stage being composed of a cascade of a plurality of convolutional layers, batch normalization, and an activation function. The first layer of convolution of each stage has a step of 2, and the other layers have the same number of convolution kernels. The output feature map size of the detail branch is 1 / 8 of the source size. For example, the preset number of levels can be 4, 5, or 6.

[0060] Step S304, fusing the semantic information and the saliency information by a fusion layer of a convolutional neural network to obtain a semantic information map containing updated semantic information and saliency information.

[0061] In an embodiment, after obtaining the semantic information and the saliency information output respectively by the semantic branch and the detail branch, further fusion processing is performed by a fusion layer. The features output by the detail branch and the semantic branch are complementary features, and more complex and accurate feature expression is obtained through the fusion layer, so that the updated semantic information and saliency information are more accurate. For example, in the saliency region, i.e. the region of interest of the human eye of the user, the weights of different object categories represented by the semantic information are: the basic weight of the human face is 1, the basic weight of the hair is 0.9, the basic weight of the microphone is 0.7, and the basic weight of the chair is 0.6; in the non-saliency region, i.e. the region of no interest of the human eye of the user, the basic weight of the human face is 0.5, the basic weight of the hair is 0.4, the basic weight of the microphone is 0.3, and the basic weight of the chair is 0.3.

[0062] Step S305, determining the pixel weight of each pixel point according to the updated semantic information and saliency information.

[0063] Specifically, the manner of determining the pixel weight of each pixel point based on the updated semantic information and saliency information is the same as the manner of determining the pixel weight of each pixel point based on the semantic information and saliency information, which will not be described here.

[0064] Step S306, calculating the image quality information based on the pixel weight, and the pixel values of the corresponding pixel points of the original image and the distorted image.

[0065] As can be seen from the above, when determining the image quality, the semantic branch of the convolutional neural network is used to extract the semantic information of the original image, and the detail branch of the convolutional neural network is used to extract the saliency information of the original image, so that more accurate and targeted information extraction is achieved. Meanwhile, more complex and accurate feature expression is obtained through the fusion layer, so that the finally obtained image quality information is more consistent with the subjective feeling of the human eye under the premise of meeting the objective requirements.

[0066] On the basis of the above technical solution, the convolutional neural network used further includes an auxiliary segmentation head node to improve the convergence speed during training. Specifically, the auxiliary segmentation head is inserted into different positions of each branch, so as to control the channels of different dimensions by adjusting the calculation complexity of the auxiliary segmentation head and the corresponding main segmentation head, thereby realizing fast convergence during training.

[0067] Figure 7 Another flowchart of an image quality determination method provided by the embodiment of the present application is given, which shows a specific process of calculating the image quality information, as shown in Figure 7 The flowchart includes:

[0068] Step S401: Obtain the distorted image and the corresponding original image, and extract semantic information and saliency information from the original image.

[0069] Step S402: Determine the pixel weight of each pixel based on the semantic information and the saliency information.

[0070] Step S403: Calculate the mean square error of the distorted image and the original image based on the pixel weights, and substitute the mean square error into the peak signal-to-noise ratio formula to calculate the image quality information.

[0071] In one embodiment, the calculation is illustrated by combining the obtained pixel weights with the PSNR algorithm. For example, taking the original image and the corresponding distorted image of size (m, n), where I is the original image, K is the distorted image, and i and j represent the coordinates of a pixel in the image, the weighted mean square error (WMSE) is first calculated, and then the weighted image quality information (WPSNR) is normalized. The calculation method is as follows:

[0072]

[0073]

[0074] As shown above, compared to the traditional PSNR which uses equal weights for the errors of each pixel (i.e., W(i,j) is set to 1 in the above formula), WPSNR, by mining the semantic and saliency features of the image, adopts an adaptive weighting strategy and introduces an attention mechanism, thus better aligning with human subjective perception. Experimental data comparison reveals that using WPSNR calculated using this scheme as an image evaluation metric is superior to other objective image quality assessment algorithms such as PSNR, SSIM, and VMAF, as detailed in the table below:

[0075] WPSNR PSNR SSIM VMAF Pearson coefficient 55% 10% 15% 20% Spearman coefficient 70% 5% 15% 10%

[0076] The percentages in the table above represent the proportion of cases where the current evaluation method has the highest correlation coefficient (Pearson coefficient or Spearman coefficient) with MOS data in a batch of typical live streaming video scenarios. It can be seen that WPSNR is significantly better than other evaluation methods.

[0077] Figure 8 This is a structural block diagram of an image quality determination device provided in an embodiment of this application. The device is used to execute the image quality determination method provided in the above embodiments, and has corresponding functional modules and beneficial effects for executing the method. For example... Figure 8 As shown, the device specifically includes: an image acquisition module 101, an information extraction module 102, a weight calculation module 103, and an image quality calculation module 104, wherein,

[0078] The image acquisition module 101 is configured to acquire a distorted image and a corresponding original image;

[0079] The information extraction module 102 is configured to extract information from the original image to obtain semantic information and saliency information;

[0080] The weight calculation module 103 is configured to determine a pixel weight of each pixel point according to the semantic information and the saliency information;

[0081] The image quality calculation module 104 is configured to calculate image quality information based on the pixel weight and pixel values of corresponding pixel points of the original image and the distorted image.

[0082] According to the above scheme, by acquiring a distorted image and a corresponding original image, extracting information from the original image to obtain semantic information and saliency information, determining a pixel weight of each pixel point according to the semantic information and the saliency information, and calculating image quality information based on the pixel weight and pixel values of corresponding pixel points of the original image and the distorted image, the accuracy of image quality evaluation is improved, and the finally determined image quality evaluation result is more consistent with human subjective experience under the premise of meeting the objective requirement.

[0083] In one possible embodiment, the weight calculation module 103 is configured to:

[0084] determine a semantic weight of each pixel point according to the semantic information and a corresponding basic weight;

[0085] calculate a saliency weight of each pixel point according to the saliency information and a corresponding predefined weight;

[0086] calculate the pixel weight of each pixel point according to the semantic weight and the corresponding saliency weight.

[0087] In one possible embodiment, the semantic information includes a semantic index value of each pixel point extracted by information extraction, different semantic index values correspond to different object categories, and the weight calculation module 103 is configured to:

[0088] determine the semantic index value of each pixel point;

[0089] determine the basic weight corresponding to the semantic index value of each pixel point as the semantic weight.

[0090] In one possible embodiment, the saliency information includes a gray value of each pixel point extracted by information extraction, the predefined weight includes a predefined minimum weight value and a predefined maximum weight value, and the weight calculation module 103 is configured to:

[0091] The saliency weight of each pixel point is calculated based on a linear difference of the gray value of each pixel point and the predefined minimum weight value and the predefined maximum weight value.

[0092] In a possible embodiment, the weight calculation module 103 is configured to:

[0093] The product of the semantic weight of each pixel point and the corresponding saliency weight is determined as the pixel weight of the pixel point.

[0094] In a possible embodiment, the information extraction module 102 is configured to:

[0095] The semantic information and the saliency information are extracted from the original image by the convolutional neural network with the trained double-branch structure.

[0096] In a possible embodiment, the information extraction module 102 is configured to:

[0097] The semantic information is extracted from the original image by the semantic branch of the convolutional neural network, the semantic branch uses global pooling as a feature weight factor, and the number of channels of each network layer is less than a preset number of channels.

[0098] The saliency information is extracted from the original image by the detail branch of the convolutional neural network, and the number of network levels of the detail branch is less than a preset number of levels.

[0099] The semantic information and the saliency information are fused by the fusion layer of the convolutional neural network to obtain a semantic information map containing updated semantic information and saliency information.

[0100] In a possible embodiment, the image quality calculation module 104 is configured to:

[0101] The mean square error of the distorted image and the original image is calculated based on the pixel weight.

[0102] The image quality information is calculated by substituting the mean square error into the peak signal-to-noise ratio formula.

[0103] Figure 9 A structural schematic diagram of an image quality determination device provided in an embodiment of the present application is shown in FIG. 1. Figure 9 As shown in FIG. 1, the device includes a processor 201, a memory 202, an input device 203, and an output device 204; the number of processors 201 in the device can be one or more, Figure 9 for example, one processor 201; the processor 201, the memory 202, the input device 203, and the output device 204 in the device can be connected through a bus or other means, Figure 9The bus is taken as an example. The memory 202 is a computer readable storage medium, which can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the image quality determination method in the embodiments of the present application. The processor 201 executes various function applications and data processing of the device by running the software programs, instructions and modules stored in the memory 202, that is, implements the image quality determination method described above. The input device 203 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device 204 can include a display device such as a display screen.

[0104] The embodiments of the present application also provide a storage medium containing computer executable instructions, which are used to execute an image quality determination method described in the above embodiments when executed by a computer processor, and the method comprises the following steps:

[0105] Obtaining a distorted image and a corresponding original image, performing information extraction on the original image to obtain semantic information and saliency information;

[0106] Determining a pixel weight of each pixel point according to the semantic information and the saliency information;

[0107] Calculating image quality information based on the pixel weight and pixel values of corresponding pixel points of the original image and the distorted image.

[0108] It is worth noting that the embodiments of the above image quality determination device include various units and modules only according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy mutual distinction, and do not limit the protection scope of the embodiments of the present application.

[0109] In some possible implementation manners, each aspect of the method provided by the present application can also be implemented in the form of a program product, which includes program codes for causing a computer device to execute the steps in the method according to various exemplary embodiments of the present application described in the specification when the program product runs on the computer device, for example, the computer device can execute the image quality determination method described in the embodiments of the present application. The program product can be realized by any combination of one or more readable media.

Claims

1. An image quality determination method, characterized in that, include: A distorted image and its corresponding original image are acquired. The original image is then processed by a trained convolutional neural network with a dual-branch structure to extract semantic information and saliency information. The dual-branch structure includes a semantic branch and a detail branch. The semantic branch is used to extract semantic information from the original image, and the detail branch is used to extract saliency information from the original image. The pixel weight of each pixel is determined based on the semantic information and the saliency information; Image quality information is calculated based on the pixel weights and the pixel values ​​of corresponding pixels in the original image and the distorted image.

2. The image quality determination method according to claim 1, characterized in that, Determining the pixel weight of each pixel based on the semantic information and the saliency information includes: The semantic weight of each pixel is determined based on the semantic information and the corresponding basic weights. The saliency weight of each pixel is calculated based on the saliency information and the corresponding predefined weights. The pixel weight of each pixel is calculated based on its semantic weight and corresponding saliency weight.

3. The image quality determination method according to claim 2, characterized in that, The semantic information includes the semantic index value of each pixel obtained from information extraction. Different semantic index values ​​correspond to different object categories. Determining the semantic weight of each pixel based on the semantic information and the corresponding basic weight includes: Determine the semantic index value for each pixel; The basic weight corresponding to the semantic index value of each pixel is determined as the semantic weight.

4. The image quality determination method according to claim 2, characterized in that, The saliency information includes the grayscale value of each pixel obtained from information extraction, and the predefined weights include a predefined minimum weight value and a predefined maximum weight value. The calculation of the saliency weight of each pixel based on the saliency information and the corresponding predefined weights includes: The saliency weight of each pixel is obtained by calculating the linear difference between the gray values ​​of each pixel based on the predefined minimum weight value and the predefined maximum weight value.

5. The image quality determination method according to claim 2, characterized in that, The step of calculating the pixel weight of each pixel based on its semantic weight and corresponding saliency weight includes: The pixel weight is determined by multiplying the semantic weight of each pixel by its corresponding saliency weight.

6. The image quality determination method according to claim 5, characterized in that, The semantic and saliency information is extracted from the original image using a trained dual-branch convolutional neural network, including: Semantic information is extracted from the original image by using the semantic branch of a convolutional neural network. The semantic branch uses global pooling as a feature weight factor, and the number of channels in each layer of the network is less than the preset number of channels. Saliency information is obtained by extracting saliency information from the original image through the detail branches of a convolutional neural network, wherein the number of network layers in the detail branches is less than a preset number of layers; The semantic information and the saliency information are fused by the fusion layer of a convolutional neural network to obtain a semantic information graph containing updated semantic information and saliency information.

7. The image quality determination method according to any one of claims 1-5, characterized in that, The process of calculating image quality information based on the pixel weights and the pixel values ​​of corresponding pixels in the original image and the distorted image includes: The mean square error of the distorted image and the original image is calculated based on the pixel weights; The image quality information is obtained by substituting the mean square error into the peak signal-to-noise ratio formula.

8. An image quality determination device, characterized in that, include: The image acquisition module is configured to acquire the distorted image and the corresponding original image. The information extraction module is configured to extract semantic information and saliency information from the original image using a trained convolutional neural network with a dual-branch structure. The dual-branch structure includes a semantic branch and a detail branch. The semantic branch is used to extract semantic information from the original image to obtain semantic information, and the detail branch is used to extract saliency information from the original image to obtain saliency information. The weight calculation module is configured to determine the pixel weight of each pixel based on the semantic information and the saliency information. The image quality calculation module is configured to calculate image quality information based on the pixel weights and the pixel values ​​of corresponding pixels in the original image and the distorted image.

9. An image quality determination device, the device comprising: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the image quality determination method according to any one of claims 1-7.

10. A storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the image quality determination method according to any one of claims 1-7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the image quality determination method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image processing method, device and equipment and storage medium

    CN112703532A

  • Dual-channel stereo image quality evaluation method based on deep learning and human visual characteristics

    CN113888515A