Image quality evaluation method and terminal device

CN117495828BActive Publication Date: 2026-09-29湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311521604.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2026-09-29
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种图像质量评价方法及终端设备,可以解决图像质量评价的精确度低的问题

Benefits of technology

[0081]在本申请的实施例中,通过获取多个失真图像,然后分别获取每个失真图像的高频分量、中频分量和低频分量,再基于所有失真图像的高频分量、中频分量和低频分量对特征编码网络进行更新,并将更新后的特征编码网络作为失真信息感知网络,然后利用失真信息感知网络获取每个失真图像的失真信息特征,再分别针对每个失真图像,获取失真图像的本质信息特征,并将本质信息特征和失真信息特征进行拼接,得到质量感知特征,然后利用所有失真图像的质量感知特征对全连接网络进行更新,并将更新后的全连接网络作为质量评价网络,最后获取待评价失真图像的质量感知特征,利用质量评价网络对待评价失真图像的质量感知特征进行评价,得到待评价失真图像的质量分数。其中,获取失真图像的三种不同频率的分量,考虑到了人类视觉系统对不同频率感知不同的成像特点,根据三种分量得到的失真信息特征能够准确描述失真图像的失真信息,将本质信息特征与准确性高的失真信息特征拼接得到质量感知特征,能够提高质量感知特征的精确度,经过更新得到的质量评价网络具有优质的性能,利用质量评价网络对失真图像的质量感知特征进行质量评价时,能够有效提高图像质量评价的精确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117495828B_ABST
    Figure CN117495828B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of image quality evaluation, and provides an image quality evaluation method and a terminal device. The image quality evaluation method comprises the following steps: obtaining a distorted image; obtaining high-frequency components, medium-frequency components and low-frequency components of the distorted image; updating a feature coding network based on all the high-frequency components, the medium-frequency components and the low-frequency components, and taking the updated feature coding network as a distorted information perception network; obtaining distorted information features of the distorted image by using the distorted information perception network; obtaining essential information features of the distorted image, splicing the essential information features and the distorted information features to obtain quality perception features; updating a full connection network by using all the quality perception features of the distorted image, and taking the updated full connection network as a quality evaluation network; and evaluating a to-be-evaluated distorted image by using the quality evaluation network to obtain a quality score of the to-be-evaluated distorted image. The method can effectively improve the accuracy of image quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image quality assessment technology, and in particular to an image quality assessment method and terminal device. Background Technology

[0002] Image quality assessment aims to automatically predict the perceived quality of images by simulating the human visual system. In the era of digital imaging, image quality assessment plays a crucial role in image compression, image enhancement, and image restoration. Based on the amount of reference information obtained from the original image, image quality assessment methods can be categorized into full-reference image quality assessment methods, half-reference image quality assessment methods, and no-reference image quality assessment methods. Among these three categories, no-reference image quality assessment is the most promising and practically valuable research direction because it can predict the perceived quality of distorted images without knowing any information from the original image.

[0003] Existing no-reference image quality assessment methods are mainly divided into two categories: those based on handcrafted features and those based on depth features. Handcrafted feature-based methods typically require the involvement of domain experts and involve a degree of subjectivity in feature selection and design. Meanwhile, thanks to the powerful feature extraction capabilities of deep learning models such as convolutional neural networks, no-reference image quality assessment methods generally exhibit higher performance. In summary, research on no-reference image quality assessment based on depth features has become a major trend in the field of image quality assessment.

[0004] With the significant success of the "self-supervised pre-training + downstream task fine-tuning" paradigm based on visual Transformer models in various visual tasks, many depth feature-based no-reference image quality assessment methods combining this paradigm have been proposed. However, these assessment methods did not take into account the characteristics of the human visual system when designing the self-supervised pre-training framework, which limited their practical applications and resulted in low accuracy in image quality assessment. Summary of the Invention

[0005] This application provides an image quality evaluation method and terminal device, which can solve the problem of low accuracy in image quality evaluation.

[0006] In a first aspect, embodiments of this application provide an image quality assessment method, which includes:

[0007] Acquire multiple distorted images;

[0008] The high-frequency, mid-frequency, and low-frequency components of each distorted image are obtained separately.

[0009] The feature coding network is updated based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, and the updated feature coding network is used as the distortion information perception network.

[0010] Distortion information features of each distorted image are obtained using a distortion information perception network;

[0011] For each distorted image, the essential information features of the distorted image are obtained, and the essential information features and the distortion information features are concatenated to obtain the quality perception features; the essential information features are used to describe the basic physical information of the distorted image.

[0012] The fully connected network is updated using the quality-aware features of all distorted images, and the updated fully connected network is used as the quality evaluation network.

[0013] The quality-perceived features of the distorted image to be evaluated are obtained, and the quality evaluation network is used to evaluate the quality-perceived features of the distorted image to be evaluated, so as to obtain the quality score of the distorted image to be evaluated.

[0014] Optionally, the high-frequency components, mid-frequency components, and low-frequency components of each distorted image are obtained separately, including:

[0015] Through the formula:

[0016]

[0017] Calculate the high-frequency components of the m-th distorted image.

[0018] Among them, I m (x,y) represents the m-th distorted image, x represents the horizontal position of pixel (x,y) in the m-th distorted image, x∈{0,1,...,R-1}, R represents the width of the m-th distorted image, y represents the vertical position of pixel (x,y) in the m-th distorted image, y∈{0,1,...,H-1}, H represents the height of the m-th distorted image, m=1,2,...,M, M represents the total number of distorted images, G() represents the Gaussian function, σ0 and σ1 are both standard deviations;

[0019] Through the formula:

[0020]

[0021] Calculate the mid-frequency component of the m-th distorted image.

[0022] Where σ² represents the standard deviation;

[0023] Through the formula:

[0024]

[0025] Calculate the low-frequency component of the m-th distorted image.

[0026] Where σ3 represents the standard deviation.

[0027] Optionally, the feature encoding network includes a feature extraction module and a feature mapping module connected in sequence;

[0028] The feature encoding network is updated based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, including:

[0029] Based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, multiple positive sample pairs are determined.

[0030] For each distorted image, the feature extraction module is used to extract features from the high-frequency, mid-frequency, and low-frequency components of the distorted image, respectively, to obtain high-frequency features, mid-frequency features, and low-frequency features.

[0031] The feature mapping module is used to map the high-frequency, mid-frequency, and low-frequency features of each distorted image to the same feature space;

[0032] A distortion information loss function is constructed based on the high-frequency, mid-frequency, and low-frequency features of all distorted images, as well as all positive sample pairs.

[0033] The feature encoding network is updated using a distortion information loss function to make the features corresponding to positive sample pairs as close as possible in the feature space, and the features corresponding to negative sample pairs as far away as possible.

[0034] Optionally, the feature encoding network can be updated using a distortion information loss function, including:

[0035] Determine whether the value of the distortion information loss function reaches the preset value of the distortion information loss function;

[0036] If so, then the feature coding network will be used as the updated feature coding network;

[0037] Otherwise, adjust the values ​​of the parameters in the feature encoding network and return the steps of extracting features from the high-frequency, mid-frequency, and low-frequency components of the distorted image using the feature extraction module for each distorted image to obtain high-frequency features, mid-frequency features, and low-frequency features.

[0038] Optionally, each distorted image has a distortion type;

[0039] Based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, multiple positive sample pairs are identified, including:

[0040] Iterate through the high-frequency, mid-frequency, and low-frequency components of all distorted images, performing the following steps:

[0041] Two high-frequency components with the same distortion type are treated as a positive sample pair, and any two high-frequency components with different distortion types are treated as a negative sample pair.

[0042] Two intermediate frequency components with the same distortion type are treated as a positive sample pair, and any two intermediate frequency components with different distortion types are treated as a negative sample pair.

[0043] Two low-frequency components with the same distortion type are treated as a positive sample pair, and any two low-frequency components with different distortion types are treated as a negative sample pair.

[0044] A high-frequency component and a mid-frequency component are used as a negative sample pair.

[0045] A high-frequency component and a low-frequency component are used as a negative sample pair.

[0046] A mid-frequency component and a low-frequency component are used as a negative sample pair.

[0047] Optionally, the loss function for distorted information is:

[0048]

[0049] in, Lpre This represents the value of the loss function for distorted information, where i = 0, 1, 2. When i = 0, This represents the high-frequency loss function, when i = 1. This represents the intermediate frequency loss function, when i = 2. L represents the low-frequency loss function. fre The discriminant loss function is:

[0050]

[0051]

[0052] Where P(m) represents the set of positive sample pairs including the i-th component of the m-th distorted image, |P(m)| represents the number of elements in P(m), and P(a) i Let |P(a)| represent the set of positive sample pairs including the i-th component of the a-th distorted image. i | represents P(a) i The number of elements in the image, a,m∈{1,2,...,M}, where M represents the total number of distorted images. When i=0, the i-th component represents the high-frequency component; when i=1, the i-th component represents the mid-frequency component; when i=2, the i-th component represents the low-frequency component. g() represents the first loss function, and f() represents the second loss function.

[0053]

[0054]

[0055] in, Let i represent the i-th component feature of the m-th distorted image. This represents the i-th component feature of the l-th distorted image. This represents the i-th component feature of the k-th distorted image. This represents the i-th component feature of the a-th distorted image. Let i represent the feature of the i-th component of the b-th distorted image. The component feature is one of the following: high-frequency feature, mid-frequency feature, and low-frequency feature. Let j represent the feature of the j-th component of the c-th distorted image, where j = 0, 1, 2, 1 represents the indicator function, which takes the value 1 when k ≠ m or j ≠ i, c ≠ a, s() represents the cosine similarity function, and τ represents the temperature parameter.

[0056] Optionally, the fully connected network can be updated using the quality-aware features of all distorted images, including:

[0057] For each distorted image, a fully connected network is used to evaluate the quality-perceived features corresponding to the distorted image, and a quality score for the distorted image is obtained.

[0058] The evaluation loss function value is calculated based on the quality score of the distorted image, and it is determined whether the evaluation loss function value reaches the preset value of the evaluation loss function.

[0059] If so, then the fully connected network will be used as the updated fully connected network;

[0060] Otherwise, adjust the parameters in the fully connected network and return to the steps of evaluating the quality-perceived features of each distorted image using the fully connected network to obtain the quality score of the distorted image.

[0061] Optionally, the evaluation loss function value is calculated based on the quality score, including:

[0062] Through the formula:

[0063]

[0064] Calculate the evaluation loss function value L;

[0065] Among them, h m t represents the quality score of the m-th distorted image. m Let m represent the true quality score of the m-th distorted image, where m = 1, 2, ..., M, and M represents the total number of distorted images.

[0066] Optionally, obtain the quality-perceived features of the distorted image to be evaluated, including:

[0067] Distortion information features of the distorted image to be evaluated are obtained using a distortion information perception network.

[0068] Obtain the essential information features of the distorted image to be evaluated;

[0069] The essential information features and distortion information features of the distorted image to be evaluated are concatenated to obtain the quality perception features of the distorted image to be evaluated.

[0070] Secondly, embodiments of this application provide an image quality evaluation device, comprising:

[0071] The acquisition module acquires multiple distorted images;

[0072] The component acquisition module acquires the high-frequency component, mid-frequency component, and low-frequency component of each distorted image, respectively.

[0073] The first update module updates the feature coding network based on the high-frequency, mid-frequency and low-frequency components of all distorted images, and uses the updated feature coding network as the distortion information perception network.

[0074] The distortion information acquisition module uses a distortion information perception network to acquire the distortion information features of each distorted image;

[0075] The stitching module extracts the essential information features of each distorted image and stitches them together with the distortion information features to obtain the quality-perceived features. The essential information features are used to describe the basic physical information of the distorted image.

[0076] The second update module updates the fully connected network using the quality-aware features of all distorted images and uses the updated fully connected network as the quality evaluation network.

[0077] The evaluation module acquires the quality-perceived features of the distorted image to be evaluated, and uses a quality evaluation network to evaluate the quality-perceived features of the distorted image to obtain a quality score for the distorted image.

[0078] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the image quality evaluation method described above.

[0079] Fourthly, embodiments of this application provide a computer-readable medium storing a computer program that, when executed by a processor, implements the image quality evaluation method described above.

[0080] The above-mentioned solution in this application has the following beneficial effects:

[0081] In the embodiments of this application, multiple distorted images are acquired, and then the high-frequency, mid-frequency, and low-frequency components of each distorted image are acquired respectively. The feature encoding network is then updated based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, and the updated feature encoding network is used as the distortion information perception network. The distortion information perception network is then used to acquire the distortion information features of each distorted image. For each distorted image, the essential information features of the distorted image are acquired, and the essential information features and the distortion information features are concatenated to obtain the quality perception features. The quality perception features of all distorted images are then used to update the fully connected network, and the updated fully connected network is used as the quality evaluation network. Finally, the quality perception features of the distorted image to be evaluated are acquired, and the quality evaluation network is used to evaluate the quality perception features of the distorted image to be evaluated to obtain the quality score of the distorted image to be evaluated. Among them, three different frequency components of the distorted image are obtained, taking into account the different imaging characteristics of the human visual system in perceiving different frequencies. The distortion information features obtained from the three components can accurately describe the distortion information of the distorted image. The essential information features are spliced ​​with the highly accurate distortion information features to obtain the quality perception features, which can improve the accuracy of the quality perception features. The updated quality evaluation network has excellent performance. When using the quality evaluation network to evaluate the quality perception features of the distorted image, it can effectively improve the accuracy of image quality evaluation.

[0082] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0083] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0084] Figure 1 A flowchart of an image quality evaluation method provided in an embodiment of this application;

[0085] Figure 2 A schematic diagram of a distortion sensing network provided in an embodiment of this application;

[0086] Figure 3 A detailed flowchart of an image quality evaluation method provided in an embodiment of this application;

[0087] Figure 4This is a schematic diagram of the structure of an image quality evaluation device provided in an embodiment of this application;

[0088] Figure 5 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0089] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0090] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0091] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0092] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0093] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0094] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0095] To address the issue of low accuracy in existing image quality assessment methods, this application provides an image quality assessment method. This method involves acquiring multiple distorted images, then obtaining the high-frequency, mid-frequency, and low-frequency components of each distorted image. The feature encoding network is then updated based on these components, and the updated feature encoding network is used as a distortion information perception network. This network is then used to acquire distortion information features for each distorted image. For each distorted image, essential information features are acquired, and these features are concatenated with the distortion information features to obtain quality-perceived features. The quality-perceived features of all distorted images are then used to update a fully connected network, which is used as a quality assessment network. Finally, the quality-perceived features of the distorted image to be evaluated are obtained, and the quality assessment network is used to evaluate these features to obtain a quality score for the distorted image. Among them, three different frequency components of the distorted image are obtained, taking into account the different imaging characteristics of the human visual system in perceiving different frequencies. The distortion information features obtained from the three components can accurately describe the distortion information of the distorted image. The essential information features are spliced ​​with the highly accurate distortion information features to obtain the quality perception features, which can improve the accuracy of the quality perception features. The updated quality evaluation network has excellent performance. When using the quality evaluation network to evaluate the quality perception features of the distorted image, it can effectively improve the accuracy of image quality evaluation.

[0096] The image quality evaluation method provided in this application will be described exemplarily below.

[0097] like Figure 1 As shown, the image quality evaluation method provided in this application includes the following steps:

[0098] Step 11: Obtain multiple distorted images.

[0099] It should be noted that the distorted images mentioned above are images that have undergone distortion resulting in quality changes, but still have a true quality score, such as overexposure, low resolution, or missing pixels.

[0100] For example, the distorted images used in this application may be synthetic distorted images from the Constance Artificially Distorted Image Quality Set (KADIS), real distorted images from the Blur dataset, the Visual Object Classes dataset (VOC), the Microsoft Common Objects in Context dataset (COCO), and the Atomic Visual Actions dataset (AVA).

[0101] It is worth mentioning that by acquiring multiple distorted images and using them as the research object, it is easier to process them in subsequent steps.

[0102] Step 12: Obtain the high-frequency component, mid-frequency component, and low-frequency component of each distorted image.

[0103] In some embodiments of this application, the steps of obtaining the high-frequency component, mid-frequency component, and low-frequency component of each distorted image specifically include:

[0104] The first step is through the formula:

[0105]

[0106] Calculate the high-frequency components of the m-th distorted image.

[0107] Among them, I m (x,y) represents the m-th distorted image, where x represents the horizontal position of pixel (x,y) in the m-th distorted image, x∈{0,1,...,R-1}, R represents the width of the m-th distorted image, y represents the vertical position of pixel (x,y) in the m-th distorted image, y∈{0,1,...,H-1}, H represents the height of the m-th distorted image, m=1,2,...,M, M represents the total number of distorted images, G() represents the Gaussian function, and σ0 and σ1 are both standard deviations.

[0108] The second step is to use the formula:

[0109]

[0110] Calculate the mid-frequency component of the m-th distorted image.

[0111] Where σ² represents the standard deviation.

[0112] The third step is to use the formula:

[0113]

[0114] Calculate the low-frequency component of the m-th distorted image.

[0115] Where σ3 represents the standard deviation.

[0116] It should be noted that the aforementioned high-frequency components refer to the parts of the distorted image where the frequency of grayscale value changes exceeds a preset upper frequency threshold, such as contours, noise, and details. The aforementioned mid-frequency components refer to the parts of the distorted image where the frequency of grayscale value changes is below the preset upper frequency threshold but above a preset lower frequency threshold, such as edge regions. The aforementioned low-frequency components refer to the parts of the distorted image where the frequency of grayscale value changes is below the preset lower frequency threshold, such as the main color region. When calculating the high-frequency, mid-frequency, and low-frequency components using the above three formulas, the standard deviation serves as the preset upper and lower frequency thresholds.

[0117] For example, computer software such as MATLAB can be used to calculate the high-frequency, mid-frequency, and low-frequency components of each distorted image.

[0118] It is worth mentioning that, compared to other frequency components in an image, the human visual system is more sensitive to high-frequency components. The three different frequency components used to acquire the distorted image take into account the different imaging characteristics of the human visual system in perceiving different frequencies.

[0119] Step 13: Update the feature coding network based on the high-frequency, mid-frequency and low-frequency components of all distorted images, and use the updated feature coding network as the distortion information perception network.

[0120] The aforementioned feature encoding network includes a feature extraction module and a feature mapping module connected in sequence.

[0121] In some embodiments of this application, the step of updating the feature coding network based on the high-frequency, mid-frequency, and low-frequency components of all distorted images specifically includes:

[0122] The first step is to identify multiple positive sample pairs based on the high-frequency, mid-frequency, and low-frequency components of all distorted images.

[0123] Each of the above distorted images has a type of distortion;

[0124] The specific determination process involves iterating through the high-frequency, mid-frequency, and low-frequency components of all distorted images and performing the following steps:

[0125] Two high-frequency components with the same distortion type are treated as a positive sample pair, and two high-frequency components with different distortion types are treated as a negative sample pair.

[0126] Two intermediate frequency components with the same distortion type are treated as a positive sample pair, and two intermediate frequency components with different distortion types are treated as a negative sample pair.

[0127] Two low-frequency components with the same distortion type are treated as a positive sample pair, and two low-frequency components with different distortion types are treated as a negative sample pair.

[0128] A high-frequency component and a mid-frequency component are used as a negative sample pair.

[0129] A high-frequency component and a low-frequency component are used as a negative sample pair.

[0130] A mid-frequency component and a low-frequency component are used as a negative sample pair.

[0131] It should be noted that the above steps are performed on all distorted images to obtain positive and negative sample pairs belonging to the high-frequency, mid-frequency, and low-frequency components of each distorted image, ultimately resulting in a large number of positive and negative sample pairs.

[0132] For example, the distortion type of a distorted image can be exposure, missing elements, blur, etc. If two high-frequency components are selected, and the distortion type of the distorted image corresponding to one high-frequency component is blur, and the distortion type of the distorted image corresponding to the other high-frequency component is also blur, then these two high-frequency components are considered a positive sample pair. If the distortion type of the distorted image corresponding to the other high-frequency component is exposure, then these two high-frequency components are considered a negative sample pair. If one of the selected components is a high-frequency component and the other is a low-frequency component, then these two components are considered a negative sample pair. That is, when two components belong to the same type and the corresponding distorted images have the same distortion type, these two components are considered a positive sample pair.

[0133] The second step involves using the feature extraction module to extract features from the high-frequency, mid-frequency, and low-frequency components of each distorted image, resulting in high-frequency features, mid-frequency features, and low-frequency features.

[0134] That is, the feature extraction module can be used to extract features from the high-frequency components of the distorted image to obtain high-frequency features; the feature extraction module can be used to extract features from the mid-frequency components of the distorted image to obtain mid-frequency features; and the feature extraction module can be used to extract features from the low-frequency components of the distorted image to obtain low-frequency features.

[0135] The third step is to use the feature mapping module to map the high-frequency, mid-frequency, and low-frequency features of each distorted image to the same feature space.

[0136] The fourth step is to construct a distortion information loss function based on the high-frequency features, mid-frequency features, low-frequency features of all distorted images, and all positive sample pairs.

[0137] The above-mentioned loss function for distorted information is:

[0138]

[0139] Among them, L pre This represents the value of the loss function for distorted information, where i = 0, 1, 2. When i = 0, This represents the high-frequency loss function, when i = 1. This represents the intermediate frequency loss function, when i = 2. L represents the low-frequency loss function. fre The discriminant loss function is:

[0140]

[0141]

[0142] Where P(m) represents the set of positive sample pairs including the i-th component of the m-th distorted image, |P(m)| represents the number of elements in P(m), and P(a) i Let |P(a)| represent the set of positive sample pairs including the i-th component of the a-th distorted image. i | represents P(a) i The number of elements in the image, a,m∈{1,2,...,M}, where M represents the total number of distorted images. When i=0, the i-th component represents the high-frequency component; when i=1, the i-th component represents the mid-frequency component; when i=2, the i-th component represents the low-frequency component. g() represents the first loss function, and f() represents the second loss function.

[0143]

[0144]

[0145] in, Let i represent the i-th component feature of the m-th distorted image. This represents the i-th component feature of the l-th distorted image. This represents the i-th component feature of the k-th distorted image. This represents the i-th component feature of the a-th distorted image. Let i represent the feature of the i-th component of the b-th distorted image. The component feature is one of the following: high-frequency feature, mid-frequency feature, and low-frequency feature. Let j represent the feature of the j-th component of the c-th distorted image, where j = 0, 1, 2, 1 represents the indicator function, which takes the value 1 when k ≠ m or j ≠ i, c ≠ a, s() represents the cosine similarity function, and τ represents the temperature parameter.

[0146] The fifth step is to update the feature encoding network using the distortion information loss function, so that the features corresponding to positive sample pairs are as close as possible in the feature space, while the features corresponding to negative sample pairs are as far apart as possible.

[0147] The specific update process is as follows: determine whether the value of the distortion information loss function has reached the preset value of the distortion information loss function;

[0148] If so, then the feature coding network will be used as the updated feature coding network;

[0149] Otherwise, adjust the values ​​of the parameters in the feature encoding network and return the steps of extracting features from the high-frequency, mid-frequency, and low-frequency components of the distorted image using the feature extraction module for each distorted image to obtain high-frequency features, mid-frequency features, and low-frequency features.

[0150] It should be noted that the feature extraction module mentioned above can be a convolutional neural network ResNet50, and the feature mapping module mentioned above can be a multilayer perceptron consisting of a fully connected layer, a batch normalization layer, a ReLU activation function layer, a fully connected layer, and a batch normalization layer connected in sequence.

[0151] For example, the distortion information loss function has a value of 2, while the preset value is 2.5. Since the value of the distortion information loss function does not reach the preset value, it indicates that the performance of the feature encoding network is not as expected. Therefore, the parameters in the feature encoding network are adjusted to update it, and the high-frequency, mid-frequency, and low-frequency features of each distorted image are re-obtained using the feature encoding network. At this point, the value of the distortion information loss function is 2.7, reaching the preset value. This indicates that the performance of the feature encoding network is now as expected. Therefore, this updated feature encoding network is used as the new feature encoding network, and then this updated feature encoding network is used as the distortion information perception network. In other words, the distortion information perception network is the trained feature encoding network.

[0152] It is worth mentioning that the distortion information features obtained from the three components can accurately describe the distortion information of the distorted image, update the feature encoding network, improve the accuracy of the high-frequency, mid-frequency and low-frequency features obtained, and at the same time make the features corresponding to positive sample pairs as close as possible in the feature space, while making the features corresponding to negative sample pairs as far apart as possible, thereby improving the feature encoding network's ability to perceive positive sample pairs.

[0153] Step 14: Use the distortion information perception network to obtain the distortion information features of each distorted image.

[0154] It should be noted that the above distortion information features describe the distortion information of the distorted image. The distortion information describes the distortion situation of the distorted image, such as the missing pixel parts, the exposed area, and the resolution after distortion.

[0155] For example, for each distorted image, the feature extraction module in the distortion information perception network extracts the high-frequency, mid-frequency, and low-frequency features of the distorted image. Then, the feature mapping module in the distortion information perception network maps the high-frequency, mid-frequency, and low-frequency features to the same feature space to obtain the distortion information features of the distorted image.

[0156] It is worth mentioning that the distortion information features obtained by using distortion information perception networks can effectively describe the distortion information of distorted images.

[0157] The steps for obtaining distortion information features described above will be illustrated below with a specific example.

[0158] like Figure 2 As shown, the distorted image 1 in the figure is decomposed by Gaussian difference to obtain high frequency component 1, mid frequency component 1 and low frequency component 1 (that is, the steps above to obtain the high frequency component, mid frequency component and low frequency component of each distorted image respectively). The distorted image 2 is decomposed by Gaussian difference to obtain high frequency component 2, mid frequency component 2 and low frequency component 2. Then all components are input into the distortion information perception network to obtain high frequency feature 1, mid frequency feature 1, low frequency feature 1, high frequency feature 2, mid frequency feature 2 and low frequency feature 2. All the obtained features are mapped to the feature space. In the feature space, all features are distributed according to the positive sample pairs and negative sample pairs to which the corresponding components belong.

[0159] Step 15: For each distorted image, obtain the essential information features of the distorted image, and concatenate the essential information features and the distortion information features to obtain the quality perception features.

[0160] The aforementioned essential information features are used to describe the basic physical information of distorted images, such as pixel grayscale, image size, and resolution.

[0161] In some embodiments of this application, pre-trained self-attention mechanism models (such as Transformer) can be used to obtain essential information features of distorted images.

[0162] It is worth mentioning that the quality perception feature obtained by splicing the essential information feature and the distortion information feature can well describe the comprehensive information of the distorted image. Moreover, the distortion information feature is based on three components and the distortion information perception network, which has high accuracy, thereby improving the accuracy of the information in the quality perception feature.

[0163] Step 16: Update the fully connected network using the quality-aware features of all distorted images, and use the updated fully connected network as the quality evaluation network.

[0164] The aforementioned fully connected network can be a multilayer perceptron.

[0165] In some embodiments of this application, the above steps specifically include:

[0166] The first step is to evaluate the quality-perceived features of each distorted image using a fully connected network to obtain a quality score for the distorted image.

[0167] The second step is to calculate the evaluation loss function value based on the quality score of the distorted image and determine whether the evaluation loss function value reaches the preset value of the evaluation loss function.

[0168] If so, then the fully connected network will be used as the updated fully connected network;

[0169] Otherwise, adjust the parameters in the fully connected network and return to the steps of evaluating the quality-perceived features of each distorted image using the fully connected network to obtain the quality score of the distorted image.

[0170] The above calculation of the evaluation loss function value based on the quality score is as follows:

[0171] Through the formula:

[0172]

[0173] Calculate the evaluation loss function value L;

[0174] Among them, h m t represents the quality score of the m-th distorted image. m Let m represent the true quality score of the m-th distorted image, where m = 1, 2, ..., M, and M represents the total number of distorted images.

[0175] It should be noted that the above fully connected network consists of two ReLU activation function layers and three sequentially connected fully connected layers. The two ReLU activation function layers are inserted between the two fully connected layers to perform nonlinear transformation on the outputs of the first two fully connected layers.

[0176] For example, if the evaluation loss function value is 1 and the preset evaluation loss function value is 3, and the evaluation loss function value does not reach the preset evaluation loss function value, it means that the accuracy of the quality score predicted by the fully connected network is low. In this case, the parameters in the fully connected network are adjusted to update the fully connected network. The updated fully connected network is used to re-acquire the quality score of each distorted image and calculate the evaluation loss function value. If the evaluation loss function value reaches the preset evaluation loss function value, it means that the accuracy of the quality score predicted by the fully connected network has reached the expectation. In this case, the fully connected network is used as the quality evaluation network.

[0177] It is worth mentioning that, based on the true quality score of the distorted image, the evaluation loss function is used to update the fully connected network, resulting in a quality evaluation network with excellent performance.

[0178] Step 17: Obtain the quality-perceived features of the distorted image to be evaluated, and use the quality evaluation network to evaluate the quality-perceived features of the distorted image to be evaluated, thereby obtaining the quality score of the distorted image to be evaluated.

[0179] The above-mentioned distorted images are those that have been distorted and need to be evaluated to obtain a quality score.

[0180] In some embodiments of this application, the steps for obtaining the quality-perceived features of the distorted image to be evaluated are specifically as follows:

[0181] The first step is to use a distortion information perception network to obtain the distortion information features of the distorted image to be evaluated.

[0182] The second step is to obtain the essential information features of the distorted image to be evaluated.

[0183] The third step is to concatenate the essential information features and distortion information features of the distorted image to be evaluated to obtain the quality perception features of the distorted image to be evaluated.

[0184] In some embodiments of this application, a pre-trained self-attention mechanism model (such as Transformer) can be used to obtain essential information features of the distorted image to be evaluated.

[0185] It should be noted that when using a distortion information sensing network to obtain the distortion information features of the distorted image to be evaluated, the steps are the same as those described above for obtaining the distortion information features of each distorted image using a distortion information sensing network. It is necessary to first obtain the high-frequency components, mid-frequency components, and low-frequency components of the distorted image to be evaluated.

[0186] It is worth mentioning that the acquisition of three different frequency components of the distorted image takes into account the different imaging characteristics of the human visual system in perceiving different frequencies. The distortion information features obtained from the three components can accurately describe the distortion information of the distorted image. By concatenating the essential information features with the highly accurate distortion information features, the quality perception features can be obtained, which can improve the accuracy of the quality perception features. The updated quality assessment network has excellent performance. When using the quality assessment network to perform quality assessment on the quality perception features of the distorted image, the accuracy of image quality assessment can be effectively improved.

[0187] The image quality evaluation method of this application will be illustrated below with a specific example.

[0188] like Figure 3 As shown, the distorted image to be evaluated is decomposed into low-frequency, mid-frequency, and high-frequency components. The low-frequency, mid-frequency, and high-frequency components are input into the distortion information perception network to obtain distortion information features. At the same time, the distorted image is input into the visual network (i.e., the self-attention mechanism model that has been pre-trained above) to obtain essential information features. The essential information features and the distortion information features are concatenated to obtain quality perception features. The quality perception features are input into the quality evaluation network to obtain the quality score of the distorted image to be evaluated (i.e., A in the figure).

[0189] Eight representative image quality assessment methods were selected: HOSA (High Order Statistics Aggregation), BIECON (Blind Image Evaluator based on a Convolutional Neural Network), WaDIQaM-NR (Weighted Average Deep Image Quarity Measure for NR-IQA), RankIQA (Image Quality Assessment method learns from Rankings), UNIQUE (Unified No-reference Image Quality and Uncertainty Evaluator) based on Deep Neural Networks, CONTRIQUE (Contrastive Image Quality Evaluator), Re-IQA (Rethinking-Image Quality Assessment Method), and LIQE (Language-Image Quality). The image quality assessment methods of this application are compared with those of the Evaluator. CONTRIQUE and Re-IQA are two deep feature-based no-reference image quality assessment methods that combine the "self-supervised pre-training + downstream task fine-tuning" paradigm. Images from the LIVE dataset, the CSIQ (categorical subjective image quality) dataset, and the Tampere Image Database 2013 (TID2013) are selected as inputs. Each of the above image quality assessment methods is used to evaluate the quality, and the Spearman rank-order correlation coefficient (SROCC) and Pearson linear correlation coefficient (PLCC) values ​​of each image quality assessment method are calculated. The results are shown in Table 1.

[0190]

[0191] Table 1

[0192] It can be seen that the image quality evaluation method of this application can effectively improve the accuracy of the quality score of the distorted image to be evaluated.

[0193] The image quality evaluation device provided in this application will be described exemplarily below.

[0194] like Figure 4 As shown, this application embodiment provides an image quality evaluation device 400, which includes:

[0195] Module 401 acquires multiple distorted images;

[0196] The component acquisition module 402 acquires the high-frequency component, mid-frequency component and low-frequency component of each distorted image respectively;

[0197] The first update module 403 updates the feature coding network based on the high-frequency, mid-frequency and low-frequency components of all distorted images, and uses the updated feature coding network as the distortion information perception network.

[0198] The distortion information acquisition module 404 uses a distortion information perception network to acquire the distortion information features of each distorted image;

[0199] The stitching module 405 acquires the essential information features of each distorted image, and stitches the essential information features and the distortion information features together to obtain the quality perception features; the essential information features are used to describe the basic physical information of the distorted image.

[0200] The second update module 406 updates the fully connected network using the quality-aware features of all distorted images and uses the updated fully connected network as the quality evaluation network.

[0201] Evaluation module 407 acquires the quality-perceived features of the distorted image to be evaluated, uses a quality evaluation network to evaluate the quality-perceived features of the distorted image to be evaluated, and obtains the quality score of the distorted image to be evaluated.

[0202] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0203] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0204] like Figure 5 As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 5 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.

[0205] Specifically, when the processor D100 executes the computer program D102, it acquires multiple distorted images, then acquires the high-frequency, mid-frequency, and low-frequency components of each distorted image, updates the feature encoding network based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, and uses the updated feature encoding network as the distortion information perception network. The distortion information perception network is then used to acquire the distortion information features of each distorted image. For each distorted image, the essential information features are acquired, and the essential information features and distortion information features are concatenated to obtain quality-perceived features. The quality-perceived features of all distorted images are then used to update the fully connected network, which is used as the quality evaluation network. Finally, the quality-perceived features of the distorted image to be evaluated are acquired, and the quality evaluation network is used to evaluate the quality-perceived features of the distorted image to be evaluated, thus obtaining the quality score of the distorted image to be evaluated. Among them, three different frequency components of the distorted image are obtained, taking into account the different imaging characteristics of the human visual system in perceiving different frequencies. The distortion information features obtained from the three components can accurately describe the distortion information of the distorted image. The essential information features are spliced ​​with the highly accurate distortion information features to obtain the quality perception features, which can improve the accuracy of the quality perception features. The updated quality evaluation network has excellent performance. When using the quality evaluation network to evaluate the quality perception features of the distorted image, it can effectively improve the accuracy of image quality evaluation.

[0206] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0207] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0208] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0209] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0210] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to the image quality evaluation method / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0211] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0212] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0213] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention.

Claims

1. An image quality assessment method, characterized in that, include: Acquire multiple distorted images; The high-frequency, mid-frequency, and low-frequency components of each distorted image are obtained separately. The feature coding network is updated based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, and the updated feature coding network is used as the distortion information perception network. The distortion information features of each distorted image are obtained using the distortion information perception network. For each distorted image, the essential information features of the distorted image are obtained, and the essential information features and the distortion information features are concatenated to obtain quality perception features; the essential information features are used to describe the basic physical information of the distorted image. The fully connected network is updated using the quality-aware features of all the distorted images, and the updated fully connected network is used as the quality evaluation network. The quality-perceived features of the distorted image to be evaluated are obtained, and the quality-perceived features of the distorted image to be evaluated are evaluated using the quality evaluation network to obtain the quality score of the distorted image to be evaluated. The feature encoding network includes a feature extraction module and a feature mapping module connected in sequence. The update of the feature coding network based on the high-frequency, mid-frequency, and low-frequency components of all distorted images includes: Based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, multiple positive sample pairs are determined. For each distorted image, the feature extraction module is used to extract features from the high-frequency, mid-frequency, and low-frequency components of the distorted image to obtain high-frequency features, mid-frequency features, and low-frequency features. The feature mapping module is used to map the high-frequency, mid-frequency, and low-frequency features of each distorted image to the same feature space; A distortion information loss function is constructed based on the high-frequency, mid-frequency, and low-frequency features of all distorted images, as well as all positive sample pairs. The feature encoding network is updated using the distortion information loss function so that the features corresponding to positive sample pairs are as close as possible in the feature space, while the features corresponding to negative sample pairs are as far away as possible. The loss function for the distorted information is: in, This represents the value of the loss function for distorted information. ,when hour, Represents the high-frequency loss function, when hour, Represents the intermediate frequency loss function, when hour, Represents the low-frequency loss function. The discriminant loss function is: in, Indicates including the first The first distorted image The set of positive sample pairs of each component. express The number of elements in the middle, Indicates including the first The first distorted image The set of positive sample pairs of each component. express The number of elements in the middle, , This represents the total number of the distorted images, when At that time, the first This component represents a high-frequency component, when At that time, the first This component represents the intermediate frequency component, when At that time, the first This component represents the low-frequency component. Denotes the first loss function. Representing the second loss function: in, Indicates the first The first distorted image Characteristic of components, Indicates the first The first distorted image Characteristic of components, Indicates the first The first distorted image Characteristic of components, Indicates the first The first distorted image Characteristic of components, Indicates the first The first distorted image The component features are one of the high-frequency features, mid-frequency features, and low-frequency features. Indicates the first The first distorted image Characteristics of the components , Indicates an indicator function, when or The value is 1 at time. Represents the cosine similarity function. This represents the temperature parameter.

2. The image quality evaluation method according to claim 1, characterized in that, The step of acquiring the high-frequency, mid-frequency, and low-frequency components of each distorted image includes: Through the formula: Calculate the first High-frequency components of a distorted image ; in, Indicates the first A distorted image, Indicates the first Pixels in a distorted image Horizontal position , Indicates the first The width of the distorted image, Indicates the first Pixels in a distorted image The vertical position, , Indicates the first The height of a distorted image , This represents the total number of the distorted images. Represents the Gaussian function. and All are standard deviations; Through the formula: Calculate the first The mid-frequency component of a distorted image ; in, Indicates standard deviation; Through the formula: Calculate the first Low-frequency components of a distorted image ; in, It represents the standard deviation.

3. The image quality evaluation method according to claim 1, characterized in that, The step of updating the feature encoding network using the distortion information loss function includes: Determine whether the value of the distortion information loss function reaches the preset value of the distortion information loss function; If so, then the feature coding network is used as the updated feature coding network; Otherwise, adjust the values ​​of the parameters in the feature encoding network and return to the step of extracting features from the high-frequency, mid-frequency, and low-frequency components of the distorted image using the feature extraction module for each distorted image to obtain high-frequency features, mid-frequency features, and low-frequency features.

4. The image quality evaluation method according to claim 1, characterized in that, Each distorted image has a distortion type; The method determines multiple positive sample pairs based on the high-frequency, mid-frequency, and low-frequency components of all distorted images, including: Iterate through the high-frequency, mid-frequency, and low-frequency components of all distorted images, performing the following steps: Two high-frequency components with the same distortion type are treated as a positive sample pair, and two high-frequency components with different distortion types are treated as a negative sample pair. Two intermediate frequency components with the same distortion type are treated as a positive sample pair, and two intermediate frequency components with different distortion types are treated as a negative sample pair. Two low-frequency components with the same distortion type are treated as a positive sample pair, and two low-frequency components with different distortion types are treated as a negative sample pair. A high-frequency component and a mid-frequency component are used as a negative sample pair. A high-frequency component and a low-frequency component are used as a negative sample pair. A mid-frequency component and a low-frequency component are used as a negative sample pair.

5. The image quality evaluation method according to claim 1, characterized in that, The step of updating the fully connected network using the quality-aware features of all the distorted images includes: For each distorted image, the quality-perceived features corresponding to the distorted image are evaluated using the fully connected network to obtain the quality score of the distorted image. The evaluation loss function value is calculated based on the quality score of the distorted image, and it is determined whether the evaluation loss function value reaches the preset value of the evaluation loss function. If so, then the fully connected network shall be used as the updated fully connected network; Otherwise, adjust the parameters in the fully connected network and return to the step of evaluating the quality-perceived features corresponding to the distorted image using the fully connected network for each distorted image to obtain the quality score of the distorted image.

6. The image quality evaluation method according to claim 5, characterized in that, The calculation of the evaluation loss function value based on the quality score of the distorted image includes: Through the formula: Calculate the evaluation loss function value ; in, Indicates the first The quality score of a distorted image. Indicates the first The true quality score of a distorted image. , This represents the total number of distorted images.

7. The image quality evaluation method according to claim 1, characterized in that, The acquisition of quality-perceived features of the distorted image to be evaluated includes: The distortion information features of the distortion image to be evaluated are obtained using the distortion information perception network. Obtain the essential information features of the distorted image to be evaluated; The essential information features and distortion information features of the distorted image to be evaluated are concatenated to obtain the quality perception features of the distorted image to be evaluated.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image quality evaluation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Distorted image quality evaluation method and device, computer equipment and storage medium

    CN112950597A

  • Quality evaluation method based on self-supervised feature extraction, storage medium and terminal

    CN116416216A