No-reference high-dynamic range image quality evaluation method
Through multi-scale Retinex processing and deep learning feature extraction, combined with principal component analysis and support vector regression model, the problem of high dynamic range image quality evaluation is solved, and more accurate and objective evaluation results are achieved.
Patent Information
- Application Number
- CN202510233241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively evaluate the quality of high dynamic range images, especially in the absence of reference images, where traditional methods cannot accurately capture the depth features and semantic information of the image.
Multi-scale Retinex processing is used to obtain gradient similarity maps, and the depth feature map is extracted in combination with the pre-trained VGG16 network. The image quality evaluation value is calculated through principal component analysis and support vector regression model.
It realizes a more accurate and objective evaluation of high dynamic range image quality, improves the model's evaluation accuracy of image quality, and enhances the adaptability and generalization ability to images of different contents.
Smart Images

Figure CN120147271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image quality assessment, and particularly to a no-reference high dynamic range image quality assessment method. Background Art
[0002] High Dynamic Range (HDR) images have a wider brightness range and richer detail representation, and can more realistically reproduce the lighting scenes in the real world. With the continuous progress of image technology and hardware support, high dynamic range imaging has gradually become an industry standard and is widely used in fields such as film production, game development, and virtual reality, and is deeply loved by consumers. However, high dynamic range images also face some challenges during generation and processing. Similar to low dynamic range images, high dynamic range content may also be distorted during acquisition, compression, and transmission. These problems not only affect the image quality but may also reduce the user's viewing experience.
[0003] Effective evaluation methods can help researchers and engineers evaluate the visual quality of images, optimize image processing algorithms, and ensure that the images finally presented to users achieve the best effects. Compared with low dynamic range images, the human eye is more sensitive to distortions in high dynamic range images. In addition, new types of distortions may be introduced during the encoding and transmission of high dynamic range content. For example, dynamic range compression during backward-compatible encoding may result in color loss or unnatural visual effects. Therefore, it is particularly important to develop new models to more accurately evaluate the quality of high dynamic range images.
[0004] The patent with the application number CN201810042480.6 discloses a no-reference high dynamic range image objective quality assessment method, which represents an image as a third-order tensor. Since chromaticity information plays an important role in high dynamic range image quality assessment, the Tucker decomposition algorithm in tensor decomposition is used to decompose the distorted high dynamic range image, and the first channel that combines luminance distortion and chromaticity distortion is obtained as the first feature image. Distortion information is extracted from the first feature image. Compared with extracting distortion information only on the luminance channel, the first feature image also contains chromaticity channel distortion, and at the same time, the data volume is the same as that of the luminance channel, without increasing additional data volume; the tensor-domain perceptual feature vector extracted from the first feature image is combined with a support vector regression training model to obtain the objective quality evaluation value of the distorted high dynamic range image, thereby realizing the objective quality assessment of no-reference high dynamic range images, significantly improving the evaluation effect, and not requiring a reference image.
[0005] In principle, a dedicated high-dynamic-range camera is required for image acquisition. However, for the general public, this type of camera is too expensive. A common alternative is to use an ordinary camera to acquire images and use dedicated algorithms to generate high-dynamic-range images. The first method is to take pictures of the same scene with multiple exposure levels, and then fuse the images with multiple exposure levels into a high-dynamic-range image. Another type of method is to perform inverse tone mapping on a single captured low-dynamic-range image to generate a high-dynamic-range image. The inverse tone mapping method for a single image requires restoring the lost dynamic information, which is very difficult and more challenging. In recent years, with the continuous application of new deep learning algorithms, the quality of high-dynamic-range image generation has also been continuously improved. However, traditional quality assessment algorithms often consider this type of image enhancement as an error. Traditional image quality assessment methods mainly target low-dynamic-range images. Accurate objective high-dynamic-range image quality assessment helps to evaluate current various high-dynamic-range content generation algorithms and can also be integrated into various inverse tone mapping operators for parameter optimization. In practical applications, it is often difficult to obtain high-dynamic-range reference images with a large data capacity. For example, high-dynamic-range images may be generated by performing inverse tone mapping on a single low-dynamic-range image or by fusing multiple low-dynamic-range images with different exposure levels. In these cases, there is no available high-dynamic-range reference image. Therefore, it becomes particularly important to develop a reference-free quality assessment method for high-dynamic-range images.
[0006] The patent with the application number CN202110166344.X discloses a high-dynamic-range image quality evaluation method based on cumulative gradient similarity. This method calculates the gradient similarity between a low-dynamic-range image sequence and a high-dynamic-range image, obtains the cumulative gradient similarity of the gradient blocks of the low-dynamic-range image and the high-dynamic-range image, and uses the K-means clustering method to perform binary classification on the cumulative gradient similarity, dividing it into static image blocks and ghost image blocks. The gradient similarities of the static image blocks and the ghost image blocks are calculated respectively, and finally they are fused to obtain the image quality evaluation result of the high-dynamic-range image; the present invention uses the cumulative gradient similarity between the low-dynamic-range image sequence and the high-dynamic-range image to divide the high-dynamic-range image into static image blocks and ghost image blocks, calculates the gradient similarities respectively and fuses them to obtain the image quality evaluation result of the high-dynamic-range image, improving the accuracy of high-dynamic-range image quality evaluation.
[0007] Reference-free methods need to rely on an in-depth understanding of image features and quality to infer the visual quality of images. However, the technical solutions of the above two inventions only calculate, correct, and train low-level perceptual features, but lack deep feature maps, thus losing rich information related to image semantics. They are more limited by information when objectively evaluating the quality of reference-free high-dynamic-range images and cannot conduct a more objective and perfect evaluation. Summary of the Invention
[0008] In view of the deficiencies in the above-mentioned existing technologies, the purpose of the present invention is to provide a method for evaluating the quality of high-dynamic-range images without reference. To achieve the above purpose, the present invention provides the following technical solutions: A method for evaluating the quality of high-dynamic-range images without reference, comprising the following steps: Step 1, based on the high-dynamic-range image, obtain the original luminance image of the high-dynamic-range image, perform multi-scale Retinex processing on the original luminance image, calculate the gradient similarity map between the original luminance image and the reflection images at different scales, and capture the gradient change information at different scales; Step 2, input the high-dynamic-range image after multi-scale Retinex with color restoration into a pre-trained VGG16 network to extract the deep feature map; Step 3, aggregate the gradient similarity map and the deep feature map to obtain a multi-dimensional vector based on channel-wise sum pooling, perform vector dimensionality reduction using the principal component analysis method, and input the reduced-dimensional feature vector into a support vector regression model to calculate the image quality evaluation value.
[0009] As a preference of the technical solution of the present invention, the high-dynamic-range image is converted from the RGB color space to the LAB color space to obtain the original luminance image, the original luminance image is decomposed into the reflection image and the illumination image, and in Step 1, the gradient similarity map and the perceptual features are calculated based on the reflection image.
[0010] As a preference of the technical solution of the present invention, in the Retinex model, the image can be described as the product of the illumination image and the reflection image, expressed as follows: where and are the spatial position coordinates of the image pixels, represents the color channel, and represent the illumination image and the reflection image respectively, and the illumination image is generally generated by Gaussian filtering.
[0011] As a preference of the technical solution of the present invention, in Step 1, the multi-scale Retinex model processing uses Gaussian filters with multiple different standard deviation values for filtering to obtain reflection images in multiple different scale cases, and image enhancement is achieved by weighted summation of the reflection images, and the formula is as follows: where, is the output of the multi-scale Retinex, is the weight of the th scale, is the th scale of the reflected image, is the number of different scales used, , and each scale corresponds to a Gaussian filter with a specific standard deviation value; The calculation method of the gradient map of the reflected map at a certain scale is as follows: where, represents the th scale, represents the original luminance map. and respectively represent the horizontal and vertical kernels of the Sobel filter.
[0012] As a preference of the technical solution of the present invention, the gradient similarity between the original luminance image and the reflected map of scale ( ∈{1, 2, 3}) can be calculated by the following formula: where, is a small constant.
[0013] As a preference of the technical solution of the present invention, MSRCR ensures the restoration of the original color ratio of the image by adjusting the color components, and introduces a color restoration factor, as shown in the following formula: where represents the three color channels of red, green, and blue, represents the gain constant, represents the adjustment factor.
[0014] As a preference of the technical solution of the present invention, based on the pre-trained VGG16 network on a large-scale dataset, the last three fully connected layers in the VGG16 network are removed, the high-dynamic range image is input, and the depth feature map is extracted from the last pooling layer. The depth feature map is: where and respectively represent the height and width of the feature map, represents the number of feature maps.
[0015] Preferably, as a technical solution of the present invention, gradient similarity maps at multiple scales are fused to obtain a two-dimensional matrix , which is used to describe the multi-scale visual perception feature map: wherein, ⊙ represents the Hadamard operator; Each of the depth feature maps describes the information of a local area of the high-dynamic range image, and all the depth feature maps together constitute the whole of the high-dynamic range image. Calculate the sum of the pixel values of all the depth feature maps: wherein, , and respectively represent the height and width of the depth feature map, and perform normalization calculation: Finally, perform weighted processing: .
[0016] Preferably, as a technical solution of the present invention, The matrix is adjusted to the same height and width as the depth feature Figure 1 : Integrate the depth feature map and the obtained after adjusting the scale: .
[0017] Preferably, as a technical solution of the present invention, perform sum pooling along the channel dimension to obtain a vector containing multiple elements: where the number of elements is the same as the number of the depth feature maps, and perform dimensionality reduction processing on the vector through principal component analysis: Finally, input the dimensionality-reduced feature vector into a support vector regression model for calculation to obtain a quality evaluation result.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: Innovatively fuse low-level perception features and high-level reasoning features, and fully consider the perception mechanism of the human visual system for image quality. This dual-level feature fusion strategy not only improves the evaluation accuracy of the model for image quality, but also enhances the adaptability and generalization ability of the algorithm to high-dynamic range images with different contents.
[0019] The multi-scale Retinex decomposition technique is adopted to generate the reflection map of the high-dynamic-range image. By calculating the gradient similarity between the reflection map and the original luminance image, the effective simulation of the human cognitive process is realized and used as the perceptual characteristic. This method not only enhances the model's ability to capture image details but also better processes the distortion caused by dynamic range compression.
[0020] 3. Extract the deep feature map from the final convolutional activation layer of the pre-trained VGG16 network, making full use of the rich semantic information it contains. These deep features can accurately represent the spatial relationships between different objects and scenes in the image, providing higher-level visual reasoning information for image quality assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings: Figure 1 It is the logic block diagram of the present invention.
[0022] Figure 2 It is the schematic diagram of the logarithmic domain Retinex decomposition of the present invention.
[0023] Figure 3 It is the comparison diagram of three different compressed HDR pictures in the embodiment of the present invention.
[0024] Figure 4 It is the average gradient similarity diagram in the embodiment of the present invention.
[0025] Figure 5 It is the structural diagram of the VGG16 deep learning network in the embodiment of the present invention.
[0026] Figure 6 It is the tone mapping result diagram of three different methods for the same high-dynamic-range image in the embodiment of the present invention.
[0027] Figure 7 It is two different content images and their corresponding deep feature maps in the last pooling layer of VGG16 in the embodiment of the present invention.
[0028] Figure 8 It is the comparison diagram of results with different feature dimensions on the Narwaria and Korshunov datasets in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0030] The components of the embodiments of the present invention that are generally described and illustrated in the accompanying drawings herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0031] Hereinafter, the terms "comprising", "having" and their cognates that can be used in various embodiments of the present invention are only intended to represent specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as precluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or the possibility of adding one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.
[0032] In addition, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be understood as indicating or implying relative importance.
[0033] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the various embodiments of the present invention belong. The terms (such as those defined in a general-use dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present invention.
[0034] The rapid development of high-dynamic-range imaging technology has brought richer visual experiences to users, but also brought new challenges. In this context, it is particularly important to develop an efficient quality assessment algorithm applicable to high-dynamic-range images. Objective quality assessment methods are generally divided into three categories: full-reference, semi-reference, and no-reference. Full-reference and semi-reference image quality assessment methods rely on all or part of the information of a high-dynamic-range reference image respectively. Full-reference algorithms have achieved certain success in high-dynamic-range image quality assessment and can provide satisfactory evaluation results. However, in practical applications, high-dynamic-range reference images with large data volumes are often difficult to obtain. For example, high-dynamic-range images may be generated by inverse tone mapping of a single low-dynamic-range image or by fusing multiple low-dynamic-range images with different exposure degrees, and there is no available high-dynamic-range reference image in these cases. Therefore, it becomes particularly important to develop a no-reference quality assessment method for high-dynamic-range images. However, the current no-reference quality assessment methods for high-dynamic-range images are still relatively scarce.
[0035] Due to the lack of reference high-dynamic-range images, no-reference methods need to rely on an in-depth understanding of image features and quality to infer the visual quality of images. This usually requires a detailed modeling of the human visual system to ensure that the extracted features can effectively reflect the true situation of visual perception. To ensure that the quality prediction is closely consistent with the subjective score, it is crucial to incorporate the perceptual factors of the human visual system into the design of the objective algorithm. The human visual perception system processes perception and reasoning in a hierarchical manner, which provides an important basis for the development of the algorithms in this chapter.
[0036] Refer to Appendix Figure 1 to Appendix Figure 7 As shown, the framework of the no-reference high-dynamic-range image quality assessment algorithm of this application. First, the high-dynamic-range image is converted from the RGB color space to the LAB color space. In this space, the luminance map (L channel) is processed using multi-scale Retinex (MSR) decomposition. The gradient similarity map between the original luminance map and the multi-sensitive reflection maps at different scales is calculated to capture the gradient change information at different scales, representing the perceptual characteristics of the process. Subsequently, multi-scale Retinex with color restoration (MSRCR) is applied to obtain the enhanced image, and this enhanced image is input into the pre-trained VGG16 network to extract the deep feature maps from the 5th pooling layer of the deep network. These feature maps are defined as inference features, providing high-level abstract information that is very important for quality assessment. Finally, the perceptual features from the gradient similarity map and the inference features from VGG16 are aggregated and input into the support vector regression model to predict the image quality score. This method effectively combines traditional image processing techniques with modern deep learning, providing a reliable framework to accurately predict the HDR image quality score.
[0037] A no-reference high-dynamic-range image quality assessment method, comprising the following steps: Step 1, based on the high-dynamic-range image, obtain the original luminance image of the high-dynamic-range image, perform multi-scale Retinex processing on the original luminance image, calculate the gradient similarity map between the original luminance image and the reflection images at different scales, and capture the gradient change information at different scales as the perceptual information related to quality.
[0038] Step 2, after the high-dynamic-range image is enhanced by multi-scale Retinex with color restoration, input it into the pre-trained VGG16 network to extract the deep feature maps.
[0039] Step 3, aggregate the gradient similarity map and the deep feature maps to obtain a multi-dimensional vector that sums and pools along the channel dimension, perform vector dimensionality reduction using the principal component analysis method, and input the reduced-dimensional feature vector into the support vector regression model to calculate and obtain the image quality assessment value.
[0040] The term "Retinex" is a combination of the words "retinal" and "cortex", reflecting the combined action of the eye and the brain in processing. This theory was initially proposed to explain color constancy under different illumination conditions. The theory holds that when the human eye observes an object, it can, to a certain extent, ignore the changes in ambient light, thus maintaining the stability of the object's color. This phenomenon is very important in visual perception because it enables people to recognize and judge the true color of an object under different lighting conditions.
[0041] After converting the high-dynamic-range image from the RGB color space to the LAB color space, the original luminance image is obtained. The original luminance image is decomposed into a reflectance image and an illumination image. In step 1, a gradient similarity map and perceptual features are calculated based on the reflectance image. An image can be represented as the product of a reflectance image and an illumination image, and this expression reveals the two basic components of the image: the reflectance image represents the inherent color and features of the object's surface, while the illumination image describes the influence of the light source on the object's surface. By decomposing the image into these two parts, the visual characteristics of the image can be understood more deeply. The illumination component of the image usually changes slowly and determines the dynamic range of the original image. In contrast, the reflectance component captures the inherent properties of the image and represents the features and details of the object's surface.
[0042] The image pixel information of the high-dynamic-range image is extracted to calculate the reflectance image. In the multi-scale Retinex model (MSR), the image can be described as the product of an illumination image and a reflectance image, and the formula is as follows: where and are the spatial position coordinates of the image pixels, represents the color channel, and represent the illumination image and the reflectance image respectively. The illumination component of the image usually changes slowly and determines the dynamic range of the original image. In contrast, the reflectance component captures the inherent properties of the image and represents the features and details of the object's surface.
[0043] To simplify the calculation, the formula is usually transformed into the logarithmic domain: In the logarithmic domain of the high-dynamic-range image, the multiplication operation can be transformed into an addition operation, which helps to reduce the computational complexity. At the same time, logarithmic compression of linear luminance also conforms to the compressive perception characteristics of the human eye's vision for linear luminance. The Retinex logarithmic decomposition is as shown in the appendix Figure 2As shown, in the Retinex decomposition process, a Gaussian filter is usually used to implement low-pass filtering to obtain a luminance map. The expression of the Gaussian filter is: In step 1, the MSR model processes and filters with Gaussian filters of multiple different standard deviation values to generate reflection images in multiple different scales. In reality, since high-dynamic-range images may be generated by inverse tone mapping of a single low-dynamic-range image or by fusing multiple low-dynamic-range images with different exposure degrees, obtaining multiple reflection images with multiple different standard deviations can perform a more realistic quality assessment in subsequent calculation steps. Image enhancement is achieved by weighted summation of the reflection images, and its formula is as follows: Where, is the output of the multi-scale Retinex, is the weight of the th scale, is the th scale of the reflection image, is the number of different scales used, , and each scale corresponds to a Gaussian filter with a specific standard deviation value. Each scale corresponds to a Gaussian filter with a specific standard deviation value, enabling the method to capture details at different scales.
[0044] In this embodiment, the MSR model adopts three different Gaussian filter scales, and their standard deviation values are empirically set to 250, 80, and 10.
[0045] The calculation method of the gradient map of the reflection map of a certain scale is as follows: Where, represents the th scale, represents the original luminance map. And respectively represent the horizontal and vertical kernels of the Sobel filter.
[0046] The gradient similarity between the original luminance image and the reflection map of scale ( ∈{1, 2, 3}) can be calculated by the following formula: Where, is a small constant used to avoid a denominator of zero in the equation.
[0047] Since high-dynamic-range content cannot be directly and accurately displayed on ordinary display devices, Figure 2 shows the tone mapping results of three high-dynamic-range images using built-in Matlab functions. These high-dynamic-range images are selected from the Narwaria database and have the same content but different coding rates. Figure 3 (a) shows a highly compressed high-dynamic-range image, where obvious blur and block artifacts are visible, as shown by the red and green borders. In addition, banding artifacts can be found in the sky, as shown by the yellow border. Figure 3 (b) shows the content of a moderately compressed high-dynamic-range image, which shows more details compared to Figure 3 (a). For example, Figure 3 the area within the red border in (a) is severely blurred or affected by block artifacts, while the corresponding area in Figure 3 (b) shows only slight blur at the bottom. Figure 3 (c) is from the high-dynamic-range reference image, showing perfect details and no visible artifacts. Figure 3 The mean opinion scores for (a), (b), and (c) are 1.077, 3.077, and 4.808 respectively, where 4.808 is the highest quality score and 1.077 is the lowest quality score.
[0048] As shown in Figure 4 shows Figure 3 the trend line of the average gradient similarity values of the high-dynamic-range images in Figure 3 (a), revealing the differences in detail retention and quality assessment of these images. For Figure 3 (a), the maximum similarity value and the flattest trend line are observed, indicating that the gradient similarity of this image is relatively high, which should be due to the fact that Figure 3 (a) is a highly compressed high-dynamic-range image lacking rich detail information, resulting in very limited similarity changes between Gaussian filters of different scales. Relatively speaking,
[0049] (c) shows the smallest similarity value but the highest score at the same time. This phenomenon indicates that a lower gradient similarity value can precisely reflect its rich details and sharpness.
[0050] Multi-Scale Retinex with Color Restoration (MSRCR) is an enhancement of the basic MSR algorithm. It adds a color restoration step on top of MSR to preserve natural colors and avoid the gray world effect that standard Retinex often produces. MSRCR ensures the restoration of the original color ratio of the image by adjusting the color components and introducing a color restoration factor, as shown in the following equation: where represents the three color channels of red, green, and blue, represents the gain constant, which is empirically set to 46. represents the adjustment factor, which is empirically set to 125.
[0051] VGG16 is a widely recognized CNN architecture. As shown in the appendix Figure 5 , VGG16 consists of 13 convolutional layers, 3 fully connected layers, and 5 pooling layers. VGG16 is known for its simplicity and effectiveness and is good at learning rich features from images, so it has excellent performance in tasks such as image classification and object detection. Nowadays, the VGG16 network pre-trained on large-scale datasets (such as ImageNet) has become the preferred choice for transfer learning and deep feature extraction. In this embodiment, the deep features extracted from the high-dynamic-range image using the pre-trained VGG16 network are used as inference features. To adapt to input images of any size, the last three fully connected layers of VGG16 are removed. Subsequently, the deep features are extracted from the last pooling layer, generating 512 feature maps.
[0052] As shown in the appendix Figure 6 , it is generated using the built-in tone mapping function in MATLAB. Although the image details are well preserved, the image lacks a natural feeling. Appendix Figure 6 (b) and appendix Figure 6 (c) are generated using linear mapping and the MSRCR method respectively, both in double-precision format. The scale settings of MSRCR are consistent with those used in the MSR part. Compared with appendix Figure 6 (b), appendix Figure 6 (c) shows better dark detail and color information. In addition, using double-precision data as input helps improve the accuracy and effect of the calculation. Therefore, MSRCR is used as a preprocessing step for feature extraction when the image is input into the VGG16 network.
[0053] As shown in the appendix Figure 7As shown, two high-dynamic range images with different contents and their depth feature maps are presented. The leftmost column shows the content of the high-dynamic range image after MSRCR processing. The image is enhanced in terms of color and brightness and presents a more natural visual experience. The following series of content shows the depth feature maps extracted from the last pooling layer of the VGG16 network. The feature maps in each column represent different levels of image information. The numbers at the bottom show the position indices of the feature maps. The symbol "……" indicates the parts not shown. Due to space limitations, not all feature maps are fully presented. Through comparison, it can be clearly seen that there are obvious differences between the depth feature maps with the same position index. The depth feature maps corresponding to different image contents are different in structure and details. These depth feature maps contain rich information related to the semantics of the images.
[0054] Based on the VGG16 network pre-trained on a large-scale dataset, the last three fully connected layers in the VGG16 network are removed, and a high-dynamic range image is input. The depth feature maps are extracted from the last pooling layer. The depth feature maps are: where and represent the height and width of the feature map respectively, represents the number of feature maps, and .
[0055] The scale settings in the multi-scale Retinex decomposition with color restoration are consistent with those in MSR.
[0056] As shown in the attached Figure 4 , there is a correlation between the average gradient similarity value of MSR at each decomposition scale and the image quality. A higher similarity value corresponds to a lower mean opinion score, while a lower similarity value corresponds to a higher mean opinion score.
[0057] By using the Hadamard operation to fuse the gradient similarity maps at three scales, a two-dimensional matrix is obtained. This matrix is used to describe the multi-scale visual perception feature map: where ⊙ represents the Hadamard operator.
[0058] From the attached Figure 7As shown, each depth feature map describes the information of a local region of the image and is an important part of the image representation. When capturing visual information at different levels of the image, each feature map may focus on different contents. These feature maps together constitute a comprehensive representation of the image, covering rich high-level semantic information. To better highlight the key features in the image, a pixel sum-value weighted feature map is adopted. This weighting operation can not only effectively focus on the key parts of the image, but also enhance the expressive power of the image representation, enabling the target features or semantic information in the image to be better reflected in subsequent tasks (such as image quality assessment, classification, etc.).
[0059] Each depth feature map describes the information of a local region of the high-dynamic range image, and all depth feature maps together constitute the whole of the high-dynamic range image. Calculate the sum of the pixel values of all depth feature maps: Among them, , and respectively represent the height and width of the depth feature map, and perform normalization calculation: Finally, perform weighting processing: .
[0060] Adjust the matrix to the same height and width as the depth feature Figure 1 : Integrate the depth feature map and the obtained after adjusting the scale: .
[0061] Perform sum pooling along the channel dimension to obtain a vector containing multiple elements: The number of elements is the same as the number of depth feature maps, and perform dimensionality reduction on the vector through principal component analysis: This method can effectively reduce the complexity of the model, avoid overfitting caused by too high dimensionality of the feature vector, and at the same time retain the most important features in the data. In this embodiment, it is preferably to finally reduce the 512-dimensional feature to a 32-dimensional reduced vector.
[0062] In the step of predicting the image quality score, regression models, including support vector regression, random forest, and neural network, etc., have been widely used to map image features to quality scores. Among these methods, SVR stands out due to its flexibility in choosing kernel functions, powerful ability to handle high-dimensional features, and remarkable generalization performance. Especially in scenarios with a small sample size and high data dimension, support vector regression shows unique advantages compared to other machine learning algorithms. Like most training-test methods, the dataset is usually divided into two subsets: the training set and the test set, usually by random sampling. In the training phase, a support vector regression evaluation model is generated based on the 32-dimensional feature vectors of the training set images and their corresponding mean opinion scores.
[0063] By collecting features related to the quality of high-dynamic-range images and constructing a support vector regression (SVR) model using a Radial Basis Function (RBF) kernel, the quality of high-dynamic-range images can be efficiently evaluated. To reduce bias, the training-test process is executed 1,000 times, and the median value is used as the final result. Each time it is executed, the entire database is randomly divided into two parts according to the image content: 80% of the samples are used as the training set, and the remaining 20% are used as the test set.
[0064] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, and the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0065] Experimental results and comparative analysis based on the above technical solutions: Algorithm evaluation criteria To evaluate the performance, three commonly used evaluation metrics were adopted: Root Mean Square Error (RMSE), Pearson Linear Correlation Coefficient (PLCC), and Spearman Rank Correlation Coefficient (SRCC). The Root Mean Square Error is a commonly used metric to measure the difference between predicted values and observed values. It quantifies the average error magnitude in a set of predictions. The lower the Root Mean Square Error value, the better the prediction accuracy. The Pearson Linear Correlation Coefficient measures the linear correlation between two datasets and provides the degree to which they vary together. The value of the Pearson Linear Correlation Coefficient ranges from -1 to 1, where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation. The Spearman Rank Correlation Coefficient evaluates the strength and direction of the association between two ranked variables. Different from the Pearson Linear Correlation Coefficient, the Spearman Rank Correlation Coefficient evaluates monotonic relationships and thus has strong robustness to non-linear correlations. Its value ranges from -1 to 1, and the higher the value, the stronger the positive correlation. Before calculating the PLCC and RMSE, it is necessary to remove the non-linearity of the target scores through logistic regression, defined as follows: where is the predicted score of the objective algorithm for the input, and 𝑎1, 𝑎2, 𝑎3, 𝑎4, and 𝑎5 are the parameters fitted through non-linear regression.
[0066] Datasets and Comparative Algorithms The algorithm performance was evaluated comparatively on two public databases: the Narwaria database and the Korshunov database. The Narwaria database contains ten reference high-dynamic-range images covering indoor and natural scenes. After the images are processed by iCAM06 tone mapping and JPEG compression, distorted HDR images are generated. By adjusting the compression settings and iCAM06 parameters, 14 different distorted images are generated. This database uses MOS scores from 1 to 5 to evaluate image quality, and a higher score indicates better quality. The Korshunov database contains twenty original HDR images with diverse content such as architecture, landscape, and portraits. The distorted images are generated according to the JPEG-XT standard, using different profiles and quality levels, covering a wider range of image content and compression quality levels. Each image is assigned a damage level from 1 to 5, and a higher value indicates greater damage.
[0067] To verify the effectiveness and reliability of the proposed algorithm, a comparative experiment was designed for two mainstream high-dynamic-range (HDR) image quality assessment models. The first type of method to be compared first converts the linear luminance values of HDR images into perceptually uniform values through PU compression, and then applies mature low-dynamic-range (LDR) image quality assessment methods to complete the quality assessment task. The LDR image quality assessment methods used include full-reference algorithms and no-reference algorithms. For the full-reference LDR image quality assessment algorithms, classic LDR image quality assessment methods such as peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and visual information fidelity (VIF) were selected. For the no-reference algorithms, methods such as blind image quality assessment (BRISQUE) and natural image quality evaluation (NIQE) were chosen. These methods do not rely on reference images and can evaluate the quality of HDR images without the original images. The second type of method focuses on specially designed HDR image quality assessment algorithms, which are directly optimized and evaluated according to the characteristics of HDR images. We selected Guan's method, high-dynamic-range visual difference prediction model (HDR-VDP-2.2), and Zhang's method as the comparison algorithms. These methods are all carefully designed to consider the special effects of factors such as luminance, contrast, and color in HDR images, and have been proven to be highly accurate and reliable in HDR image quality assessment. The tone mapping image quality assessment algorithm (hereinafter referred to as TMIQA) extracts image features based on the perceptual process and pays special attention to the image features in the highlight and lowlight regions. According to the damage characteristics of high-dynamic images, TMIQA is applicable to the quality assessment of tone-mapped images, so it is also evaluated as a comparison algorithm in this study.
[0068] Through the comparative experiment, it aims to comprehensively evaluate the performance of the proposed algorithm under different evaluation frameworks, and compare its accuracy, robustness, and computational efficiency when processing HDR images. The first type of comparative experiment will help us understand the applicability and limitations of traditional LDR quality assessment methods in HDR images, while the second type of comparative experiment can demonstrate the effects and advantages of the evaluation methods specially designed for HDR images. Finally, by comprehensively comparing these methods, the superiority of the proposed algorithm in processing HDR image quality assessment can be better verified.
[0069] Data and Analysis: The following table shows the comparison results on two public datasets, and the best-performing metrics are shown in bold: From the result data, it can be obtained that: According to the comparison of experimental results, the proposed method not only shows excellent prediction accuracy in the most challenging high-dynamic-range image quality assessment task, but also maintains high consistency and robustness on different datasets. Compared with other competing methods, the method proposed in this chapter is more consistent with the results of subjective evaluation, providing a more reliable and accurate high-dynamic-range image quality assessment.
[0070] Ablation experiments for detecting aggregation effects: High-dynamic-range image quality assessment is achieved by fusing two features. Specifically, the gradient similarity representing perceptual features is used to weight the depth feature map representing reasoning features, thereby extracting features related to image quality. To verify the complementarity of these features, experimental comparisons were made with and without perceptual feature weighting. The following table shows the results, where "Yes" indicates that the depth features are weighted by gradient similarity, and "No" indicates that only the depth features are used. The results show that after fusing these two features, both the PLCC and SRCC metrics on the Narwaria and Korshunov datasets have been significantly improved. These findings emphasize the strong complementary effect between perceptual features and reasoning features.
[0071] Using PCA (Principal Component Analysis) to reduce the dimension of high-dimensional features can not only significantly reduce the computational cost, but also avoid overfitting problems caused by too high feature dimensions, ensuring the generalization ability of the model. Figure 8 Shows the experimental results on the Narwaria and Korshunov datasets when using different feature dimensions. Figure 8 (a) represents the results on the Narwaria dataset. Figure 8 (b) represents the results on the Korshunov dataset. From the image results, on the Narwaria dataset, the 32-dimensional feature combination achieved the best performance, indicating that on this dataset, moderate dimensionality reduction can most fully retain useful information while avoiding excessive computation redundancy. For the Korshunov dataset, as the feature dimension increases, the overall performance has a certain improvement, especially the incremental change is more obvious below 32 dimensions. But when the number of dimensions exceeds 32, the trend of performance improvement flattens out, indicating that after this point, more feature dimensions have a gradually weakened effect on performance improvement. Considering the scale of the dataset and computational efficiency comprehensively, a 32-dimensional feature dimension is considered a balanced choice, which can not only ensure good evaluation results but also maintain low computational overhead.
[0072] Cross-dataset evaluation: To evaluate the generalization ability of the proposed learning algorithm, the model is trained on one dataset (Narwaria or Korshunov) and evaluated on another dataset. The following table shows the cross-dataset evaluation results. All PLCC and SRCC values are close to 0.85, indicating that the algorithm has good generalization performance. In addition, each functional module or unit in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0073] If the above-mentioned functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0074] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.
Claims
1. A method for evaluating the quality of a high dynamic range image without reference, characterized in that: The following steps are involved: Step 1, based on the high dynamic range image, obtaining the original brightness image of the high dynamic range image, performing multi-scale Retinex processing on the original brightness image, calculating the gradient similarity map between the original brightness image and the reflection images of different scales, and capturing the gradient change information at different scales; Step 2, the high dynamic range image is subjected to multi-scale Retinex with color restoration and then input into a pre-trained VGG16 network to extract a deep feature map; Step 3: Aggregate the gradient similarity map and the depth feature map to obtain a multidimensional vector based on summation and pooling along the channel dimension, use principal component analysis to reduce the dimension of the vector, input the reduced dimension feature vector into a support vector regression model, and calculate the image quality assessment value.
2. The method for evaluating quality of a non-reference high dynamic range image according to claim 1, characterized in that: The high dynamic range image is converted from the RGB color space to the LAB color space to obtain the original brightness image, the original brightness image is decomposed into the reflection image and the illumination image, and the gradient similarity map and the perceptual features are calculated based on the reflection image in the step 1.
3. The method for evaluating quality of a non-reference high dynamic range image according to claim 2, wherein: In the Retinex model, the image can be described as the product of the illumination image and the reflectance image, which is expressed as follows: in and is the spatial position coordinate of the image pixel, represents the color channel, and They represent the illumination image and the reflection image respectively. The illumination image is generally generated by Gaussian filtering.
4. The method for evaluating quality of a non-reference high dynamic range image according to claim 3, characterized in that: In step 1, the multi-scale Retinex uses different standard deviations The Gaussian filter of the value is filtered to obtain reflection images at multiple different scales, and the image enhancement is achieved by weighting the reflection images. The formula is as follows: in, is the output of multi-scale Retinex, It is The weight of the scale, It is The reflection image of the scale, is the number of different scales used, , each scale corresponds to a Gaussian filter of value; The calculation method of the gradient map of the reflection map at a certain scale is as follows: in, Indicates scale, Represents the original brightness map. and Represent the horizontal and vertical kernels of the Sobel filter respectively.
5. The method for evaluating quality of a non-reference high dynamic range image according to claim 4, characterized in that: The original brightness image and the scale ( ∈{1,2,3}) The gradient similarity between reflection maps can be calculated by the following formula: in, is a small constant.
6. The method for evaluating quality of a non-reference high dynamic range image according to claim 5, characterized in that: MSRCR adjusts the color components to ensure that the color ratio of the original image is restored and introduces a color restoration factor, as shown in the following formula: in Represents the three color channels of red, green and blue. represents the gain constant, Represents the adjustment factor.
7. The method for evaluating quality of a non-reference high dynamic range image according to claim 5, characterized in that: Based on the VGG16 network pre-trained with a large-scale dataset, the last three computational connection layers in the VGG16 network are removed, the image processed by the MSRCR is input, and the depth feature map is extracted from the last pooling layer. The depth feature map is: middle and Represent the height and width of the feature map respectively, Indicates the number of feature maps.
8. The method for evaluating quality of a non-reference high dynamic range image according to claim 7, characterized in that: The gradient similarity graphs at multiple scales are fused to obtain a two-dimensional matrix , which is used to describe the multi-scale visual perception feature map: Where, ⊙ represents the Hadamard operator; Each of the depth feature maps describes information of a local area of the high dynamic range image, and all of the depth feature maps together constitute the whole of the high dynamic range image. The sum of the pixel values of all of the depth feature maps is calculated: in, , and Respectively represent the height and width of the depth feature map, and perform normalization calculations: Finally, weighted processing is performed: 。 9. The method for evaluating quality of a non-reference high dynamic range image according to claim 8, characterized in that: Will The matrix is adjusted to the same height and width as the depth feature map: The depth feature map and the rescaled To integrate: 。 10. The method for evaluating quality of a non-reference high dynamic range image according to claim 9, characterized in that: Sum pooling is performed along the channel dimension to obtain a vector containing multiple elements: The number of elements is consistent with the number of the depth feature map, and the vector is reduced in dimension by principal component analysis: Finally, the dimension-reduced feature vector is input into a support vector regression model for calculation to obtain a quality assessment result.
Citation Information
Patent Citations
Non-reference high-dynamic range image objective quality evaluation method
CN108322733A
High-dynamic-range image quality evaluation method based on cumulative gradient similarity
CN112734754A