Image Quality Assessment Method, Device and Medium Based on Semi-Reference Edge Detection
Through the image quality evaluation method based on semi-reference edge detection, the cascading two-layer attention HED architecture and deep learning technology are used to solve the adaptability and accuracy of image quality evaluation, and flexible image quality evaluation and edge detection are achieved.
Patent Information
- Application Number
- CN202510398126.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing reference-free image quality evaluation methods have shortcomings in adaptability, accuracy and detail retention, and rely on manual assistance, resulting in unstable evaluation results.
The image quality evaluation method based on semi-reference edge detection is adopted, and edge features of RGB feature images are extracted by obtaining the image to be evaluated for format consistency preprocessing, and edge features of RGB feature images are extracted, and edge distortion detection and evaluation are performed using cascaded bilayer attention HED architecture and cross-entropy loss function analysis, combined with deep learning and attention mechanism.
It realizes flexible image quality evaluation in different scenarios, improves the adaptability and accuracy of evaluation, reduces the dependence on manual assistance, and enhances the robustness and detail retention ability of edge detection.
Smart Images

Figure CN119904462B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image quality assessment method, device, and medium based on semi-reference edge detection. Background Art
[0002] With the rapid development of image processing technologies, image quality assessment (IQA) plays an important role in fields such as image processing and computer vision. Image quality assessment methods are divided into two categories: reference-based and reference-free. The reference-based quality assessment method calculates the quality score by comparing the original image and the image to be evaluated, while the reference-free image quality assessment method only relies on the characteristics of the image to be evaluated for quality assessment.
[0003] There are still problems with existing reference-free quality assessment image quality assessment methods. First, the adaptability to distortion types is poor: many reference-free quality assessment methods cannot effectively handle different types of distortions in images (such as compression, noise, blur), and the evaluation results are not accurate enough. Second, the evaluation method is fixed: the existing evaluation methods cannot dynamically select the most suitable evaluation strategy according to the image characteristics, resulting in unstable evaluation results on different types of images. Third, the lack of detail retention evaluation: the existing technologies tend to ignore the detailed parts of the image, especially the loss of edge information, which affects the accuracy of image quality assessment. Fourth, the dependence on artificial assistance: the existing technologies still rely more or less on artificial assistance, making the objective evaluation method have a gap from a completely intelligent objective evaluation. Summary of the Invention
[0004] The embodiments of this application provide an image quality assessment method, device, and medium based on semi-reference edge detection, which solve the technical problems of poor adaptability, poor accuracy, and single evaluation method in image quality assessment.
[0005] In a first aspect, the embodiments of this application provide an image quality assessment method based on semi-reference edge detection, which is characterized in that the method includes: obtaining an image to be evaluated, and performing format-uniform preprocessing on the image to be evaluated to obtain an RGB feature image; extracting image edge features from the RGB feature image to obtain edge feature data to be detected; constructing a cascaded double-layer attention HED architecture based on an evolutionary algorithm, and inputting the edge feature data to be detected into the cascaded double-layer attention HED architecture based on the evolutionary algorithm to obtain a multi-stage edge prediction map; performing cross-entropy loss function analysis on the multi-stage edge prediction map to obtain an edge information map; performing edge distortion detection on the edge information map to obtain edge detail loss evaluation data; based on the edge detail loss evaluation data, determining the edge feature vector of the distorted image through distorted image edge feature extraction; and obtaining an image quality assessment result through image quality assessment according to the edge feature vector of the distorted image.
[0006] In an implementation manner of the present application, preprocessing is performed on the image to be evaluated to make its format consistent, so as to obtain an RGB feature image, which specifically includes: performing noise removal processing on the image to be evaluated to obtain denoised image data; based on the denoised image data, determining a visually enhanced image through denoised image enhancement; wherein, the denoised image enhancement includes: histogram equalization and contrast enhancement; converting the visually enhanced image into the RGB format to obtain the RGB feature image.
[0007] In an implementation manner of the present application, image edge feature extraction is performed on the RGB feature image to obtain edge feature data to be detected, which specifically includes: performing color feature extraction on the RGB feature image to obtain image color feature data; wherein, the color feature extraction includes: HSV color space conversion and color histogram calculation; performing texture feature extraction on the RGB feature image through a gray-level co-occurrence matrix and a Gabor filter to obtain image texture feature data; based on the image color feature data and the image texture feature data, obtaining the edge feature data to be detected.
[0008] In an implementation manner of the present application, the edge feature data to be detected is input into a cascaded double-layer attention HED architecture based on an evolutionary algorithm to obtain a multi-stage edge prediction map, which specifically includes: inputting the edge feature data to be detected into the cascaded double-layer attention HED architecture based on the evolutionary algorithm to obtain a stage edge prediction map; performing upsampling fusion on the stage edge prediction map to determine a fusion prediction map that conforms to the image to be evaluated; based on the stage edge prediction map and the fusion prediction map, obtaining the multi-stage edge prediction map.
[0009] In an implementation manner of the present application, cross-entropy loss function analysis is performed on the multi-stage edge prediction map to obtain an edge information map, which specifically includes: constructing a cross-entropy loss function based on the multi-stage edge prediction map; wherein, the calculation formula of the cross-entropy loss function is:
[0010]
[0011]
[0012] Wherein, and W are the height and width of the image to be evaluated, Y is the true edge map, is the edge detection map of the th stage, is the fusion prediction map; wherein, and W are the height and width of the image to be evaluated, Y is the true edge map, is the edge detection map of the th stage, is the fusion prediction map, represents the pixel value of the real edge map at the position (h, w), and represent the pixel values of the predicted map and the fused predicted map of the i-th stage at the position (h, w); according to the cross-entropy loss function, the losses corresponding to the multi-stage edge prediction maps are weighted and summed to determine the total loss function of the multi-stage edge prediction maps; among them, the calculation formula of the total loss function of the multi-stage edge prediction maps is:
[0013]
[0014] Among them, is the loss weight coefficient of the stage prediction map of the -th stage, is the loss weight coefficient of the fused prediction map of the -th stage; input the multi-stage edge prediction map into the total loss function of the multi-stage edge prediction map to obtain the edge information map.
[0015] In an implementation manner of the present application, edge distortion detection is performed on the edge information map to obtain edge detail loss evaluation data, which specifically includes: converting the edge information map into edge information data, and performing edge structure similarity analysis on the edge information data to obtain the edge structure similarity; among them, the calculation formula of the edge structure similarity analysis is:
[0016]
[0017] Among them, is the intersection part of the predicted edge region and the real edge region, represents the union part of the predicted edge region and the real edge region; perform edge intensity difference analysis on the edge information map to obtain edge intensity difference data; among them, the edge intensity difference analysis includes: image gradient magnitude analysis, error analysis, and the calculation formula of the image gradient magnitude analysis is:
[0018]
[0019]
[0020] Among them, is the gradient in the x direction, is the gradient in the y direction, is the image to be evaluated; perform edge position offset analysis on the edge information map to obtain edge position offset data; among them, the calculation formula of the edge position offset analysis is:
[0021]
[0022] Among them, is the total number of pixels in the edge information map, For the th pixel in the real edge map of the image, For the position of the th pixel in the predicted edge map; According to the edge structure similarity, edge intensity difference data, and edge position offset data, through comprehensive threshold score analysis, edge detail loss evaluation data is obtained.
[0023] In an implementation manner of the present application, according to the edge feature vector of the distorted image, through image quality evaluation, an image quality evaluation result is obtained, specifically including: based on the preset ResNet convolutional neural network with self-attention mechanism, advanced feature extraction is performed on the edge feature vector to obtain edge advanced features; The advanced edge features are used as the evaluation learning background data of the ResNet convolutional neural network, and are trained to convergence through the mean square error loss function to obtain an image quality evaluation model; The image to be evaluated is input into the image quality evaluation model to obtain an image quality evaluation result.
[0024] In an implementation manner of the present application, after obtaining the image quality evaluation result according to the edge feature vector of the distorted image through image quality evaluation, the method further includes: converting the quality evaluation result into an image quality score and an image distortion situation; Outputting the image quality score and the image distortion situation to a preset visualization system.
[0025] In a second aspect, an image quality evaluation device based on semi-reference edge detection is further provided in an embodiment of the present application, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: obtain an image to be evaluated, and perform preprocessing of consistent format on the image to be evaluated to obtain an RGB feature image; perform image edge feature extraction on the RGB feature image to obtain edge feature data to be detected; construct a cascaded double-layer attention HED architecture based on an evolutionary algorithm, and input the edge feature data to be detected into the cascaded double-layer attention HED architecture based on the evolutionary algorithm to obtain a multi-stage edge prediction map; perform cross-entropy loss function analysis on the multi-stage edge prediction map to obtain an edge information map; perform edge distortion detection on the edge information map to obtain edge detail loss evaluation data; based on the edge detail loss evaluation data, determine the edge feature vector of the distorted image through distorted image edge feature extraction; According to the edge feature vector of the distorted image, through image quality evaluation, an image quality evaluation result is obtained.
[0026] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for image quality assessment based on semi-reference edge detection, storing computer-executable instructions, characterized in that the computer-executable instructions are set to: obtain an image to be evaluated, and perform preprocessing of format normalization on the image to be evaluated to obtain an RGB feature image; extract image edge features from the RGB feature image to obtain edge feature data to be detected; construct a cascaded double-layer attention HED architecture based on an evolutionary algorithm, and input the edge feature data to be detected into the cascaded double-layer attention HED architecture based on the evolutionary algorithm to obtain a multi-stage edge prediction map; perform cross-entropy loss function analysis on the multi-stage edge prediction map to obtain an edge information map; perform edge distortion detection on the edge information map to obtain edge detail loss evaluation data; based on the edge detail loss evaluation data, determine the edge feature vector of the distorted image through edge feature extraction of the distorted image; and obtain an image quality evaluation result through image quality evaluation according to the edge feature vector of the distorted image.
[0027] An embodiment of the present application provides an image quality assessment method, device, and medium based on semi-reference edge detection. Through image feature extraction, deep edge detection, edge distortion detection, and edge feature extraction of distorted images, combined with deep learning and attention mechanism technologies, the technical problems of poor adaptability, poor accuracy, and single evaluation method in image quality assessment are solved, and flexible image quality assessment in different scenarios is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0029] Figure 1 It is a flowchart of an image quality assessment method based on semi-reference edge detection provided by an embodiment of the present application;
[0030] Figure 2 It is a schematic internal structure diagram of an image quality assessment device based on semi-reference edge detection provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] The embodiments of the present application provide an image quality evaluation method, device, and medium based on semi-reference edge detection. By extracting image features, performing depth edge detection, edge distortion detection, and extracting edge features of distorted images, and combining deep learning and attention mechanism technologies, the technical problems of poor adaptability, poor accuracy, and single evaluation method in image quality evaluation are solved, and flexible image quality evaluation in different scenarios is realized.
[0033] The technical solutions proposed in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0034] Figure 1 FIG. is a flowchart of an image quality evaluation method based on semi-reference edge detection provided by an embodiment of the present application. As Figure 1 shown, an image quality evaluation method based on semi-reference edge detection provided by an embodiment of the present application specifically includes the following steps:
[0035] Step 101: Obtain the image to be evaluated, and perform preprocessing of format normalization on the image to be evaluated to obtain an RGB feature map.
[0036] Exemplarily, preprocessing operations such as noise removal and image enhancement are performed on the input image to ensure the quality of the image in subsequent evaluations. In the present application, the image is converted into an RGB image for subsequent feature extraction processing.
[0037] Specifically, performing preprocessing of format normalization on the image to be evaluated to obtain an RGB feature image includes: performing noise removal processing on the image to be evaluated to obtain denoised image data; based on the denoised image data, determining a visually enhanced image through denoised image enhancement; wherein, the denoised image enhancement includes: histogram equalization, contrast enhancement; converting the visually enhanced image into an RGB format to obtain an RGB feature image.
[0038] In one embodiment, noise removal processing is performed on the image to be evaluated, and Gaussian filtering is used to remove the noise in the image, especially for salt-and-pepper noise and Gaussian noise in the image. After the image to be evaluated is input, a suitable filter (such as median filtering or Gaussian filtering) is selected, and filtering processing is performed according to the characteristics of the image noise to remove the noise and maintain the quality of the image. By convolving the image with a Gaussian kernel, the image is smoothed and high-frequency noise is reduced. Most of the image details can be effectively retained.
[0039] Image enhancement uses histogram equalization and contrast enhancement methods to improve the overall visual quality of the image. First, the contrast of the image is adjusted through histogram equalization to make the brightness distribution of the image more uniform. By modifying the gray levels of the image, the brightness contrast of the image becomes more obvious, enhancing the image details. Second, contrast enhancement is performed. By adjusting the brightness and contrast of the image, the key information of the image, especially the details such as edges and textures, is highlighted. By performing histogram equalization or contrast enhancement on the image, the visual effect of the image is improved, making subsequent edge detection and quality assessment more accurate.
[0040] Finally, format conversion is performed to convert the image into an RGB image for subsequent feature extraction processing, so that the subsequent processing module can perform efficient processing.
[0041] Step 102: Extract the image edge features from the RGB feature image to obtain the edge feature data to be detected.
[0042] Exemplarily, extracting the image edge features from the RGB feature image can extract the basic features of the image's color, texture, and edges, providing information for subsequent edge detection and quality assessment, and realizing the targeted extraction of image edge features.
[0043] Specifically, extracting the image edge features from the RGB feature image to obtain the edge feature data to be detected includes: extracting the color features of the RGB feature image to obtain the image color feature data; among them, color feature extraction includes: HSV color space conversion and color histogram calculation; through the gray-level co-occurrence matrix and Gabor filter, extracting the texture features of the RGB feature image to obtain the image texture feature data; based on the image color feature data and the image texture feature data, obtaining the edge feature data to be detected.
[0044] In one embodiment, first, color feature extraction provides support for subsequent image analysis by performing color space conversion on the image and calculating the color histogram or other color features. And, through color histogram calculation, the occurrence frequency of each color in the image is calculated to generate a color distribution map. The histogram can capture the color distribution information in the image and provide data support for subsequent image analysis.
[0045] Then, through color space conversion, the RGB image is converted into the HSV color space to extract color features from different color models. The extracted color features are used as features.
[0046] The gray-level co-occurrence matrix (GLCM) and Gabor filter are used to extract the texture features of the image. The gray-level co-occurrence matrix specifically refers to: by calculating the spatial correlation of the pixel gray values in the image, extracting texture features such as contrast, homogeneity, and entropy.
[0047] Using a Gabor filter, texture information in the image is extracted through a convolution operation, especially having a good capturing effect on edge textures and details. Then, the extracted features are combined into a texture feature quantity for subsequent input into the depth evaluation model.
[0048] Step 103: Construct a cascaded double-layer attention HED architecture based on an evolutionary algorithm, and input the edge feature data to be detected into the cascaded double-layer attention HED architecture based on the evolutionary algorithm to obtain a multi-stage edge prediction map.
[0049] Exemplarily, a cascaded HED network based on an evolutionary algorithm is used to extract edge information from the image and enhance the robustness against factors such as noise and blur. However, due to the large dimension of image features and complex calculations, existing HED networks generally have the problem of overfitting, resulting in low accuracy. To improve the accuracy and robustness of edge detection, this module combines a spatial attention mechanism and a channel attention mechanism, enabling the network to focus more on important regions and feature channels in the image, thereby improving the edge detection ability of the image. To improve the convergence speed of HED edge detection and reduce the risk of overfitting, this module combines an evolutionary algorithm to optimize the network hyperparameters, enabling the hyperparameters to achieve high-speed parameter optimization in a probabilistic and selective manner. To further improve the edge detection accuracy and be able to detect more complex or subtle structures, based on the multi-scale depth supervision of HED, this module introduces a top-down and bottom-up two-way information flow mechanism and a cascaded double-layer attention mechanism HED network architecture.
[0050] Specifically, inputting the edge feature data to be detected into the cascaded double-layer attention HED architecture based on the evolutionary algorithm to obtain a multi-stage edge prediction map includes: inputting the edge feature data to be detected into the cascaded double-layer attention HED architecture based on the evolutionary algorithm to obtain a stage edge prediction map; performing upsampling fusion on the stage edge prediction map to determine a fusion prediction map that conforms to the image to be evaluated; and obtaining a multi-stage edge prediction map based on the stage edge prediction map and the fusion prediction map. During this process, an evolutionary algorithm is used to optimize the network architecture parameters, reduce the risk of overfitting, and greatly improve the convergence speed of the network. Introducing a top-down and bottom-up two-way information flow mechanism and a cascaded double-layer attention mechanism HED network architecture to improve the edge detection accuracy and enhance the edge detection effect.
[0051] In one embodiment, the single-layer double-layer attention HED architecture is explained by Table 1 below.
[0052]
[0053] Table 1
[0054] The output layer first constructs 6 1x1 convolutional layers (conv1_dsn, conv2_dsn, conv3_dsn, conv4_dsn, conv5_dsn, fuse), each with 1 convolutional kernel, to generate edge prediction maps at different stages. Then, the prediction results of conv1_dsn to conv5_dsn are upsampled and fused to the original image size. Finally, the upsampled results of conv1_dsn to conv5_dsn are weighted and summed to obtain the final edge prediction map.
[0055] The second layer of the cascaded double-layer HED architecture is symmetric to the structure shown in Table 1, with the same number of specific stages, and the output layer is an additional output layer identical to the fifth stage in Table 1. It realizes "the upper layer guiding the lower layer" in the form of symmetric cascading to help the underlying features better focus on meaningful regions during training; and "the lower layer's details feeding back to the upper layer" to help the higher layers identify more refined edges.
[0056] Step 104: Analyze the multi-stage edge prediction map using the cross-entropy loss function to obtain the edge information map.
[0057] Exemplarily, the present application uses a pixel-level cross-entropy loss function to calculate the losses for 6 edge prediction maps (the prediction maps of 5 stages + the fused prediction map) respectively, and then weights and sums them to obtain the total loss, achieving further improvement in the edge detection performance of the HED network when processing images with high noise or blurred edges.
[0058] Specifically, analyzing the multi-stage edge prediction map using the cross-entropy loss function to obtain the edge information map includes: constructing the cross-entropy loss function based on the multi-stage edge prediction map; where the calculation formula of the cross-entropy loss function is:
[0059]
[0060]
[0061] Among them, and W are the height and width of the image to be evaluated, Y is the ground truth edge map, is the edge detection map of the th stage, is the fused prediction map; where, and W are the height and width of the image to be evaluated, Y is the ground truth edge map, is the edge detection map of the th stage, is the fused prediction map, represents the pixel value of the ground truth edge map at the position (h, w), and Denote the pixel value at position (h, w) of the prediction map and the fused prediction map in the $i$-th stage; according to the cross-entropy loss function, sum the losses weighted by the multi-stage edge prediction maps to determine the total loss function of the multi-stage edge prediction maps. The calculation formula of the total loss function of the multi-stage edge prediction maps is as follows:
[0062] (3)
[0063] Among them, is the loss weight coefficient of the stage prediction map in the -th stage, is the loss weight coefficient of the fused prediction map in the -th stage; input the multi-stage edge prediction map into the total loss function of the multi-stage edge prediction map to obtain the edge information map.
[0064] In one embodiment, assume the size of the input image is $X$, with size $H\times W$, and the ground truth edge map is $Y$.
[0065] For the multi-stage edge prediction map, construct the cross-entropy loss function. Among them, the cross-entropy loss function is explained by the following formula:
[0066] (1)
[0067] (2)
[0068] Among them, is the edge detection map in the -th stage;
[0069] is the fused prediction map;
[0070] denotes the pixel value (0 or 1) of the ground truth edge map at position (h, w);
[0071] and denote the pixel values (probability values between 0 and 1) of the prediction map and the fused prediction map in the $i$-th stage at position (h, w).
[0072] Sum the losses of the prediction maps and the fused prediction maps in 5 stages weighted to determine the total loss function of the multi-stage edge prediction maps. Among them, the total loss function of the multi-stage edge prediction maps is explained by the following formula:
[0073] (3)
[0074] Among them, is the loss weight coefficient of the stage prediction map in the -th stage, is the The loss weight coefficient of the fusion prediction map for each stage; input the multi-stage edge prediction map into the total loss function of the multi-stage edge prediction map to obtain the edge information map.
[0075] The improved HED network architecture introduces a Spatial Attention Module (SAM) and a Channel Attention Module (CAM) at each stage, which are used to enhance the network's attention to important regions and feature channels in the image. SAM highlights significant regions in the image by generating a spatial attention map; CAM emphasizes feature channels with a large amount of information by generating channel attention weights. The combination of these two attention mechanisms can effectively improve the edge detection performance of the HED network, especially when dealing with images with high noise or blurred edges.
[0076] Furthermore, this application adopts a spatial attention mechanism. Since the spatial attention mechanism helps the model selectively focus on important regions in the image spatially. It can assign more "attention" weights to key regions in the image, enabling the model to concentrate on processing these regions instead of uniformly processing the entire image, thereby reducing overfitting. The specific implementation method is to generate a two-dimensional attention map and calculate the importance of each spatial position in the image.
[0077] Furthermore, the channel attention mechanism dynamically adjusts the importance of each channel by assigning a weight to each channel. In the HED network, different channels represent different edge features, and the channel attention mechanism enables the network to pay more attention to those feature channels that are crucial for edge detection. Calculate the features of each channel through global average pooling or max pooling, and then generate the weight of each channel through a small fully connected network. Finally, multiply the channel attention weight with each channel of the input feature map to enhance the features of important channels.
[0078] It should be noted that through the double attention mechanism layer, deep important correlation vectors of the feature vectors are extracted, improving the learning efficiency of the HED network and further reducing its training complexity. Since the output of the corresponding module in this part is the edge map of the image, which is an abstract representation of the main edge information in the input image, usually an edge map represented by binary or grayscale values. The size of the output image is the same as that of the input image. The output edge map shows the edge information of the object contours and important structures in the image.
[0079] Step 105: Perform edge distortion detection on the edge information map to obtain edge detail loss evaluation data.
[0080] Exemplarily, perform distortion detection on the edge information in the image to evaluate whether important edge details are lost during processes such as image compression and transmission, realizing the monitoring of the edge information distortion state of the image and improving the detection accuracy of edge details.
[0081] Specifically, edge distortion detection is performed on the edge information map to obtain edge detail loss evaluation data, including: converting the edge information map into edge information data, and performing edge structural similarity analysis on the edge information data to obtain the edge structure similarity; wherein, the calculation formula for edge structural similarity analysis is:
[0082]
[0083] Among them, is the intersection part of the predicted edge region and the true edge region, represents the union part of the predicted edge region and the true edge region; edge intensity difference analysis is performed on the edge information map to obtain edge intensity difference data; wherein, edge intensity difference analysis includes: image gradient magnitude analysis, error analysis, and the calculation formula for image gradient magnitude analysis is:
[0084]
[0085]
[0086] Among them, is the gradient in the x direction, is the gradient in the y direction, is the image to be evaluated; edge position offset analysis is performed on the edge information map to obtain edge position offset data; wherein, edge position offset analysis is explained by the following formula:
[0087]
[0088] Among them, is the total number of pixels in the edge information map, is the th pixel in the image in the true edge map, is the th pixel in the image in the predicted edge map; according to the edge structure similarity, edge intensity difference data, and edge position offset data, edge detail loss evaluation data is obtained through comprehensive threshold scoring analysis.
[0089] In one embodiment, first, the fidelity of the edge is evaluated by calculating the structural similarity (EdgeStructural Similarity, ESSIM) between two edge maps.
[0090] ESSIM takes into account features such as the continuity, directionality, and intensity of edges. It is usually measured using edge overlap (e.g., IoU (Intersection over Union) or Dice coefficient). In this application, IoU is used for measurement because IoU measurement is intuitive and simple to calculate, making it more suitable for actual image analysis applications. By the degree of overlap with the true edges, the similarity of the edge structure can be clearly evaluated. The edge structure similarity analysis is explained by the following formula:
[0091] (4)
[0092] Where, is the intersection part of the predicted edge region and the true edge region, represents the union part of the predicted edge region and the true edge region.
[0093] Furthermore, the value of IoU is between 0 and 1. When it is close to 1, it means that the predicted edge and the true edge completely coincide, indicating that the edge structure is very similar (indicating that there is a significant edge distortion in the image, and further optimization or quality assessment may be required).
[0094] When it is close to 0, it means that the predicted edge and the true edge hardly overlap, indicating that the edge structures are quite different (indicating that the edge structure of the image is relatively intact and the distortion is small).
[0095] In this application, the threshold of IoU is set to 0.7. When IoU > 0.7, it is considered that there is no significant distortion in the edge structure; when IoU ≤ 0.7, it is considered that there is significant edge distortion in the image, and subsequent quality assessment and optimization may be required.
[0096] Then, by calculating the edge intensity difference (Edge Intensity Difference, EID) between corresponding positions of two edge maps, the degree of loss of edge details is evaluated. EID can reflect the changes in the blurriness and sharpness of edges. The gradient magnitude is calculated using the Canny edge detector, and then the mean square error (MSE) or absolute error of the two images is calculated.
[0097] The edge intensity difference analysis includes: image gradient magnitude analysis, error analysis. The image gradient magnitude analysis is explained by the following formula:
[0098] (5)
[0099] (6)
[0100] Where, is the gradient in the x direction, is the gradient in the y direction, is the image to be evaluated.
[0101] Furthermore, calculate the gradient magnitude G of the Canny edge detector (representing the edge intensity of each pixel in the image), which is explained by the following formula:
[0102] (7)
[0103] Perform edge position offset analysis on the edge information map to obtain edge position offset data; among them, the edge position offset analysis is explained by the following formula:
[0104] (8)
[0105] Among them, is the total number of pixels in the edge information map, is the th pixel in the image in the true edge map, is the th pixel in the image on the predicted edge map (the marked position of the edge); according to the edge structure similarity, edge intensity difference data, and edge position offset data, through comprehensive threshold scoring analysis, obtain the edge detail loss evaluation data.
[0106] It should be noted that the smaller the above MSE value, the smaller the displacement between the predicted edge and the true edge, and the more accurate the edge detection result
[0107] Finally, according to the edge distortion metric result, determine whether the image to be evaluated has significant edge distortion. This method uses comprehensive threshold scoring to set different thresholds for the three metrics. When one of the three edge distortion metrics exceeds the threshold, it is considered that the image to be evaluated has edge distortion and subsequent quality evaluation and optimization are required.
[0108] Therefore, the setting of the threshold is very crucial. Different threshold settings will lead to different distortion detection results, which will in turn affect the performance of the entire image quality evaluation system. To avoid the subjectivity of manually setting the threshold, this application searches for the optimal threshold corresponding to the three metrics through grid search.
[0109] For each set of threshold combinations, the deep edge detection model HED will use these thresholds to calculate the edge distortion metric and evaluate the performance of the model on the training set. For example, the HED model will calculate the edge structure similarity, edge intensity difference, and edge position offset, and combine the above metrics to judge the performance of the model.
[0110] Furthermore, the threshold setting is specifically set in the following way:
[0111] First, define the range of thresholds: According to the general setting method, the threshold of edge overlap is defined as the interval from 0.5 to 0.9, the edge sharpness is set as the interval from 0.4 to 0.8, and the edge position offset is set as the interval from 0.1 to 0.9.
[0112] Then, search by traversing the combinations of these thresholds. For example, the threshold of edge overlap can be set to 0.7, the threshold of edge sharpness can be set to 0.6, and the threshold of edge error can be set to 0.1 to combine multiple thresholds.
[0113] Each time a set of thresholds is set, evaluate the performance of the HED edge detection model built in S3 on the training set, and select the optimal threshold combination. Since HED is an edge detection model, the better its edge detection effect, the better the model's extraction effect of edge information. This is because the aforementioned HED model essentially generates an edge map, which is a continuous probability map that represents the intensity or confidence of each pixel belonging to the edge, and the value is usually between 0 and 1. For example, when a pixel value is close to 1, it indicates a high probability of belonging to the edge; close to 0 means it does not belong to the edge. Thresholds are used to convert this continuous value map into a binary map. By setting a threshold, it is determined which pixels are considered "edges" and which are not. Each metric (such as edge structure similarity, edge intensity difference, edge position shift) will have its own independent threshold. By grid search, traverse all possible threshold combinations and apply them to the edge map output by the HED model to evaluate the effects of these threshold combinations. Different thresholds will result in different edge maps (i.e., binarized edge maps), and then calculate the specific values of the metrics according to each edge map.
[0114] Calculate the comprehensive score for each threshold combination. The threshold combination with a higher comprehensive score is the optimal combination. The comprehensive score corresponding to the evaluation criterion is explained by the following formula:
[0115] (9)
[0116] Among them, (Intersection over Union)is the edge overlap;
[0117] (Edge Intensity Difference)is the edge intensity difference;
[0118] (Edge Location Shift)is the edge position shift.
[0119] The higher the value, the better the threshold evaluation score, the closer the edge detection map of the model is to the reference map, and the more accurate the setting of the characterization threshold.
[0120] Step 106: Based on the edge detail loss evaluation data, determine the edge feature vector of the distorted image through the extraction of the edge features of the distorted image.
[0121] In one embodiment, the HED model that has been trained is directly used to perform edge detection on the distorted image. The trained HED model has strong robustness by learning the edge features in the image and can effectively process various types of distorted images, such as compression distortion, noise, blur, etc. The HED model can automatically extract edge information without explicit threshold setting and is particularly suitable for detecting edges in distorted images.
[0122] It should be noted that the input of the dual-attention mechanism HED model is the distorted image. After the preprocessing stage of the image data to be evaluated, the input to the model is the trained HED edge detection model based on the dual-attention mechanism, and the output is the edge feature vector of the distorted image.
[0123] Step 107: According to the edge feature vector of the distorted image, obtain the image quality evaluation result through image quality evaluation.
[0124] According to the edge feature vector of the distorted image, obtaining the image quality evaluation result through image quality evaluation specifically includes: performing high-level feature extraction on the edge feature vector based on the preset self-attention mechanism ResNet convolutional neural network to obtain high-level edge features; using the high-level edge features as the evaluation learning background data of the ResNet convolutional neural network and training the model to convergence through the mean square error loss function to obtain the image quality evaluation model; inputting the image to be evaluated into the image quality evaluation model to obtain the image quality evaluation result.
[0125] After obtaining the image quality evaluation result according to the edge feature vector of the distorted image through image quality evaluation, the method further includes: converting the quality evaluation result into an image quality score and the image distortion situation; outputting the image quality score and the image distortion situation to a preset visualization system.
[0126] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiments of this application also provide an image quality evaluation device based on semi-reference edge detection, and its structure is as Figure 2 shown.
[0127] Figure 2 This is a schematic diagram of the internal structure of an image quality evaluation device based on semi-reference edge detection provided by the embodiments of this application. As Figure 2 shown, the device includes:
[0128] at least one processor 201;
[0129] and a memory 202 communicatively connected to the at least one processor;
[0130] wherein the memory 202 stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to:
[0131] obtain an image to be evaluated, and perform preprocessing for format normalization on the image to be evaluated to obtain an RGB feature image; extract image edge features from the RGB feature image to obtain edge feature data to be detected; construct a two-layer attention HED architecture, and input the edge feature data to be detected into the two-layer attention HED architecture to obtain a multi-stage edge prediction map; perform cross-entropy loss function analysis on the multi-stage edge prediction map to obtain an edge information map; perform edge distortion detection on the edge information map to obtain edge detail loss evaluation data; based on the edge detail loss evaluation data, determine an edge feature vector of the distorted image by extracting edge features of the distorted image; and obtain an image quality evaluation result through image quality evaluation according to the edge feature vector of the distorted image.
[0132] Some embodiments of the present application provide a non-volatile computer storage medium corresponding to Figure 1 for image quality evaluation based on semi-reference edge detection, storing computer-executable instructions, and the computer-executable instructions are configured to:
[0133] obtain an image to be evaluated, and perform preprocessing for format normalization on the image to be evaluated to obtain an RGB feature image; extract image edge features from the RGB feature image to obtain edge feature data to be detected; construct a two-layer attention HED architecture, and input the edge feature data to be detected into the two-layer attention HED architecture based on an evolutionary algorithm to obtain a multi-stage edge prediction map; perform cross-entropy loss function analysis on the multi-stage edge prediction map to obtain an edge information map; perform edge distortion detection on the edge information map to obtain edge detail loss evaluation data; based on the edge detail loss evaluation data, determine an edge feature vector of the distorted image by extracting edge features of the distorted image; and obtain an image quality evaluation result through image quality evaluation according to the edge feature vector of the distorted image.
[0134] The various embodiments in the present application are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0135] The system, medium, and method provided by the embodiments of the present application correspond one by one. Therefore, the system and the medium also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system and the medium will not be elaborated here.
[0136] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0137] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0138] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0140] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0141] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0142] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0143] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0144] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An image quality assessment method based on semi-reference edge detection, characterized in that: The method comprises: Acquire an image to be evaluated, and perform format consistency preprocessing on the image to be evaluated to obtain an RGB feature image; Extracting edge features of the RGB feature image to obtain edge feature data to be detected; Constructing a cascaded double-layer attention HED architecture based on an evolutionary algorithm, and inputting the edge feature data to be detected into the cascaded double-layer attention HED architecture based on an evolutionary algorithm to obtain a multi-stage edge prediction graph; Performing a cross entropy loss function analysis on the multi-stage edge prediction graph to obtain an edge information graph; Performing edge distortion detection on the edge information graph to obtain edge detail loss assessment data; Based on the edge detail loss evaluation data, determining an edge feature vector of the distorted image by extracting edge features of the distorted image; An image quality assessment result is obtained by performing image quality assessment according to the edge feature vector of the distorted image.
2. The image quality assessment method based on semi-reference edge detection according to claim 1, characterized in that: The image to be evaluated is preprocessed to obtain an RGB feature image, specifically including: Performing noise removal processing on the image to be evaluated to obtain denoised image data; Based on the denoised image data, a visually enhanced image is determined by denoising image enhancement; wherein the denoising image enhancement includes: histogram equalization and contrast enhancement; The visual enhancement image is converted into RGB format to obtain the RGB feature image.
3. The image quality assessment method based on semi-reference edge detection according to claim 1, characterized in that: Performing image edge feature extraction on the RGB feature image to obtain edge feature data to be detected specifically includes: Color feature extraction is performed on the RGB feature image to obtain image color feature data; wherein the color feature extraction includes: HSV color space conversion and color histogram calculation; Extracting texture features from the RGB feature image using a gray-level co-occurrence matrix and a Gabor filter to obtain image texture feature data; The edge feature data to be detected is obtained based on the image color feature data and the image texture feature data.
4. The image quality assessment method based on semi-reference edge detection according to claim 1, characterized in that: The edge feature data to be detected is input into the dual-layer attention HED architecture to obtain a multi-stage edge prediction map, specifically including: Inputting the edge feature data to be detected into the double-layer attention HED architecture to obtain a stage edge prediction map; Upsampling and fusing the edge prediction graphs of the stages to determine a fused prediction graph that conforms to the image to be evaluated; Based on the stage edge prediction map and the fusion prediction map, the multi-stage edge prediction map is obtained.
5. The image quality assessment method based on semi-reference edge detection according to claim 1, characterized in that: Performing a cross entropy loss function analysis on the multi-stage edge prediction graph to obtain an edge information graph specifically includes: Based on the multi-stage edge prediction graph, a cross entropy loss function is constructed; wherein the calculation formula of the cross entropy loss function is: in, and W are the height and width of the image to be evaluated, Y is the true edge map, For the Edge detection graph of each stage, is the fusion prediction graph, represents the pixel value of the true edge map at position (h,w), and Represents the pixel value of the prediction map and the fused prediction map at the position (h, w) of the i-th stage; According to the cross entropy loss function, the losses corresponding to the multi-stage edge prediction graph are weighted and summed to determine the total loss function of the multi-stage edge prediction graph; wherein the calculation formula of the total loss function of the multi-stage edge prediction graph is: in, For the The loss weight coefficient of the stage prediction graph of each stage, For the The loss weight coefficient of the fusion prediction graph at each stage; The multi-stage edge prediction map is input into the multi-stage edge prediction map total loss function to obtain the edge information map.
6. The image quality assessment method based on semi-reference edge detection according to claim 1, characterized in that: Performing edge distortion detection on the edge information graph to obtain edge detail loss assessment data specifically includes: The edge information graph is converted into edge information data, and edge structure similarity analysis is performed on the edge information data to obtain edge structure similarity; wherein the calculation formula of the edge structure similarity analysis is: in, To predict the intersection of the edge area and the real edge area, represents the union of the predicted edge region and the true edge region; Perform edge strength difference analysis on the edge information graph to obtain edge strength difference data; wherein the edge strength difference analysis includes: image gradient amplitude analysis and error analysis, and the calculation formula of the image gradient amplitude analysis is: in, is the gradient in the x direction, is the gradient in the y direction, is the image to be evaluated; An edge position offset analysis is performed on the edge information graph to obtain edge position offset data; wherein the calculation formula for the edge position offset analysis is: in, is the total number of pixels in the edge information map, For the image pixels in the real edge map, For the image The position of pixels on the predicted edge map; The edge detail loss evaluation data is obtained through comprehensive threshold score analysis based on the edge structure similarity, the edge strength difference data and the edge position offset data.
7. The image quality assessment method based on semi-reference edge detection according to claim 1, characterized in that: According to the edge feature vector of the distorted image, an image quality assessment result is obtained through image quality assessment, which specifically includes: Based on a ResNet convolutional neural network with a preset self-attention mechanism, high-level feature extraction is performed on the edge feature vector to obtain edge high-level features; The high-level edge features are used as evaluation learning background data of the ResNet convolutional neural network, and trained by a mean square error loss function until the model converges to obtain an image quality evaluation model; The image to be evaluated is input into an image quality evaluation model to obtain the image quality evaluation result.
8. The image quality assessment method based on semi-reference edge detection according to claim 1, characterized in that: After obtaining an image quality assessment result by performing image quality assessment according to the edge feature vector of the distorted image, the method further includes: Converting the quality assessment result into an image quality score and an image distortion condition; The image quality score and the image distortion condition are output to a preset visualization system.
9. An image quality assessment device based on semi-reference edge detection, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Acquire an image to be evaluated, and perform format consistency preprocessing on the image to be evaluated to obtain an RGB feature image; Extracting edge features of the RGB feature image to obtain edge feature data to be detected; Constructing a cascaded double-layer attention HED architecture based on an evolutionary algorithm, and inputting the edge feature data to be detected into the cascaded double-layer attention HED architecture based on an evolutionary algorithm to obtain a multi-stage edge prediction graph; Performing a cross entropy loss function analysis on the multi-stage edge prediction graph to obtain an edge information graph; Performing edge distortion detection on the edge information graph to obtain edge detail loss assessment data; Based on the edge detail loss evaluation data, determining an edge feature vector of the distorted image by extracting edge features of the distorted image; An image quality assessment result is obtained by performing image quality assessment according to the edge feature vector of the distorted image.
10. A non-volatile computer storage medium for image quality assessment based on semi-reference edge detection, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Acquire an image to be evaluated, and perform format consistency preprocessing on the image to be evaluated to obtain an RGB feature image; Extracting edge features of the RGB feature image to obtain edge feature data to be detected; Constructing a cascaded double-layer attention HED architecture based on an evolutionary algorithm, and inputting the edge feature data to be detected into the cascaded double-layer attention HED architecture based on an evolutionary algorithm to obtain a multi-stage edge prediction graph; Performing a cross entropy loss function analysis on the multi-stage edge prediction graph to obtain an edge information graph; Performing edge distortion detection on the edge information graph to obtain edge detail loss assessment data; Based on the edge detail loss evaluation data, determining an edge feature vector of the distorted image by extracting edge features of the distorted image; An image quality assessment result is obtained by performing image quality assessment according to the edge feature vector of the distorted image.
Citation Information
Patent Citations
Non-reference screen content image quality evaluation method based on multi-scale edge feature fusion
CN114897884A
Photovoltaic panel crack detection method based on dual-channel multi-scale attention mechanism
CN116402761A
Cited By
Semi-reference image quality evaluation method and system for motion blur scene
CN121415226A