An Image Inpainting Quality Evaluation Method Based on Ranking Learning and Siamese Neural Network

Through the image quality sorting method based on sorting learning and twin neural network, combined with the local face film consistency evaluation module, the problem of scarcity of data sets and evaluation consistency consistency in image repair quality evaluation is solved, and higher prediction accuracy is achieved.

CN113269680BActive Publication Date: 2025-05-27BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110174118.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-09
Publication Date
2025-05-27
Estimated Expiration
2041-02-09

AI Technical Summary

Technical Problem

The prior art has problems in the evaluation of image repair quality that data sets are scarce and it is difficult to accurately evaluate the coherence and consistency between the repaired area and the known area.

Method used

Using the image quality sorting method based on sorting learning and twin neural networks, a rich and diverse training data set was created, and a local face coherence evaluation module (PCAM) was added to the twin network to evaluate the local structural coherence and texture color consistency of the repaired area.

Benefits of technology

The prediction accuracy of the image repair quality sorting model is improved, and the quality of image repair results can be evaluated more accurately, which is 21.64% higher than that of the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113269680B_ABST
    Figure CN113269680B_ABST
Patent Text Reader

Abstract

The present invention relates to an image inpainting quality evaluation method based on ranking learning and siamese neural network, specifically including the generation of training data set, the construction of siamese network, the design of local structure coherence and color texture consistency evaluation module, the training and testing of the model, etc. The present invention is based on deep learning and has end-to-end learning ability, and can realize the quality ranking task of two pairs of images to be evaluated. The training data set created by the present invention makes up for the lack of data set in this field; the network structure proposed by the present invention can well evaluate the coherence of the inpainted area with the surrounding known areas and the internal structure of the inpainted area, as well as the consistency in terms of color texture and other contents. Through experimental tests, the network model proposed by the present invention has a higher accuracy in the ranking result of the inpainted image quality than the existing methods, thus proving the effectiveness and practicability of the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing and computer vision, and relates to an image restoration quality evaluation method based on ranking learning and Siamese neural network. Background Art

[0002] Image quality evaluation mainly analyzes and studies the characteristics of the image to be evaluated, and uses a certain algorithm to evaluate the quality of the image. According to the different images to be evaluated, image quality evaluation can be divided into fields such as fidelity quality evaluation, aesthetics quality evaluation, and image restoration quality evaluation involved in the present invention. Image restoration quality evaluation targets the result images restored by different image restoration algorithms. Currently, there are mainly three mainstream evaluation methods: methods based on similarity calculation, methods based on saliency maps, and methods based on learning.

[0003] The method based on similarity calculation calculates the similarity between the restored result image and the original image in the experimental environment. The higher the similarity, the higher the quality of the image to be evaluated. Commonly used similarity calculation metrics include SSIM and PSNR. However, this method based on similarity calculation has certain limitations. On the one hand, in real image restoration tasks, we often cannot obtain the original image, only the damaged image to be restored. On the other hand, the method of similarity calculation is only applicable to the region filling task in image restoration and not applicable to object removal, because the result image generated by region filling tends to be the same as the original image, but this is not the case for object removal.

[0004] The method based on saliency maps first calculates the saliency map of the image to be evaluated using a certain algorithm. This map represents the saliency degree of different objects in the image to be evaluated compared to other objects. If the saliency of the restored area in the saliency map is strong, then we can understand that the restored area is less natural compared to other areas, and the quality of this map is also lower. This method also has limitations. The generation of the saliency map highly depends on a certain saliency map calculation algorithm. If the saliency map calculated by this algorithm is rough, then we cannot accurately evaluate the quality of the restored result image through this method.

[0005] Learning-based methods are currently a research trend in image inpainting quality assessment and are also the methods used in this invention. The aim of this method is to train a model that can automatically evaluate the quality of the image to be evaluated, and then rank the image quality relationships of two or more images. Isogawa et al. proposed a ranking model based on learning to rank and SVM in the paper "Image quality assessment for inpainted images via learning to rank" published in 2018, and created their own dataset and feature extraction method for model training. However, the types and quantities of the datasets they created are relatively limited and are not sufficient to comprehensively simulate real inpainted result images. In addition, the feature extraction method they designed only focuses on the edges of the inpainted areas and ignores the internal features.

[0006] In summary, current image inpainting quality assessment technologies have problems such as scarce datasets and difficulty in accurately evaluating the coherence and consistency between the inpainted areas and known areas. How to simultaneously improve the above problems is a current research difficulty. Summary of the Invention

[0007] In view of the deficiencies in the above-mentioned prior art, the present invention proposes an image quality ranking method based on learning to rank and siamese neural networks, and creates a dataset that can be used for model training. This method includes the following steps:

[0008] S1: Dataset creation. First, collect the original data, manually circle the areas to be inpainted in the original images, then apply a local degradation strategy to simulate the inpainting effect on the areas to be inpainted, and obtain the inpainting effect scores and rankings by quantitatively comparing the degraded areas and the original images. The finally synthesized dataset takes image pairs as a basic unit, and the label of each pair of images is their quality ranking relationship.

[0009] S2: Siamese network construction. Since the present invention adopts a learning to rank strategy and needs to input two images simultaneously, the present invention uses a siamese network as the backbone network. This network has a total of two branches, and the network structures of each branch are the same and the parameters are shared to ensure that when two images are input into the network, they can use the same method for feature extraction.

[0010] S3: Local coherence assessment of the inpainted area. A Patch-wise Coherence Assessment Module (PCAM) is added to the main framework of the siamese network to measure the local structural coherence and texture color consistency of the inpainted area at the feature level.

[0011] S4: Network model training. We use the above-mentioned dataset and network structure to train the model, stop training when the loss function converges completely, obtain the trained model parameters, and use these parameters to extract the features of two images, and then obtain the quality ranking relationship of these two images to be evaluated.

[0012] Compared with the prior art, the present invention has the following advantages: (1) The training dataset created by the present invention has a large number and rich types, and can more comprehensively simulate real image restoration result maps. (2) The present invention uses a convolutional neural network to extract image features, and specifically measures the local structural coherence of the repaired area and the consistency in terms of texture and color. Compared with existing methods, the prediction accuracy of the image quality ranking model trained by the present invention is higher. Description of the Drawings

[0013] Figure 1 is the flowchart of the method involved in the present invention;

[0014] Figure 2 are the parameters of the network structure designed by the present invention;

[0015] Figure 3 is the local patch-based coherence evaluation module designed by the present invention;

[0016] Figure 4 is the quantitative comparison chart of the prediction results of the present invention and the results of other algorithms;

[0017] Figure 5 is the visual comparison chart of the prediction results of the present invention and the results of other algorithms; Detailed Embodiments

[0018] To make the objectives, technical solutions, and advantages of the present invention more clearly visible, the following further elaborates in detail and completely on the technical solutions of the embodiments of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] The present invention proposes an image restoration quality evaluation technology based on a siamese neural network. The flowchart is as Figure 1 shown, and specifically includes the following steps:

[0020] S1: Image restoration quality evaluation based on a learning method must rely on relevant training datasets, but currently there are no publicly available datasets. Therefore, the present invention first needs to create a dataset.

[0021] S1-1: Creation of dataset based on local area degradation. First, 300 original images are selected from some public datasets, and the area to be repaired of each image is manually marked. Then, different types of degradation operations of different degrees are performed on the area to be repaired of each image (four types of degradation, Gaussian blur, JPEG compression, salt and pepper noise, and distortion deformation, are used in this invention). The higher the degradation level, the worse the image quality. In this way, image pairs with sorting relationship labels can be created to simulate the real image repair result map.

[0022] S1-2: Creation of dataset for search-matching based image restoration algorithm. The original images used in this step are all from S1-1. The general process of search-matching based image restoration algorithm is as follows: First, a restoration area φ is delineated along the edge of the restoration area. p , and then find a point with φ from the known area p The most similar φ q , and then φ q The content covers φ p , thus completing a repair, constantly replacing φ p All the areas to be repaired can be repaired. When the present invention uses this technology to create a data set, it does not use the most similar area to fill the area to be repaired, but uses 5 areas that are not very similar. The lower the similarity, the worse the quality of the repaired result image. In this way, repair result images of different qualities can be obtained.

[0023] S1-3: Test data set creation. In order to detect the ranking effect of the model on the real image restoration result map, the present invention also creates a real test data set, and the steps are as follows: first, 200 original images different from the above steps are selected, and the area to be restored is manually circled, and then these images are restored using three algorithms with different image restoration capabilities. Each original image can obtain three restoration result maps of different quality, and thus form a test image pair with a ranking label.

[0024] S2: By analyzing the actual image restoration result graph, it can be found that relatively low-level features (such as texture, lines, and colors) are the main factors affecting the quality of the evaluated image, and high-level semantic information will also affect the quality. Therefore, the present invention designs a 5-layer convolutional neural network as the branch structure of the backbone network. The structures of the two branches of the twin network are the same and the parameters are shared. There is a global average pooling layer before the fully connected layer of the branch network. The purpose is to reduce the number of parameters and improve the efficiency of training. It is followed by 3 fully connected layers, and the final output is a single scalar representing the overall structural coherence of the image. In addition to being passed to the next layer, the output features of each convolutional layer will also be sent to PCAM for measuring local coherence. Figure 2These are the network parameters designed by the present invention. The first column is the network layer, including 5 convolutional layers, 5 activation layers, 3 pooling layers, and 3 fully connected layers. The second column shows information such as the convolutional kernel size and stride in each layer. For example, in the first layer, the convolutional kernel size is 11x11, the stride is 2, and the padding size is 5. The third and fourth columns represent the input feature size and output feature size of each layer. The size of the original input image is 224x224. After each pooling, the size of the feature is reduced to half of the original. The final output of the network is a single scalar, which represents the quality score of the image.

[0025] S3: As Figure 3 shown, the present invention designs a module that can measure the local structure coherence and texture color consistency of the image to be evaluated. The working principle of this module is as follows: First, according to the mask image, the bounding box of the repaired area is calculated. Then, according to the bounding box, the multi-channel features output by each convolutional layer are cropped. The features obtained after cropping are the features we need to measure. Then, this feature is divided into blocks (the size of the block set in the present invention is 3x3). Subsequently, the cosine similarity between each block and its surrounding 8 blocks is calculated. Finally, a similarity map representing the structure coherence and texture color consistency of the repaired area can be obtained. After adjusting the size of this similarity map, it is directly spliced with the last layer of the two network branches.

[0026] S4: The above are the preparatory work before training, including the creation of the dataset and the design of the network structure. Next, the network needs to be trained until the loss function converges completely and then stop training to obtain the model parameters suitable for the image quality ranking task. Using these parameters for calculation, the quality ranking relationship between two images can be finally obtained.

[0027] S4-1: In the training stage, the mixed dataset created in the above steps is selected in this example. Its composition includes both image pairs synthesized based on regional degradation and image pairs generated based on traditional image restoration algorithms.

[0028] S4-2: The sorting loss function is used for training: Loss rank = max(0, δ(x 1 , x 2 )(y 2 - y 1 ) + m), where x 1 and x 2 represent the input images, δ(x 1 , x 2 ) represents the sorting label of the input image pair. In the setting of this example, if the quality of x 1 is better than that of x 2 , then its value is 1, otherwise the value is -1. y1 and y 2 refer to two single scalars output by the two branches of the Siamese network, representing the quality scores of two input images. m represents the interval used to enlarge the difference between the quality scores of the two images. In this task, we set m to 0.05.

[0029] S4-3: During the training process, the means of data augmentation include resizing each image to a unified size (the size of the image in this example is 224x224) and normalization operations.

[0030] S4-4: In this embodiment, the Adam optimizer is selected during the network training process, where the optimizer parameters β 1 = 0.9 and β 2 = 0.999. The Batch size is equal to 1, the initial learning rate is 3e-5, and the weight decay coefficient is 1e-4.

[0031] Usage phase

[0032] Prepare the data training set and construct the network structure according to the aforementioned method. After the training is completed, input the pair of pictures to be sorted into the trained overall network to obtain the quality sorting relationship of these two images.

[0033] Method testing

[0034] The method disclosed in the present invention is tested in the validation set created in step S1-3 and compared with the currently disclosed methods, so as to verify that the network framework proposed by the present invention can better evaluate the structural coherence and texture color consistency of the images to be evaluated, and at the same time, the accuracy rate of the obtained image quality sorting is relatively higher.

[0035] The quantitative evaluation index is the prediction accuracy rate of the sorting result. The calculation method is to divide the number of image pairs predicted correctly by the model by the total number of image pairs to be evaluated. As Figure 4 shown, the present invention is compared with the SVM-based evaluation algorithm proposed by Isogawa et al. The results show that the prediction accuracy of the present invention can be up to 21.64% higher than their method. At the same time, the present invention also conducts ablation experiments on the performance of PCAM. The results show that adding PCAM to the network can improve the prediction accuracy by 5.79%.

[0036] Figure 5It is a visual comparison between the experimental results of the present invention and those of Isogawa et al. (a) represents the original image to be repaired, and the masked part is the object to be removed. (b) and (c) represent the result images of different image repair algorithms with different qualities. The sorting result of the method of Isogawa et al. for the two images (b) and (c) is that (b) is worse than (c), while the result of the present method is consistent with the true sorting result, that is, (b) is better than (c). This is because Isogawa et al. only considered the coherence of the edges of the repaired area when designing the feature extraction method. If the edges of a repaired image are repaired well while the interior is not, then errors will occur when using the method proposed by Isogawa et al. to predict the relationship between image qualities. The technology proposed in the present invention can not only evaluate the coherence of the edges of the repaired area, but also evaluate the consistency of the texture colors inside the repaired area, so the prediction result is relatively better.

[0037] In summary, the present invention discloses an image repair quality evaluation method based on ranking learning and siamese neural network, mainly elaborating on the construction method of the training data set, the composition of the network model and the training process in this technology. The data set altogether includes three parts: the data set generated based on regional degradation, the data set generated based on traditional image repair algorithms, and the data set generated based on the real image repair results. The entire technical framework is divided into two parts: the basic siamese network framework, and the local structure coherence and color texture consistency measurement module. By testing the existing methods, it is verified that the present invention is more targeted than the existing methods in feature extraction of the repaired result images, and has higher accuracy in the sorting results of the qualities of two images to be evaluated.

Claims

1. An image inpainting quality evaluation method based on sorting learning and Siamese neural network, characterized in that, it includes the following steps: S1: Dataset creation; First, collect the original data, manually circle the area to be inpainted in the original image, then apply the local degradation strategy to simulate the inpainting effect on the area to be inpainted, and obtain the inpainting effect score and sorting by quantitatively comparing the degraded area and the original image; The finally synthesized dataset takes image pairs as a basic unit, and the label of each pair of images is their quality sorting relationship; S2: Siamese network construction; Adopt the Siamese network as the backbone network; This network has a total of two branches, and the network structures of each branch are the same and the parameters are shared to ensure that when two images are input into the network, they can extract features in the same way; S3: Local coherence evaluation of the inpainted area; A local patch coherence evaluation module is added to the main framework of the Siamese network to measure the local structural coherence and texture color consistency of the inpainted area from the feature level; S4: Network model training; Use the above dataset and network structure to train the model, stop training when the loss function converges completely, obtain the trained model parameters, and use these parameters to extract the features of two images, and then obtain the quality sorting relationship of these two images to be evaluated; S2 is specifically as follows: Design a 5-layer convolutional neural network as the branch structure of the backbone network, and the structures of the two branches of the Siamese network are the same and the parameters are shared; The first column is the network layer, including 5 convolutional layers, 5 activation layers, 3 pooling layers and 3 fully connected layers; They are convolutional layer 1, activation layer 1, convolutional layer 2, activation layer 2, max pooling layer in turn; Convolutional layer 3, activation layer 3; Convolutional layer 4, activation layer 4, max pooling layer; Convolutional layer 5, activation layer 5, global average pooling layer and 3 fully connected layers; After each pooling, the size of the feature will be reduced to half of the original, and the final output of the network is a single scalar, which is used to represent the quality score of the image; S3 is specifically as follows: Design a module that can measure the local structural coherence and texture color consistency of the image to be evaluated. The working principle of this module is as follows: First, calculate the bounding box of the inpainted area according to the mask image, and then crop the multi-channel features output by each convolutional layer according to the bounding box. The cropped features are the features to be measured; Then cut this feature into blocks, and then calculate the cosine similarity between each block and its surrounding 8 blocks. Finally, a similarity map representing the structural coherence and texture color consistency of this inpainted area can be obtained; After adjusting the size of this similarity map, it is directly spliced with the last layer of the two network branches; S4 is specifically as follows: S4-1: In the training stage, the dataset created in the above steps is selected, and its composition includes both image pairs synthesized based on regional degradation and image pairs generated based on traditional image inpainting algorithms; S4-2: The training uses a ranking loss function: Loss rank = max(0, δ(x 1 , x 2 )(y 2 - y 1 ) + m), where x 1 and x 2 represent the input images, δ(x 1 , x 2 ) represents the ranking label of the input image pair. For example, if the quality of x 1 is better than that of x 2 , its value is 1; otherwise, the value is -1. y 1 and y 2 refer to two single scalars output by the two branches of the siamese network, representing the quality scores of the two input images. m represents the margin, which is used to widen the difference between the quality scores of the two images. m is set to 0.05; S4-3: The means of data augmentation during training include scaling each image to a unified size and normalization operation; S4-4: The Adam optimizer is selected during the network training process, where the optimizer parameters β 1 = 0.9 and β 2 = 0.999; The Batchsize is equal to 1, the initial learning rate is 3e-5, and the weight decay coefficient is 1e-4.

Citation Information

Patent Citations

  • A contrast learning image quality evaluation method based on a twin network

    CN109727246A

  • Enhanced image quality evaluation method based on twin network

    CN110033446A