An Active Learning Method for Graph-to-Graph Transformation Tasks Based on Submodular Functions

By using deep convolutional neural networks to extract the local area features of the image in the graph-graph transformation task and constructing representative and diverse submodule functions, the information redundancy problem of existing algorithms when selecting local areas of the image is solved, achieving more efficient annotation and better model accuracy.

CN114283161BActive Publication Date: 2025-07-01ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111589805.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2025-07-01
Estimated Expiration
2041-12-23

Smart Images

  • Figure CN114283161B_ABST
    Figure CN114283161B_ABST
Patent Text Reader

Abstract

The present invention discloses an active learning method for graph-graph transformation tasks based on submodular functions, including: training a deep model for graph-graph transformation tasks using a set of local regions of images with existing labels; extracting deep features of unlabeled local regions on the images using the trained model; constructing a submodular function for measuring the representativeness and diversity of the extracted set of deep features of local regions; optimizing and solving the problem of maximizing the submodular function to obtain the set of local regions of images to be labeled selected in this round of iteration; obtaining the labels of the local regions of images to be labeled, adding them to the set of local regions of images with existing labels, and then retraining the deep model for graph-graph transformation tasks using all labeled data sets; repeating the above steps until the selected local regions of images meet the requirements, and obtaining the most valuable local regions for labeling. Using the present invention, more valuable local regions of images can be selected from unlabeled images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of active learning in machine learning, and particularly relates to an active learning method for graph-to-graph transformation tasks based on submodular functions. Background Art

[0002] Whether it is traditional machine learning methods or deep learning methods, a machine learning model with good performance requires a large amount of labeled data. However, when manually annotating data, due to reasons such as the large scale of the annotated data and the high difficulty of annotating target task data, the cost required for data annotation is extremely high. To address the above problems, many researchers have tried to solve them using active learning methods. The goal of active learning methods is to automatically identify the most valuable data for annotation, so as to make the machine learning model achieve the best performance as much as possible under a certain annotation cost. Various studies have shown that active learning methods can indeed obtain a satisfactory machine learning model under the limitation of annotation cost by selecting the most valuable data for annotation.

[0003] The concept of the graph-to-graph conversion task was proposed by the article "Image-to-image translation with conditional adversarial networks" in the 2017 IEEE Conference on Computer Vision and Pattern Recognition. The graph-to-graph transformation task is a general term for a class of common computer vision tasks. The goal of this type of task is to "transform" the input image into an output image with different semantics, such as the semantic map of each region of the output image in the image semantic segmentation task, and the edge information map of the output image in the image edge detection task, etc. Since the prediction target of the graph-to-graph transformation task is also an image containing various semantics, the data annotation for such computer vision tasks requires pixel-level annotation of the image. Therefore, the annotation cost of this type of task is more substantial than that of general computer vision tasks, and active learning algorithms are more needed to reduce the annotation cost and improve the model accuracy under a certain amount of annotated data.

[0004] The existing active learning algorithms for graph-to-graph transformation tasks can basically be divided into two categories: image-level based active learning methods and region-level based active learning methods. The former regards the image as the smallest unit and selects a subset of images for model training and iteration.

[0005] At the 2017 International Conference on Medical Image Computing and Computer-Assisted Intervention, the article "Suggestive annotation: A deep active learning framework for biomedical image segmentation" proposed a deep active learning framework called suggestive annotation method. This framework is for the image segmentation task in the graph-to-graph transformation task. By training a series of models on the annotated data and selecting the sample set with the largest variance in the unannotated data as the data to be annotated, it selects the most valuable image set for annotation. However, such methods regard a single image as the smallest unit and ignore the information redundancy within a single image, thus marking many unnecessary regions of the image. Therefore, some scholars have also studied region-level active learning methods, which not only select images but also specify which regions within the images need to be annotated.

[0006] At the 2020 IEEE Conference on Computer Vision and Pattern Recognition, the article "ViewAL: Active Learning With Viewpoint Entropy for Semantic Segmentation" utilizes the viewpoint consistency in the multi-view dataset and measures the uncertainty through the inconsistency of model predictions, thereby selecting the most valuable region set for annotation. However, although this algorithm utilizes the information redundancy within the image, it does not utilize multiple active learning data selection criteria, so it does not achieve the most ideal effect. Summary of the Invention

[0007] Based on the deficiencies of the prior art, the present invention provides an active learning method for graph-to-graph transformation tasks based on submodular functions, which can select more valuable local regions of images from unannotated images and be used for subsequent annotation.

[0008] An active learning method for graph-to-graph transformation tasks based on submodular functions includes the following steps:

[0009] (1) Train a deep model for graph-to-graph transformation tasks using the set of local regions of images with existing labels;

[0010] (2) For images without annotation labels for local regions, use the trained deep model to extract the deep features of the unannotated local regions on the images;

[0011] (3) For the extracted set of local region depth features, construct a submodular function to measure its representativeness and diversity;

[0012] (4) Optimize and solve the maximization problem of the constructed submodular function to obtain the set of local regions of the image to be labeled selected in this round of iteration;

[0013] (5) Obtain the labels of the local regions of the image to be labeled, add them to the set of local regions of the image with existing labels, and then retrain the deep model for the graph-to-graph transformation task using all the labeled datasets;

[0014] (6) Repeat the above steps (2) to (5) until the selected local regions of the image meet the requirements or limitations, and finally obtain the preset number of local regions with the most labeling value as the target regions for later labeling.

[0015] The present invention extracts the local region features of the image through a deep convolutional neural network, and constructs a submodular function based on the representativeness and diversity criteria for the local regions of the image on this feature; then, by solving the maximization problem of this submodular function, a set of local regions of the image with both representativeness and diversity is obtained as the target regions to be labeled.

[0016] In step (1), the set of local regions of the image with initial existing labels is denoted as D l , and the corresponding set of local regions of the image without labels is denoted as D i ; use the set of local regions of the image with initial existing labels D l to train the deep model f for the graph-to-graph transformation task.

[0017] In step (2), for an image whose local regions have no labeled tags, divide the image into k1×k2 local regions. At this time, the convolutional layer features corresponding to the deep model f are also divided into k1×k2 local regions, so as to obtain the depth features corresponding to the local regions.

[0018] To eliminate the feature changes caused by the displacement of the local regions of the image, calculate the average value in the channel direction for the obtained depth features corresponding to the local regions of the image. For each local region of the image in D u , obtain the corresponding local region feature S i .

[0019] In step (3), denote the features of the set of all unlabeled local regions of the image as S u , when selecting a subset S u of S a , the specific process of constructing the submodular function is as follows:

[0020] (3-1) First, define the selected set S aFor S u For any element I u in it, its representativeness is defined as That is, for S u any element I u in it, the element with the highest similarity to I a can be found in S u , and the similarity between them is defined as the representativeness of S a to I u ;

[0021] (3-2) After defining the selected set S a For the representativeness of S u set, sum up the representativeness of all elements in S a For S u , then the representativeness of S a For S u set is obtained.

[0022] The obtained representativeness formula of S a for S u set is:

[0023]

[0024] In the formula, sim(I a , I u ) The similarity measure is the cosine similarity between two local features I a and I u .

[0025] In step (4), the submodular function maximization problem is:

[0026]

[0027] In the formula, select the target subset S u from the set of image local region features S a , so that the corresponding submodular function is maximized, and the number of local regions of S a does not exceed c.

[0028] After that, solve the submodular function maximization problem through the greedy acceleration algorithm. Through the greedy acceleration algorithm, the obtained solution of the submodular function maximization problem has a lower bound of 1-1 / e. The greedy acceleration algorithm can achieve almost linear time complexity by acceleration and at the same time ensure the optimization performance.

[0029] In step (5), obtain the label of the selected image local region, add it to the set D l of the selected local regions, and then update the model f.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] 1. The present invention uses the submodular function as a tool to study how to use the active learning algorithm to select local regions of images in the graph-to-graph conversion task, and can be applied to a variety of graph-to-graph conversion tasks with a convolutional neural network as the backbone network.

[0032] 2. The present invention constructs a submodular function that combines representativeness and diversity criteria, so that the selected local regions of images are more valuable regions for annotation. At the same time, through experiments on the graph-to-graph conversion task, it is proved that the local active learning algorithm proposed by the present invention can obtain a model superior to the baseline algorithm and some other similar methods on data with the same area of annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart of an active learning method for graph-to-graph conversion tasks based on submodular functions according to the present invention;

[0034] Figure 2 is the ODS comparison result with other methods on the dataset BSDS500 in the image edge detection task according to the embodiment of the present invention;

[0035] Figure 3 is the OIS comparison result with other methods on the dataset BSDS500;

[0036] Figure 4 is the AP comparison result with other methods on the dataset BSDS500;

[0037] Figure 5 is the ODS comparison result with other methods on the dataset NYUD;

[0038] Figure 6 is the OIS comparison result with other methods on the dataset NYUD;

[0039] Figure 7 is the AP comparison result with other methods on the dataset NYUD. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The present invention will be further described in detail below with reference to the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention, but do not limit it in any way.

[0041] The task of the present invention is described as follows: Assume that there are n images in the dataset corresponding to the target graph-to-graph transformation task, and each image will be divided into k1×k2 local regions. Therefore, there are a total of n×k1×k2 image local regions. The goal of the local active learning algorithm is to select C most valuable local regions for annotation from all the image local regions. The local active learning algorithm is an iterative selection algorithm process, and we will select c local regions in each iteration. We denote the set of unselected local regions as D u , and the selected local active regions as D l .

[0042] As Figure 1 shown, an active learning method for graph-to-graph transformation tasks based on submodular functions includes the following steps:

[0043] S01, at the beginning stage of the local active learning algorithm, we use the labeled data in the initial D l to train the deep model for the graph-to-graph transformation task. We denote the trained deep model as f.

[0044] S02, perform feature extraction of image local regions. When inputting the image I into the convolutional neural network model f, we regard the feature map of a specific convolutional layer as the feature of the input image. When the image is divided into k1×k2 local regions, we also divide the corresponding convolutional layer features into k1×k2 local regions, so as to obtain the deep features corresponding to the local regions. In addition, in order to eliminate the feature changes caused by the displacement of the image local regions, we calculate the average value in the channel direction of the obtained image local region features. At this time, for each image local region in D u , we obtain the corresponding local region feature S u .

[0045] S03, for the deep features of the extracted local regions, we construct a submodular function to measure their representativeness and diversity. We denote the set feature of all unlabeled image local regions as S u , and when selecting a subset S u of S a , we define the representativeness of the S a set with respect to the S u set as:

[0046]

[0047] where the similarity measure sim(I a ,I u ) is the cosine similarity between two local features I a and I u .

[0048] S04, we transform the local region selection of active learning into a submodular function maximization problem:

[0049]

[0050] This submodular function maximization optimization objective means that in this iteration process, from the set S of unlabeled images u select the target set S a such that the corresponding submodular function is maximized, and the number of local regions of S a does not exceed c. Then, we solve this submodular function maximization problem through a greedy acceleration algorithm. Through the greedy acceleration algorithm, the obtained solution of this submodular function maximization problem will have a lower bound of 1 - 1 / e. Moreover, the greedy acceleration algorithm can achieve almost linear time complexity through acceleration while ensuring the optimization performance.

[0051] S05, we add the selected image local regions and corresponding labels to the set D of selected local regions l , and then use the updated D l to retrain and update the model f.

[0052] S06, continuously iterate steps S02 to S05 until the number of selected image local regions D l reaches the limit C, obtaining the dataset and model selected by the final active learning algorithm.

[0053] To verify the effectiveness of the present invention, in this embodiment, a common edge detection task is used as a specific example of the graph - graph transformation task. In this task, the Holistically-Nested Edge Detection (HED) model is used as the edge detection model, which is also one of the most commonly used deep edge detection models. This model is an edge detection model based on a convolutional neural network, so it is suitable for the local feature extraction method of the present invention. More specifically, the present invention uses the output of the convolutional layer conv4_3 in the deep model as the feature extraction layer. After obtaining the model results, the standard non-maximum suppression algorithm is applied to the edge map to obtain a more appropriate edge detection result.

[0054] The present invention conducts experiments on the Berkeley Segmentation Dataset (BSDS500) and the NYU Depth Dataset (NYUD). The BSDS500 dataset contains 200 training samples, 100 validation samples, and 200 test samples; the NYUD dataset contains 381 training samples, 414 validation samples, and 654 test samples. Each image in the two datasets has manually annotated image edge information. In the experiment, all images are cropped and resized to the same size. For each dataset, it is divided into an unlabeled set and a test set of the same number. In each active learning iteration, images or image patches are selected from the unlabeled image pool and their labels are queried. The test set is used to evaluate the performance of the edge detection model.

[0055] The present invention samples three most common metrics for evaluating the accuracy of the image edge detection algorithm: Optimal Dataset Scale (ODS), Optimal Image Scale (OIS), and Average Precision (AP). A total of various image active learning algorithms such as Random, Suggestive Annotation, Random-Patch, and CEREALS are compared. For active learning methods that need to utilize uncertainty, a pre-trained edge detection model is required. Therefore, the present invention randomly selects 20% of the images from the training set as the initial labeled set. For all the compared methods, this initial image set is the same. Then, the Holistically-Nested Edge Detection (HED) model is trained on the initial data. For the algorithms that select the whole image and the algorithms that select local regions of the image, the present invention uses the number of selected pixels as the annotation cost. After selecting the same number of image pixels, the deep model is trained on the updated labeled dataset and the model results are evaluated again. For the active learning algorithms that select the whole image, the loss function of the Holistically-Nested Edge Detection model is calculated on all pixels in the image. For the methods that select local regions of the image, it is very likely to obtain an image with only a partially labeled region for training the deep model. At this time, only the loss function on all labeled pixels is calculated as the final loss for model training.

[0056] Figures 2 to 7 The curves representing the number of active learning iterations (x-axis) and the corresponding evaluation metrics (y-axis) under different active learning methods are shown. From Figures 2 to 7 observing the curves, it can be found that in the datasets of these two experiments, the local active learning algorithm (solid red line, named PASFo) proposed by the present invention has better performance than other active learning methods. In addition, overall, the algorithms that select local regions of the image are superior to the algorithms that select the whole image (except for the method of randomly selecting local regions).

[0057] In addition, in order to prove the effectiveness of the representativeness and diversity criteria, a comparative experiment was conducted between the present invention and the CEREALS method. The experimental results show that the active learning method proposed in the present invention, which combines the principles of representativeness and diversity, is superior to the active learning method that only adopts the uncertainty criterion. The reason is that the number of samples in local active learning has increased significantly, which is the number of images multiplied by the number of local regions divided for each image. Therefore, for the active learning method of selecting local regions for the graph-to-graph transformation task, the number of image patches in each active learning iteration is much larger than the number of images when selecting the entire image. Therefore, diversity becomes more important when selecting local regions of images. If an active learning algorithm only considers uncertainty, it will lead to the selection of many similar regions in each active learning iteration, resulting in a decline in the performance of the algorithm.

[0058] The above-described embodiments have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modification, supplement, and equivalent replacement made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An active learning method for graph-graph transformation tasks based on submodular functions, characterized in that, Including the following steps: (1) Training a deep model for a graph-to-graph transformation task using the set of local regions of images with existing labels; (2) For an image with unlabeled local regions, using the trained deep model to extract the deep features of the unlabeled local regions on the image; Specifically: For an image without labeled tags in a local area, the image is divided into k1×k2 local areas. At this time, the convolutional layer features corresponding to the depth model f are also divided into k1×k2 local areas, so as to obtain the depth features corresponding to the local areas; calculate the average value in the channel direction for the depth features corresponding to the obtained local areas of the image. For each local area of the image in D u , the corresponding local area feature S is obtained u ; (3) For the set of deep features of the extracted local regions, constructing a submodular function to measure its representativeness and diversity; the specific process of constructing the submodular function is as follows: (3-1) First, define the selected set S a For any element I u in S u the representativeness of which is defined as that is, for any element I u in S u , the similarity between I a and the element with the highest similarity to I u found in S a is defined as the representativeness of I u ; Define the selected set S after (3-2). a For S u Regarding the representativeness of the set, for S a For S u Sum up the representativeness of all elements in it, that is, obtain S a For S u The representativeness of the set; (4) Optimally solving the maximization problem of the constructed submodular function to obtain the set of local regions of the image to be labeled selected in this round of iteration; (5) Obtaining the labels of the local regions of the image to be labeled, adding them to the set of local regions of the image with existing labels, and then retraining the deep model for the graph-to-graph transformation task using all the labeled datasets; (6) Repeating the above steps (2) to (5) until the selected local regions of the image meet the requirements or limitations, and finally obtaining a preset number of local regions with the most labeling value.

2. The active learning method for graph-graph transformation tasks based on submodular functions according to claim 1, wherein In step (1), the set of local regions of the image with initial existing labels is denoted as D l , and the corresponding set of local regions of the image without labels is denoted as D u ; Using the set of local regions D of the image with initial existing labels l The deep model f for the graph-graph conversion task is trained 3. The active learning method for graph-graph transformation tasks based on submodular functions according to claim 1. In step (3-2), the obtained S a For S u The representative formula of the set is: where sim(I a , I u ) the similarity measure is the cosine similarity between two local features I a and I u .

4. The active learning method for graph-graph transformation tasks based on submodular functions according to claim 3, wherein In step (4), the submodular function maximization problem is: where the set feature S of the local regions of the image u selects a target subset S a from it, such that the corresponding submodular function is maximized, and the number of local regions in S a does not exceed c.

5. The active learning method for graph-graph transformation tasks based on submodular functions according to claim 4, wherein Solving the submodular function maximization problem through a greedy acceleration algorithm. Through the greedy acceleration algorithm, the obtained solution of the submodular function maximization problem has a lower bound of 1 - 1 / e.