An Improved Instance Segmentation Method Based on Sorting and Semantic Consistency Constraints
By introducing sorting loss and semantic consistency constraint loss in the instance segmentation algorithm, the shortcomings of existing algorithms in mask quality are solved, and higher mAP indicators and better mask quality are achieved.
Patent Information
- Application Number
- CN202110608265.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-06-01
AI Technical Summary
The existing two-stage and single-stage instance segmentation algorithms still have room for improvement in mask quality on public data sets, especially in the balance between real-time and precision.
New loss functions are introduced into the existing instance segmentation algorithm framework, including sorting loss and semantic consistency constraint loss, for semantic constraints on subregions of instances and pixel-level labels of instances that may contain instances.
By introducing these new loss functions, the mask quality of the instance segmentation algorithm on the public data set is significantly improved, the mAP indicator is improved, and its effectiveness is verified on the representative algorithms of instance segmentation Mask-RCNN and Yolact.
Smart Images

Figure CN113409327B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision image segmentation, and specifically is a novel improvement of instance segmentation. Background Art
[0002] Instance segmentation is a very important task in the field of computer vision. The several main tasks in modern computer vision include: image classification, object detection, semantic segmentation, instance segmentation, etc., and the complexity of these tasks increases gradually. Instance segmentation not only requires segmenting the mask of an object and identifying its category, but also needs to distinguish different instances of the same category. Therefore, instance segmentation can be defined as a technology that simultaneously solves the object detection problem and the semantic segmentation problem.
[0003] For a given image, an instance segmentation algorithm predicts a set of data labels {C j , B j , I j , S j}, where C j represents the category of the j-th predicted instance, B j represents the position information of the j-th predicted instance, I j represents the segmentation mask information of the j-th instance, and S j is the confidence of the category of the j-th predicted instance.
[0004] In recent years, with the development of the field of artificial intelligence, numerous methods and techniques for instance segmentation have emerged. The current mainstream instance segmentation methods can be divided into two-stage instance segmentation methods and one-stage instance segmentation methods.
[0005] Among them, two-stage instance segmentation includes two major categories: top-down and bottom-up, which are basically derived from two-stage object detection algorithms. As Figure 1 shown, the first stage is a simple foreground-background binary classification and regression process, aiming to extract sub-regions. The sub-regions are represented by a rectangular box (prediction box) with two pairs of coordinates, and the two pairs of coordinates represent the two corner points of the rectangular box. The rectangular box may contain an object instance. The second stage mainly performs precise regression on the corner point coordinates of the rectangular box for each sub-region extracted in the first stage, classifies the instance of the rectangular box and performs pixel-level annotation, and at the same time obtains the mask of the instance through a segmentation sub-network. The classic work of this type of algorithm is the Mask-RCNN algorithm, which reaches 37.1% mAP (mean average precision) on the coco dataset.
[0006] According to the existing literature, the two-stage instance segmentation algorithm achieves the highest accuracy on the public dataset. There are two drawbacks to this type of method: First, the two-stage processing scheme affects the speed of the algorithm, and the inability to ensure real-time performance has become the biggest obstacle to the practical application of the algorithm; Second, the quality of the masks obtained by such algorithms is still uneven. Taking Mask-RCNN as an example, its final mask is upsampled from a small area of 28×28, resulting in poor quality of the restored instance mask, manifested as insufficient or excessive coverage of the instance area by the instance mask.
[0007] Single-stage instance segmentation algorithms such as Figure 2 shown, these algorithms basically come from single-stage object detection algorithms. Compared with two-stage algorithms, there is no sub-region extraction process. The classic of this type of algorithm is the Yolact series of algorithms, which achieved a mAP of 29.8% on the coco dataset and a speed of 33.5 FPS.
[0008] The real-time performance of single-stage algorithms is sufficient to meet the requirements for practical application. The problems are that the accuracy is more degraded compared to two-stage algorithms, and the problem of low mask quality is more serious.
[0009] In summary, the best mAP metrics of the existing two-stage and single-stage instance segmentation algorithms on the public dataset indicate that there is room for further improvement in the mask quality of instance segmentation.
[0010] The idea of this invention to improve the mask quality is to add a new loss function to the above framework, define semantic constraints on the sub-regions that may contain instances and the pixel-level labels of instances. The specific forms are ranking loss and semantic consistency constraint loss. The loss function defined in this invention can be directly applied to the existing two-stage and single-stage instance segmentation algorithms to improve the mAP metrics of the original algorithms. This invention has conducted verification experiments on the representative two-stage and single-stage instance segmentation algorithms Mask-RCNN and Yolact respectively, and the experimental results have shown the effectiveness of the loss function proposed in this invention. Summary of the Invention
[0011] Aiming at the problem of poor mask quality obtained by the existing instance segmentation algorithms, this invention proposes to add a new loss function (ranking loss and semantic consistency constraint loss) to the existing algorithm framework, thereby forming a new instance segmentation method to improve the mask quality obtained by instance segmentation.
[0012] This invention proposes an improved instance segmentation method based on ranking and semantic consistency constraints. The implementation schemes on the single-stage and two-stage instance segmentation frameworks are introduced separately below.
[0013] 1. Improved Instance Segmentation Scheme Based on Two-Stage Network
[0014] AsFigure 3 As shown in Figure 3 , in the two-stage instance segmentation algorithm framework, the first stage mainly completes the extraction of regions of interest. Here, for the sorting and selection operation of sub-regions, on the basis of the original classification and regression losses, a sorting loss is added. For the segmentation operation in the second stage, on the basis of the original classification, regression, and segmentation losses, a semantic consistency loss is added.
[0015] 2. Improved Scheme for Instance Segmentation Based on Single-Stage Network
[0016] As Figure 4 shown in Figure 4 , in the single-stage algorithm, on the basis of the original classification, regression, and segmentation losses, the sorting loss and the semantic consistency loss are respectively added to the regression head and the segmentation head.
[0017] 1. Introduction to the Basic Model
[0018] The present invention adds a sorting loss and a semantic consistency loss to the original instance segmentation framework.
[0019] Next, the original basic modules of the representative works Mask-RCNN and Yolact of the two types of algorithms are introduced respectively.
[0020] Mask-RCNN Basic Module:
[0021] Feature extraction network Backbone: ResneXt101 + FPN is used as the Backbone, and the pre-trained weights are the resnet101 file, which has been pre-trained on the ImageNet dataset.
[0022] RPN network: Used to generate region proposals. This layer judges whether the anchors belong to positive or negative through softmax, and then uses bounding box regression to correct the anchors to obtain accurate ROIs (i.e., sub-regions).
[0023] Fully connected layer FC: Classification operation to obtain instance labels, and regression operation to obtain instance boxes.
[0024] Fully convolutional network FCN: Mask segmentation operation to obtain instance masks.
[0025] Each module is as Figure 5 shown:
[0026] Yolact Basic Module:
[0027] Feature extraction network Backbone: Resnet101 + FPN is used as the Backbone, and the pre-trained weight file is still the resnet101 file.
[0028] Protonet Network: Generate k prototype masks for each image. It is connected to the output of FPN and is a fully convolutional network that predicts a set of prototype masks.
[0029] Prediction Head: Performs regression classification operations, equivalent to the sub-region extraction process in two-stage methods, and outputs sub-regions and their classifications; simultaneously predicts k linear combination coefficients for linearly combining prototype masks, and performs a Crop operation on the linear combination result to obtain instance masks.
[0030] Subsection Rank Loss (SRLoss)
[0031] The core idea of this constraint is: Sort each predicted bounding box in non-increasing order according to the score, and encourage positive sample predicted bounding boxes to be as far in front of negative sample predicted bounding boxes as possible, so that more accurate sub-regions can be obtained.
[0032] The rank loss is defined as Equation (1):
[0033]
[0034] Where P is the set of positive samples, and the positive and negative samples are determined by calculating the intersection over union (IoU) between the prior box or anchor box and the GT Bbox and according to thresholds (such as 0.7 and 0.3, samples with IoU higher than 0.7 are classified as positive samples, samples with IoU lower than 0.3 are classified as negative samples, and the remaining samples are not trained or processed); the positive and negative samples are sorted from large to small according to the IoU, and the sorting serial number is the r value; the samples are sorted from large to small according to the classification confidence score, and the sorting serial number is the value of sort(r). The more forward the sorting result of positive samples, the smaller the SRLoss.
[0035] Subsection Semantic Consistency Loss (SSCLoss)
[0036] The purpose of this loss is to constrain the semantic consistency of pixels within the mask region: Constrain the number of labeled categories of pixels within the sub-region to be as small as possible; Constrain the pixels within the sub-region to belong to the same category as much as possible. The former counts the number of segmentation categories, and this part drops to the lowest when the mask quality is the best; the latter calculates the ratio of the number of pixels with correct semantic categories within the mask region to the total number of pixels within the mask region.
[0037] Let M i be the number of pixels labeled as the i-th category within the mask region, and the final form of the loss function is:
[0038]
[0039]
[0040]
[0041] Among them, c is the total number of categories, and c = 80 on the MS COCO dataset, and α is a hyperparameter.
[0042] Model Training and Testing
[0043] During the training process, the actual process of Mask-RCNN is as Figure 7 shown. The names of functional blocks are within the rectangular boxes, and the polyline arrows point to the network losses. The newly added losses are all shown in bold. The pre-trained model used in the present invention has been trained on the imagenet dataset.
[0044] Use the standard label file annotation including [id, image_id, category_id, segmentation, area, bbox, iscrowd], category label category_id, mask label segmentation, and instance box label bbox. category_id is the category label, where the segmentation is polygon format data (coordinate values of adjacent pairs of data as the contour edge points of the instance), and the bbox data format is [x 1 , y 1 , x 2 , y 2 .
[0045] In the first stage, the loss function is mainly obtained from the training of the RPN network, including the RPN foreground and background classification loss l rpn_cls , the true label t ∈ {1, 0, -1}, and the anchors with the true label of 0 do not participate in the construction of the loss function. The ones with the label of -1 are converted to 0 for cross-entropy calculation; the RPN object box regression loss l rpn_reg , and the sorting loss SRLoss proposed by the present invention, denoted as l rpn_rank in this example.
[0046] In the second stage, the loss function is mainly the classification loss l cls , the regression loss l reg , the segmentation loss l seg , and the semantic consistency loss SSLoss proposed by the present invention, denoted as l sc in this example.
[0047] Among them, the classification loss and the segmentation loss are usually cross-entropy losses, while the regression loss is often the Smooth_L1 loss. Their general forms are:
[0048] Cross-entropy loss:
[0049] y is the true value, and y i is the predicted probability value, and n is the number of samples.
[0050] Smooth_L1 loss:
[0051] y * is the predicted value, y is the label, and x is the difference between the two.
[0052] The overall training is divided into two major parts. First, train the RPN network part, and the loss function L to be optimized is 1 as follows:
[0053] L 1 = l rpn_cls + l rpn_reg + l rpn_rank (5)
[0054] After the ROI region is screened and enters the second stage, the loss function L to be optimized is 2 as follows:
[0055] L 2 = l cls + l reg + l seg + l sc (6)
[0056] In the test part, use the Mask prediction branch to process the top 100 detection boxes with the highest scores. To improve the inference efficiency, the network predicts K mask images for each ROI, but only the mask image with the highest class probability needs to be used here. Resize this mask image back to the ROI size and binarize it with a set threshold of 0.5. Those above the threshold are retained, and finally, an image-level add operation is performed on the segmented mask and the original image to obtain the final instance segmentation visualization result.
[0057] The network structure of the single-stage instance segmentation representative algorithm Yolact after adding losses is as Figure 8 shown. The names of the functional blocks are within the rectangular frames, and the broken-line arrows point to the network losses. The newly added losses are all in bold. The pre-trained model used was pre-trained on the ImageNet dataset. The network training and testing principles are basically similar to those in the previous text. The difference is that all losses are directly optimized without stages. The total losses of the network are listed below:
[0058] L = l cls + l reg + l seg + l segm + lrank +l sc
[0059] Among them, l cls and l reg are the classification loss and the regression loss, and l seg is the segmentation loss, and l segm is the semantic segmentation loss. The sorting loss is denoted as l rank , and the semantic consistency loss is denoted as l sc Both are the cross-entropy loss L ce and the Smooth_L1 loss L smooth_l1 Excluding the regression and classification losses, the rough semantic segmentation loss l segm (is still a loss in the form of cross-entropy). Description of the Drawings
[0060] Figure 1 Two-stage instance segmentation method
[0061] Figure 2 Single-stage instance segmentation method
[0062] Figure 3 Two-stage network structure diagram with added sorting and semantic consistency losses
[0063] Figure 4 Single-stage network structure diagram with added sorting and semantic consistency losses
[0064] Figure 5 Schematic diagram of Mask-RCNN
[0065] Figure 6 Schematic diagram of Yolact
[0066] Figure 7 Improved structure diagram of Mask-RCNN
[0067] Figure 8 Improved structure diagram of Yolact Detailed Implementation Manner
[0068] The present invention conducts experiments using the MS COCO series datasets (COCO2015, COCO2016). The COCO dataset is a large and rich object detection, segmentation, and caption dataset. This dataset aims at scene understanding and mainly captures from complex daily scenes. The positions of the objects in the images are calibrated through precise segmentation. The images include 91 types of objects, 328,000 images, and 2,500,000 labels. So far, it is the largest dataset for semantic segmentation, providing 80 categories, with more than 330,000 images, among which 200,000 are annotated, and the number of individuals in the entire dataset exceeds 1.5 million. The COCO dataset now has 3 types of annotations: object instances, object keypoints, and image captions, and stores the labels using JSON files. The training set and the validation set are used for training and testing respectively. Since the task completed is instance segmentation, the object instance type of annotation is used.
[0069] Evaluation metrics. The present invention follows the evaluation standard protocol for instance segmentation and uses mAP to evaluate at different Intersection over Union (IOU) thresholds. The present invention conducts experiments using the official provided evaluation code.
[0070] Experimental settings. In the experiment, the effectiveness of the present invention is tested by sequentially adding the losses defined by the present invention to each framework.
[0071] The experiments for Mask-RCNN and Yolact are carried out on the mmdetection 2.3 version launched by SenseTime.
[0072] The hyperparameters are set as follows in the Mask-RCNN experiment: α = 1; set diverse input sizes; the Non-Maximum Suppression (NMS) threshold is 0.7; the initial learning rate is 0.333, and it decays dynamically during the training process; the batchsize is set to 4, and the training is carried out on four GPUs respectively; the weights are stored once per epoch.
[0073] The hyperparameters are set on Yolact as: α = 1; set the initial size of the input image to 550×550; the Non-Maximum Suppression (NMS) threshold is 0.7; the initial learning rate is 0.333, and it decays dynamically during the training process; the batchsize is set to 4, and the training is carried out on four GPUs respectively, and the weights are stored once per epoch.
[0074] In all of the above, validation is performed once every 3 rounds of training and then the training continues.
[0075] The performance score of the model in the present invention is compared with the baseline performance published by the original author. Table 1 shows the experimental results on the COCO dataset. It can be seen that the experimental results have all been improved.
[0076] Table 1 Experimental Results on the COCO Dataset
[0077]
Claims
1. An improved instance segmentation method based on sorting and semantic consistency constraints, characterized in that it includes the following steps: Model training and testing The pre-trained model used has been trained on the ImageNet dataset; the algorithms used in the pre-trained model include the original basic modules of Mask-RCNN and Yolact; Mask-RCNN basic module: Feature extraction network Backbone: ResneXt101+FPN is the Backbone, and the pre-trained weights used are the resnet101 file, which has been pre-trained on the ImageNet dataset; RPN network: used to generate region proposals; judge whether the anchors belong to positive or negative through softmax, and then use bounding box regression to correct the anchors to obtain accurate ROIs, that is, sub-regions; Fully connected layer FC: classification operation to obtain instance labels, and regression operation to obtain instance boxes; Fully convolutional network FCN: mask segmentation operation to obtain instance masks; Yolact basic module: Feature extraction network Backbone: Resnet101+FPN is the Backbone, and the pre-trained weight file is still the resnet101 file; Protonet network: generate k prototype Masks for each image; k = 32; connected to the output of FPN, it is a fully convolutional structure network, and a set of prototype Masks are predicted; Prediction Head: regression classification operation, equivalent to the sub-region extraction process in the two-stage method, outputting sub-regions and their classifications; at the same time, predicting k linear combination coefficients for linearly combining the prototype Masks, and performing Crop operation on the linear combination result to obtain instance masks; Using the standard label file annotation including [id, image_id, category_id, segmentation, area, bbox, iscrowd], category label category_id, mask label segmentation, instance box label bbox; among which the segmentation is polygon format data, with adjacent pairs of data as the coordinate values of the edge points of the instance contour, and the bbox data format is [x 1 , y 1 , x 2 , y 2 ; In the first stage, the loss function is obtained from the training of the RPN network, including the RPN foreground and background classification loss l rpn_cls , the true label t ∈ {1, 0, -1}, the anchors with the true label of 0 do not participate in the construction of the loss function, and the ones with the label of -1 are converted to 0 for cross-entropy calculation; the RPN bounding box regression loss l rpn_reg , and the sorting loss SRLoss, denoted as l rpn_rank ; Sub-region sorting loss The sorting loss is defined as in Equation (1): where P is the set of positive samples, and the positive and negative samples are determined according to the threshold after calculating the intersection over union of the prior box or anchor box anchor and the GT Bbox; samples with an intersection over union higher than 0.7 are classified as positive samples, samples with an intersection over union lower than 0.3 are classified as negative samples, and the remaining samples are not trained and processed; the positive and negative samples are sorted from large to small according to the intersection over union, and the sorting serial number is the r value; the samples are sorted from large to small according to the classification confidence score, and the sorting serial number is the value of sort(r); the more forward the sorting result of the positive samples, the smaller the SRLoss; In the second stage, the loss function includes the classification loss l cls , the regression loss l reg , the segmentation loss l seg , and the semantic consistency loss SSLoss, denoted as l sc ; Sub-region semantic consistency loss Denote M i as the number of pixel points labeled as the i-th category within the mask region, and the form of the final loss function is as follows: where c is the total number of categories, c = 80 on the MS COCO dataset, and α is a hyperparameter; Among them, the classification loss and segmentation loss are cross-entropy losses, while the regression loss is the Smooth_L1 loss; their forms are: Cross-entropy loss: y is the true value, y i is the predicted probability value, and n is the number of samples; Smooth_L1 Loss: y * is the predicted value, y3 is the label, and x is the difference between the two; The overall training is divided into two major parts. First, the RPN network part is trained, and the loss function L to be optimized is 1 as follows: L 1 =l rpn_cls +l rpn_reg +l rpn_rank (5) After the ROI region screening enters the second stage, the loss function L to be optimized is as follows: 2 As follows: L 2 =l cls +l reg +l seg +l sc (6) In the test section, the Mask prediction branch is used to process the top 100 detection boxes with the highest scores; to accelerate the inference efficiency, the network predicts K mask images for each ROI, but only the mask image with the highest class probability needs to be used here. Resize this mask image back to the ROI size and binarize it with a set threshold of 0.
5. Those higher than the threshold are retained. Finally, an image-level add operation is performed on the segmented mask and the original image to obtain the final visualization result of instance segmentation; List the total loss of the network: L=l cls +l reg +l seg +l segm +l rank +l sc Among them, l cls and l reg are the classification loss and the regression loss, l seg is the segmentation loss, l segm is the semantic segmentation loss, and the sorting loss is denoted as l rank .
Citation Information
Patent Citations
Digital elevation model production method based on deep learning
CN110084817A
Semantic information extraction method based on Mask-RCNN
CN111862119A