Interactive semantic instance joint segmentation method based on dual-branch iterative correction network

Through the interactive semantic instance joint segmentation method of the dual-branch iterative correction network, the problem of insufficient interactive intent perception in iterative training is solved, higher segmentation accuracy and robustness are achieved, and the image segmentation results of multi-task learning are output.

CN116778163BActive Publication Date: 2025-09-16NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310784455.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-09-16
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

Existing iterative training methods cannot accurately perceive the interaction intentions of different interaction point types, and lack temporal consistency during the iterative process, which affects segmentation accuracy and robustness.

Method used

An interactive semantic instance joint segmentation method based on a dual-branch iterative correction network is adopted. By generating instance selection points and semantic selection points, the parallel symmetric ASPP module and upsampling module are combined for feature extraction, and the interactive point feature map is updated in the iterative process until the prediction accuracy reaches 85%.

Benefits of technology

It improves the accuracy and robustness of segmentation, can better perceive the user's segmentation intention, output two image segmentation results, and improves the training quality and segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778163B_ABST
    Figure CN116778163B_ABST
Patent Text Reader

Abstract

This invention discloses an interactive semantic instance joint segmentation method based on a dual-branch iterative correction network. The method comprises the following steps: in a data preprocessing phase, instance interaction points and corresponding semantic interaction points are generated according to the user's segmentation intent; when a point is generated, the interaction point is classified into active and passive types according to the point generation method and added to another branch; three RGB channels of the image, two channels of semantic points, and two channels of instance points are linked as input to the segmentation network; the input information is passed through a feature extraction backbone and dual segmentation branches to obtain semantic prediction results and instance prediction results; error regions are generated based on the prediction results, and correction points are obtained, which are updated to the interaction point feature map and re-input into the network; iteration is stopped when the prediction accuracy reaches 85%, and the final result is obtained. The invention can output two image segmentation results and improve the training quality and segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of interactive image segmentation, and in particular to an interactive semantic instance joint segmentation method based on a dual-branch iterative correction network. Background Art

[0002] Interactive image segmentation refers to separating the target of interest from the complex image background environment based on certain similarity criteria under the prior knowledge provided by the user. It is a key issue in fields such as image analysis, pattern recognition, and computer vision. The quality of segmentation will directly affect subsequent related applications.

[0003] In recent years, as deep learning models have achieved excellent results in numerous computer vision tasks, interactive segmentation methods based on deep learning have attracted increasing attention from scholars both domestically and internationally. These methods break through the limitations of traditional interactive methods, which can only extract low-level local features, and leverage the feature extraction capabilities of convolutional neural networks to achieve outstanding segmentation results, becoming a mainstream interactive image segmentation method. The iterative interactive segmentation method, based on deep learning methods, proposes a new training process that simulates a true interactive form. During training, click points are iteratively generated and the input data is dynamically expanded, resulting in even better segmentation results.

[0004] However, while existing iterative training methods simulate authentic interactions, they fail to perceive the varying intent of different interaction point types, making it difficult to achieve satisfactory segmentation results with a limited number of interaction points. Furthermore, the complete independence of all iterations during iterative training renders the model unable to perceive temporal order and struggles to maintain spatiotemporal consistency across iterations. These shortcomings severely impact the accuracy and robustness of these methods, limiting their practical application. Summary of the Invention

[0005] The purpose of the present invention is to provide an interactive semantic instance joint segmentation method based on a dual-branch iterative correction network with high accuracy, good robustness and strong practicality.

[0006] The technical solution to achieve the purpose of the present invention is: an interactive semantic instance joint segmentation method based on a two-branch iterative correction network, comprising the following steps:

[0007] Step 1: In the data preprocessing stage, instance selection points and corresponding semantic selection points are generated according to the user's segmentation intention. These points are converted into disks and used as the initial elements in the interaction point feature map.

[0008] Step 2: When an instance active selection point is generated, a corresponding semantic passive selection point is added to the interaction point feature map of the semantic branch. When a semantic active selection point is generated, a corresponding instance passive selection point is added to the interaction point feature map of the instance branch.

[0009] Step 3: The segmentation network uses a parameter-sharing residual network as the encoder, and the parallel symmetrical ASPP module, upsampling module, and segmentation module are connected in sequence as the decoder; the three RGB channels of the image, the two channels of semantic points, and the two channels of instance points are concatenated as the input of the segmentation network;

[0010] Step 4: The input information passes through the feature extraction trunk and the parallel symmetrical semantic branch and instance branch to obtain the semantic prediction results and instance prediction results;

[0011] Step 5: Generate the error area based on the prediction results, and then obtain the correction point, update it to the interaction point feature map of step 2, and re-input it into the network;

[0012] Step 6: When the prediction accuracy reaches 85%, stop the iteration and get the final result.

[0013] Furthermore, in step 1, instance selection points and corresponding semantic selection points are generated, specifically:

[0014] During data preprocessing, all instances in the graphics to be segmented are traversed, and a selected instance is obtained each time. The selected instance is used as the reference instance, and the geometric center of the reference instance is used as the instance selection point.

[0015] Set the category of the benchmark instance individual as the benchmark category, and use the geometric center of the benchmark category in the semantic segmentation ground_truth as the semantic selection point;

[0016] Interaction points are divided into foreground interaction points and background interaction points, which correspond to the two types of error areas, false negative and false positive, respectively.

[0017] Furthermore, step 2 implements the sharing of same-coordinate information of semantic instances in a dual-branch network. It is stipulated that when an interaction point is added to a semantic branch or an instance branch, the corresponding point of the location point in the foreground and background type of the other branch is calculated and added to the other branch. The interaction point is defined as an active point and the corresponding point is a passive point. Specifically:

[0018] Let instance_mask be the mask of the reference instance, and class_mask be the mask of the reference class;

[0019] (1) Each time an active point is added to the instance foreground channel, a passive point is added to the same position in the semantic foreground channel;

[0020] (2) Each time an active point is added to the instance background channel, it is first determined whether the point with the same coordinate in the class_mask belongs to the foreground or the background. If it belongs to the foreground, a semantic foreground passive point is added; if it belongs to the background, a semantic background passive point is added;

[0021] (3) Each time an active point is added to the semantic background channel, a passive point is added to the same position in the instance background channel;

[0022] (4) Each time an active point is added to the semantic foreground channel, it is first determined whether the point with the same coordinate in instance_msk belongs to the foreground or the background. If it belongs to the foreground, an instance foreground passive point is added; if it belongs to the background, an instance background passive point is added.

[0023] Furthermore, in step 3, the three RGB channels of the image, the two channels of the semantic points, and the two channels of the instance points are linked as the input of the segmentation network, specifically:

[0024] The two channels of semantic points are the semantic foreground point channel and the semantic background point channel, and the two channels of instance points are the instance foreground point channel and the instance background point channel. The image resolution is set to 320*480*3 pixels, and each point channel is set to a feature information map of 320*480*2 pixels.

[0025] Furthermore, in step 4, the input information passes through the feature extraction trunk and the parallel symmetrical semantic branch and instance branch to obtain the semantic prediction results and instance prediction results, specifically:

[0026] The feature extraction backbone is ResNet, which includes a convolutional layer I, a residual module I, and a residual module II. The extracted features are input into two segmentation branches respectively; the segmentation branch specifically includes a dilated spatial convolution pooling pyramid I, an upsampling module I, an upsampling module II, an upsampling module III, and a classification head;

[0027] The feature splicing layer I of the connection channel module is connected to the convolution layer I and the convolution layer of the classification head; the convolution layer I is connected to the Maxpool layer of the residual module I and the convolution layer of the upsampling module III; the convolution layer of the residual module I is connected to the convolution layer of the residual module II and the convolution layer of the upsampling module II; the convolution layer of the residual module I is connected to the convolution layer of the void spatial convolution pooling pyramid I of each branch; the feature splicing layer II of the void spatial convolution pooling pyramid I is connected to the convolution layer of the upsampling module I; the feature splicing layer III of the upsampling module I is connected to the convolution layer of the upsampling module II; the feature splicing layer IV of the upsampling module II is connected to the convolution layer of the upsampling module III; the feature splicing layer V of the upsampling module III is connected to the convolution layer of the classification head.

[0028] Furthermore, in step 5, the error area is generated according to the prediction results, and then the correction point is obtained, which is updated to the interaction point feature map of step 2 and re-input into the network. Specifically:

[0029] The current segmentation result obtained by the model is compared with the true label of the target area. Each time, the largest area of ​​incorrect segmentation is selected, and the centroid of the largest area is located as the segmentation coordinate of the newly generated interaction point. This process is repeated in a cycle, which is called iterative training. At the same time, the center of the largest error area is used as the active point to generate the corresponding passive point for the other branch.

[0030] Furthermore, in step 6, the iteration is stopped when the prediction accuracy reaches 85%, and the final result is obtained, which is specifically:

[0031]

[0032] Among them, IoU is the degree of overlap between the predicted instance mask and the true instance mask, that is, the prediction accuracy;

[0033] I(X) is the pixel count of the intersection of the predicted instance mask and the true instance mask;

[0034] U(X) is the pixel count of the union of the predicted instance mask and the true instance mask;

[0035] I(X) and U(X) are approximated as follows:

[0036]

[0037]

[0038] Where I(X) is the pixel count of the intersection of the predicted instance mask and the true instance mask;

[0039] U(X) is the pixel count of the union of the predicted instance mask and the true instance mask;

[0040] v is a pixel in the image to be segmented;

[0041] V is the set of all pixels in the image to be segmented;

[0042] X is the pixel probability on the set V output by the network;

[0043] Y∈{0, 1} V is the ground truth assignment of the set V, where 0 represents background pixels and 1 represents object pixels.

[0044] Compared with the existing technology, the present invention has the following significant advantages: (1) interactive semantic instance joint segmentation takes into account the fact that traditional interactive image segmentation cannot accurately perceive the user's segmentation intention and can output two image segmentation results; (2) semantic segmentation and instance segmentation tasks are combined for multi-task learning. Compared with traditional interactive image segmentation technology, the present invention improves the training quality and segmentation accuracy, and has good robustness and strong practicality.

[0045] The present invention is further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flow chart of the interactive semantic instance joint segmentation method based on the dual-branch iterative correction network of the present invention.

[0047] Figure 2 It is a schematic diagram of iterative correction of the interactive semantic instance joint segmentation method based on the dual-branch iterative correction network of the present invention.

[0048] Figure 3 This is a network structure diagram of the dual-branch network applied in the deep model in the present invention.

[0049] Figure 4 This is the segmentation result diagram of the model of the present invention. DETAILED DESCRIPTION

[0050] This paper proposes an interactive semantic-instance joint segmentation method based on a dual-branch iterative correction network. This method improves on existing iterative training methods based on deep learning by introducing a selection-correction network during training, splitting the training process into two stages: selection and correction. It then uses two branches to simultaneously perform semantic-instance joint segmentation. This method takes into account the inability of traditional interactive image segmentation to accurately perceive user segmentation intent, outputs two image segmentation results, and combines semantic and instance segmentation tasks for multi-task learning. Compared with traditional interactive image segmentation techniques, this method improves training quality and segmentation accuracy.

[0051] The present invention provides an interactive semantic instance joint segmentation method based on a dual-branch iterative correction network, comprising the following steps:

[0052] Step 1: In the data preprocessing stage, instance selection points and corresponding semantic selection points are generated according to the user's segmentation intention. These points are converted into disks and used as the initial elements in the interaction point feature map.

[0053] Step 2: When an instance active selection point is generated, a corresponding semantic passive selection point is added to the interaction point feature map of the semantic branch. When a semantic active selection point is generated, a corresponding instance passive selection point is added to the interaction point feature map of the instance branch.

[0054] Step 3: The segmentation network uses a parameter-sharing residual network as the encoder, and the parallel symmetrical ASPP module, upsampling module, and segmentation module are connected in sequence as the decoder; the three RGB channels of the image, the two channels of semantic points, and the two channels of instance points are concatenated as the input of the segmentation network;

[0055] Step 4: The input information passes through the feature extraction trunk and the parallel symmetrical semantic branch and instance branch to obtain the semantic prediction results and instance prediction results;

[0056] Step 5: Generate the error area based on the prediction results, and then obtain the correction point, update it to the interaction point feature map of step 2, and re-input it into the network;

[0057] Step 6: When the prediction accuracy reaches 85%, stop the iteration and get the final result.

[0058] Furthermore, in step 1, instance selection points and corresponding semantic selection points are generated, specifically:

[0059] During data preprocessing, all instances in the graphics to be segmented are traversed, and a selected instance is obtained each time. The selected instance is used as the reference instance, and the geometric center of the reference instance is used as the instance selection point.

[0060] Set the category of the benchmark instance individual as the benchmark category, and use the geometric center of the benchmark category in the semantic segmentation ground_truth as the semantic selection point;

[0061] Interaction points are divided into foreground interaction points and background interaction points, which correspond to the two types of error areas, false negative and false positive, respectively.

[0062] Furthermore, step 2 implements the sharing of same-coordinate information of semantic instances in a dual-branch network. It is stipulated that when an interaction point is added to a semantic branch or an instance branch, the corresponding point of the location point in the foreground and background type of the other branch is calculated and added to the other branch. The interaction point is defined as an active point and the corresponding point is a passive point. Specifically:

[0063] Let instance_mask be the mask of the reference instance, and class_mask be the mask of the reference class;

[0064] (1) Each time an active point is added to the instance foreground channel, a passive point is added to the same position in the semantic foreground channel;

[0065] (2) Each time an active point is added to the instance background channel, it is first determined whether the point with the same coordinate in the class_mask belongs to the foreground or the background. If it belongs to the foreground, a semantic foreground passive point is added; if it belongs to the background, a semantic background passive point is added;

[0066] (3) Each time an active point is added to the semantic background channel, a passive point is added to the same position in the instance background channel;

[0067] (4) Each time an active point is added to the semantic foreground channel, it is first determined whether the point with the same coordinate in instance_msk belongs to the foreground or the background. If it belongs to the foreground, an instance foreground passive point is added; if it belongs to the background, an instance background passive point is added.

[0068] Furthermore, in step 3, the three RGB channels of the image, the two channels of the semantic points, and the two channels of the instance points are linked as the input of the segmentation network, specifically:

[0069] The two channels of semantic points are the semantic foreground point channel and the semantic background point channel, and the two channels of instance points are the instance foreground point channel and the instance background point channel. The image resolution is set to 320*480*3 pixels, and each point channel is set to a feature information map of 320*480*2 pixels.

[0070] Furthermore, in step 4, the input information passes through the feature extraction trunk and the parallel symmetrical semantic branch and instance branch to obtain the semantic prediction results and instance prediction results, specifically:

[0071] The feature extraction backbone is ResNet, which includes a convolutional layer I, a residual module I, and a residual module II. The extracted features are input into two segmentation branches respectively; the segmentation branch specifically includes a dilated spatial convolution pooling pyramid I, an upsampling module I, an upsampling module II, an upsampling module III, and a classification head;

[0072] The feature splicing layer I of the connection channel module is connected to the convolution layer I and the convolution layer of the classification head; the convolution layer I is connected to the Maxpool layer of the residual module I and the convolution layer of the upsampling module III; the convolution layer of the residual module I is connected to the convolution layer of the residual module II and the convolution layer of the upsampling module II; the convolution layer of the residual module I is connected to the convolution layer of the void spatial convolution pooling pyramid I of each branch; the feature splicing layer II of the void spatial convolution pooling pyramid I is connected to the convolution layer of the upsampling module I; the feature splicing layer III of the upsampling module I is connected to the convolution layer of the upsampling module II; the feature splicing layer IV of the upsampling module II is connected to the convolution layer of the upsampling module III; the feature splicing layer V of the upsampling module III is connected to the convolution layer of the classification head.

[0073] Furthermore, in step 5, the error area is generated according to the prediction results, and then the correction point is obtained, which is updated to the interaction point feature map of step 2 and re-input into the network. Specifically:

[0074] The current segmentation result obtained by the model is compared with the true label of the target area. Each time, the largest area of ​​incorrect segmentation is selected, and the centroid of the largest area is located as the segmentation coordinate of the newly generated interaction point. This process is repeated in a cycle, which is called iterative training. At the same time, the center of the largest error area is used as the active point to generate the corresponding passive point for the other branch.

[0075] Furthermore, in step 6, the iteration is stopped when the prediction accuracy reaches 85%, and the final result is obtained, which is specifically:

[0076]

[0077] Among them, IoU is the degree of overlap between the predicted instance mask and the true instance mask, that is, the prediction accuracy;

[0078] I(X) is the pixel count of the intersection of the predicted instance mask and the true instance mask;

[0079] U(X) is the pixel count of the union of the predicted instance mask and the true instance mask;

[0080] I(X) and U(X) are approximated as follows:

[0081]

[0082]

[0083] Where I(X) is the pixel count of the intersection of the predicted instance mask and the true instance mask;

[0084] U(X) is the pixel count of the union of the predicted instance mask and the true instance mask;

[0085] v is a pixel in the image to be segmented;

[0086] V is the set of all pixels in the image to be segmented;

[0087] X is the pixel probability on the set V output by the network;

[0088] Y∈{0, 1} V is the ground truth assignment of the set V, where 0 represents background pixels and 1 represents object pixels.

[0089] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0090] Example

[0091] Combine Figures 1 and 2 The interactive semantic instance joint segmentation method based on a dual-branch iterative correction network of the present invention comprises the following steps:

[0092] Step 1: In the data preprocessing stage, instance interaction points and corresponding semantic interaction points are generated according to the user's segmentation intention;

[0093] Step 2: When a point is generated, the interaction point is divided into two types: active and passive according to the point generation method and another branch is added;

[0094] Step 3: Concatenate the three RGB channels of the image, the two channels of the semantic points, and the two channels of the instance points as the input of the segmentation network;

[0095] Step 4: The input information is processed through the feature extraction trunk and the dual segmentation branch to obtain semantic prediction results and instance prediction results;

[0096] Step 5: Generate the error area based on the prediction results, and then obtain the correction point, update it to the interaction point feature map, and re-enter the network;

[0097] Step 6: When the prediction accuracy reaches 85%, stop the iteration and get the final result;

[0098] Furthermore, in step 1, instance interaction points and corresponding semantic interaction points are generated, specifically:

[0099] During data preprocessing, all individuals in the image are traversed, and the geometric center of each selected individual is used as the instance interaction point. Based on the type of individual where the instance interaction point is located, all individuals of that type are traversed in the image, and the geometric center of each individual is used as the semantic interaction point. Interaction points are divided into foreground interaction points and background interaction points, corresponding to false negative and false positive error areas, respectively.

[0100] Furthermore, in step 2, the interaction points are divided into active and passive types and another branch is added, specifically:

[0101] The goal is to enable semantic instances in a dual-branch network to share information about the same coordinates. When an interaction point is added to a semantic branch or instance branch, the corresponding point of the foreground and background type of the location point in the other branch is calculated and added to the other branch. The interaction point is defined as an active point, and the corresponding point is a passive point. Specifically:

[0102] Let instance_msk be the mask for selecting a single instance within the instance ground truth, and class_msk be the mask for the single class to which the above single instance belongs within the semantic ground truth.

[0103] (1) Each time a point is actively added to the instance foreground channel, a passive point is added to the same position in the semantic foreground channel.

[0104] (2) Each time a point is actively added to the instance background channel, it is first determined whether the point with the same coordinate in class_msk belongs to the foreground / background, and then the semantic foreground / background passive point is added.

[0105] (3) Each time a point is actively added to the semantic background channel, a passive point is added to the same position in the instance background channel.

[0106] (4) Each time a point is actively added to the semantic foreground channel, it is first determined whether the point with the same coordinate in instance_msk belongs to the foreground / background, and then the instance foreground / background passive point is added.

[0107] Furthermore, in step 3, the three RGB channels of the image, the two channels of the semantic points, and the two channels of the instance points are linked, specifically:

[0108] The two channels for semantic points are the semantic foreground point channel and the semantic background point channel, and the two channels for instance points are the instance foreground point channel and the instance background point channel. Set the image resolution to 320*480*3 pixels, and set each point channel to a 320*480*2 pixel feature information map.

[0109] Furthermore, in step 4, the input information is processed through the feature extraction trunk and the dual segmentation branch to obtain the semantic prediction results and instance prediction results, specifically:

[0110] The feature extraction backbone is ResNet, which includes a convolutional layer I, a residual module I, and a residual module II, and inputs the extracted features into two segmentation branches respectively. The segmentation branch specifically includes a void space convolution pooling pyramid I, an upsampling module I, an upsampling module II, an upsampling module III, and a classification head. The feature splicing layer I of the connection channel module is connected to the convolutional layer I and the convolutional layer of the classification head. The convolutional layer I is connected to the Maxpool layer of the residual module I and the convolutional layer of the upsampling module III. The convolutional layer of the residual module I is connected to the convolutional layer of the residual module II and the convolutional layer of the upsampling module II. The convolutional layer of the residual module I is connected to the convolutional layer of the void space convolution pooling pyramid I of each branch. The feature splicing layer II of the void space convolution pooling pyramid I is connected to the convolutional layer of the upsampling module I. The feature splicing layer III of the upsampling module I is connected to the convolutional layer of the upsampling module II. The feature concatenation layer IV of the upsampling module II is connected to the convolution layer of the upsampling module III. The feature concatenation layer V of the upsampling module III is connected to the convolution layer of the classification head.

[0111] Furthermore, in step 5, the correction point is obtained, updated to the interaction point feature map, and re-input into the network, specifically:

[0112] The current segmentation result obtained by the model is compared with the true label of the target area. Each time, the largest area of ​​incorrect segmentation is selected, and the centroid of the largest area is located as the segmentation coordinate of the newly generated interaction point. This process is repeated in a cycle, which is called iterative training. At the same time, the center of the largest error area is used as the active point to generate the corresponding passive point for the other branch.

[0113] Furthermore, in step 6, the iteration is stopped when the prediction accuracy reaches 85%, and the final result is obtained, which is:

[0114] Let V = {1, 2, ..., N} be the set of all pixels of all images in the training set, X be the output of the network (outside the sigmoid layer) representing the pixel probabilities on the set V, and Y ∈ {0, 1} V is the ground truth assignment of the set V, where 0 represents background pixels and 1 represents object pixels. Then, the IoU count can be defined as:

[0115]

[0116] Among them, I(X) and U(X) can be approximated as follows:

[0117]

[0118]

[0119] The IoU is the prediction accuracy.

[0120] This embodiment uses an RGB three-dimensional image as input, uses an iterative training method to simulate interaction during the training phase, and accepts user input of interaction point information, including foreground points and background points, during the testing phase. It ultimately generates a semantic instance foreground-background segmentation result in the form of a two-dimensional vector of the same size as the RGB image, where a pixel value of 1 represents the foreground and a pixel value of 0 represents the background.

[0121] (1) The SBD dataset provided in the paper "Semantic contours from inverse detectors" is an image segmentation dataset, with a training set containing 8498 images and a validation set containing 2820 images. This invention uses the SBD dataset training set as the training dataset, uniformly transforming the input images to a size of 320*480 and performing standardization and normalization. During the training process, an iterative training strategy is used to generate interaction points. That is, the ground truth label is compared with the segmentation result of the previous iteration, and the centroid of the area with the largest error is taken as the newly added interaction point.

[0122] (2) During the training process, the goal is to achieve the sharing of information on the same coordinates of semantic instances in the dual-branch network. It is stipulated that when an interaction point is added to the semantic branch or the instance branch, the information of the point is first calculated based on the correct value of the data set to determine whether it belongs to the foreground point or the background point. At the same time, the corresponding point of the location point in the foreground and background type of the other branch is calculated and added to the other branch. Among them, the interaction point is defined as the active point, and the corresponding point is the passive point.

[0123] (3) Link the three RGB channels of the image, the two channels of semantic points, and the two channels of instance points. The two channels of semantic points are the semantic foreground point channel and the semantic background point channel, and the two channels of instance points are the instance foreground point channel and the instance background point channel. Set the image resolution to 320*480*3 pixels, and set each point channel to a 320*480*2 pixel feature information map.

[0124] (4) The input information passes through the feature extraction trunk and the dual segmentation branches to obtain semantic prediction results and instance prediction results. The feature extraction trunk is ResNet, which includes a convolution layer I, a residual module I, and a residual module II. The extracted features are input into the two segmentation branches respectively. The segmentation branch specifically includes a void space convolution pooling pyramid I, an upsampling module I, an upsampling module II, an upsampling module III, and a classification head. The feature splicing layer I of the connection channel module is connected to the convolution layer I and the convolution layer of the classification head. The convolution layer I is connected to the Maxpool layer of the residual module I and the convolution layer of the upsampling module III. The convolution layer of the residual module I is connected to the convolution layer of the residual module II and the convolution layer of the upsampling module II. The convolution layer of the residual module I is connected to the convolution layer of the void space convolution pooling pyramid I of each branch. The feature splicing layer II of the void space convolution pooling pyramid I is connected to the convolution layer of the upsampling module I. The feature concatenation layer III of upsampling module I is connected to the convolution layer of upsampling module II. The feature concatenation layer IV of upsampling module II is connected to the convolution layer of upsampling module III. The feature concatenation layer V of upsampling module III is connected to the convolution layer of the classification head.

[0125] (5) Compare the segmentation result currently obtained by the model with the true label of the target area, select the largest area of ​​incorrect segmentation each time, locate the centroid of the largest area as the segmentation coordinate of the newly generated interaction point, and repeat this process, which is called iterative training; at the same time, the center of the largest error area is used as the active point to generate the corresponding passive point for the other branch. Figure 3 As shown, Figure 3 The first and fourth rows show the data instances and semantically correct values, respectively. The second and fifth rows show the network segmentation results. In the colored image, red dots represent active points, and green dots represent automatically generated passive points. × represents background points. The five columns represent the segmentation results after adding correction points for five iterations.

[0126] (6) When the prediction accuracy reaches 85%, the iteration is stopped and the final result is obtained. Let V = {1, 2, ..., N} be the set of all pixels of all images in the training set, X be the output of the network (outside the sigmoid layer) representing the pixel probability on the set V, and Y∈{0, 1} V is the ground truth assignment of the set V, where 0 represents background pixels and 1 represents object pixels. Then, the IoU count can be defined as:

[0127]

[0128] Among them, I(X) and U(X) can be approximated as follows:

[0129]

[0130]

[0131] The IoU is the prediction accuracy.

[0132] (7) The final experimental results are as follows Figure 4 shown. Figure 4 This demonstrates the complete iterative interaction process of the present invention. As the number of interaction points increases, the model can achieve better segmentation results at a faster speed, meaning less interaction cost. On average, adding five interaction points can yield a segmentation result with an IoU > 85%.

[0133] The above-mentioned semantic instance dual-branch joint segmentation design enables the model to fully explore the actual user intentions implied by the interaction points during the iterative training process, thereby obtaining better segmentation results at a lower cost; compared with the existing technology, the present invention has the following significant advantages: interactive semantic instance joint segmentation takes into account the characteristics of traditional interactive image segmentation that cannot accurately perceive the user's segmentation intentions, outputs two image segmentation results, and at the same time, combines the semantic segmentation and instance segmentation tasks for multi-task learning, which improves the training quality and segmentation accuracy compared with traditional interactive image segmentation technology.

Claims

1. An interactive semantic instance joint segmentation method based on a dual-branch iterative correction network, characterized by: The following steps are involved: Step 1: In the data preprocessing stage, instance selection points and corresponding semantic selection points are generated according to the user's segmentation intention. These points are converted into disks and used as the initial elements in the interaction point feature map. Step 2: When an instance active selection point is generated, a corresponding semantic passive selection point is added to the interaction point feature map of the semantic branch. When a semantic active selection point is generated, a corresponding instance passive selection point is added to the interaction point feature map of the instance branch. Step 3: The segmentation network uses a parameter-sharing residual network as the encoder, and the parallel symmetrical ASPP module, upsampling module, and segmentation module are connected in sequence as the decoder; the three RGB channels of the image, the two channels of semantic points, and the two channels of instance points are concatenated as the input of the segmentation network; Step 4: The input information passes through the feature extraction trunk and the parallel symmetrical semantic branch and instance branch to obtain the semantic prediction results and instance prediction results; Step 5: Generate the error area based on the prediction results, and then obtain the correction point, update it to the interaction point feature map of step 2, and re-input it into the network; Step 6: When the prediction accuracy reaches 85%, stop the iteration and get the final result; In step 1, instance selection points and corresponding semantic selection points are generated, specifically: During data preprocessing, all instances in the graphics to be segmented are traversed, and a selected instance is obtained each time. The selected instance is used as the reference instance, and the geometric center of the reference instance is used as the instance selection point. Set the category of the benchmark instance individual as the benchmark category, and use the geometric center of the benchmark category in the semantic segmentation ground_truth as the semantic selection point; Interaction points are divided into foreground interaction points and background interaction points, which correspond to the two types of error areas, false negative and false positive, respectively; Step 2 implements the sharing of same-coordinate information of semantic instances in a dual-branch network. It stipulates that when an interaction point is added to a semantic branch or an instance branch, the corresponding point of the point in the foreground and background type of the other branch is calculated and added to the other branch. The interaction point is defined as an active point and the corresponding point is a passive point. Specifically: Let instance_mask be the mask of the reference instance, and class_mask be the mask of the reference class; (1) Each time an active point is added to the instance foreground channel, a passive point is added to the same position in the semantic foreground channel; (2) Each time an active point is added to the instance background channel, first determine whether the point with the same coordinate in class_mask belongs to the foreground or the background. If it belongs to the foreground, add a semantic foreground passive point; if it belongs to the background, add a semantic background passive point; (3) Each time an active point is added to the semantic background channel, a passive point is added to the same position in the instance background channel; (4) Each time an active point is added to the semantic foreground channel, first determine whether the point with the same coordinate in instance_msk belongs to the foreground or the background. If it belongs to the foreground, add the instance foreground passive point; if it belongs to the background, add the instance background passive point.

2. The interactive semantic instance joint segmentation method based on a dual-branch iterative correction network according to claim 1 is characterized in that In step 3, the three RGB channels of the image, the two channels of the semantic points, and the two channels of the instance points are concatenated as the input of the segmentation network, specifically: The two channels of semantic points are the semantic foreground point channel and the semantic background point channel, and the two channels of instance points are the instance foreground point channel and the instance background point channel. The image resolution is set to 320*480*3 pixels, and each point channel is set to a feature information map of 320*480*2 pixels.

3. The interactive semantic instance joint segmentation method based on a dual-branch iterative correction network according to claim 1 is characterized in that In step 4, the input information passes through the feature extraction trunk and the parallel symmetrical semantic branch and instance branch to obtain the semantic prediction results and instance prediction results, specifically: The feature extraction backbone is ResNet, which includes a convolutional layer I, a residual module I, and a residual module II. The extracted features are input into two segmentation branches respectively; the segmentation branch specifically includes a dilated spatial convolution pooling pyramid I, an upsampling module I, an upsampling module II, an upsampling module III, and a classification head; The feature splicing layer I of the connection channel module is connected to the convolution layer I and the convolution layer of the classification head; the convolution layer I is connected to the Maxpool layer of the residual module I and the convolution layer of the upsampling module III; the convolution layer of the residual module I is connected to the convolution layer of the residual module II and the convolution layer of the upsampling module II; the convolution layer of the residual module I is connected to the convolution layer of the void spatial convolution pooling pyramid I of each branch; the feature splicing layer II of the void spatial convolution pooling pyramid I is connected to the convolution layer of the upsampling module I; the feature splicing layer III of the upsampling module I is connected to the convolution layer of the upsampling module II; the feature splicing layer IV of the upsampling module II is connected to the convolution layer of the upsampling module III; the feature splicing layer V of the upsampling module III is connected to the convolution layer of the classification head.

4. The interactive semantic instance joint segmentation method based on a dual-branch iterative correction network according to claim 1, characterized in that In step 5, the error area is generated based on the prediction results, and then the correction point is obtained. The interaction point feature map of step 2 is updated and re-input into the network. Specifically: The current segmentation result obtained by the model is compared with the true label of the target area. Each time, the largest area of ​​incorrect segmentation is selected, and the centroid of the largest area is located as the segmentation coordinate of the newly generated interaction point. This process is repeated in a cycle, which is called iterative training. At the same time, the center of the largest error area is used as the active point to generate the corresponding passive point for the other branch.

5. The interactive semantic instance joint segmentation method based on a dual-branch iterative correction network according to claim 1 is characterized in that In step 6, when the prediction accuracy reaches 85%, the iteration is stopped and the final result is obtained, which is: ; Among them, IoU is the degree of overlap between the predicted instance mask and the true instance mask, that is, the prediction accuracy; is the pixel count of the intersection of the predicted instance mask and the true instance mask; is the pixel count of the union of the predicted instance mask and the ground-truth instance mask; and It is approximately as follows: ; in, is a pixel in the image to be segmented; is the set of all pixels in the image to be segmented; is the set of network outputs Pixel probability on ; is a collection where 0 represents background pixels and 1 represents object pixels.

Citation Information

Patent Citations

  • Interactive image segmentation method based on iterative selection-correction network

    CN113689437A

  • Interactive image segmentation method based on unsupervised learning

    CN116109656A