Instance Segmentation Method Based on Detection Enhancement and Multi-Stage Bounding Box Feature Refinement

By constructing a two-branch object detection head for decoupled classification and regression in the instance segmentation model and the instance segmentation head of the multi-stage refinement and enhancement spatial hollow fusion module, the detection branch is used to enhance segmentation branches, and the instance segmentation accuracy reduction caused by inaccurate bounding boxes in the two-stage model is solved, and high-quality instance segmentation mask results are achieved.

CN115797629BActive Publication Date: 2025-06-27FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211502719.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-06-27
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The two-stage instance segmentation model results in a degraded instance segmentation accuracy when the bounding box predicted by the object detection is not accurate enough or the category is wrong, and the existing model ignores the potential enhancement of the segmentation branch by detecting branches.

Method used

An instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement is proposed. By constructing a dual-branch object detection head for decoupled classification and regression and an instance segmentation head for multi-stage refinement and enhancement spatial hollow fusion module, the detection branch is enhanced and segmented.

Benefits of technology

The segmentation is enhanced by indirect and direct methods using detection, which significantly improves the accuracy and quality of instance segmentation and generates high-quality instance segmentation mask results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797629B_ABST
    Figure CN115797629B_ABST
Patent Text Reader

Abstract

The present invention proposes an instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement, comprising the following steps: performing data preprocessing on the images in the training set, including data augmentation and normalization; constructing a decoupled classification and regression dual-branch object detection head; constructing an instance segmentation head that fuses a multi-stage refinement and enhanced spatial hole module, and enhancing the refinement process of the instance segmentation head in multiple stages using the bounding box features of the object detection head; constructing an instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement; training the instance segmentation network using the images in the training set, generating instance segmentation results and calculating the loss function, and backpropagating to optimize the parameters of the entire network to obtain a trained instance segmentation network; inputting the image to be processed into the trained instance segmentation network to obtain the instance segmentation result. This method can improve the performance of instance segmentation through object detection and generate high-quality instance segmentation masks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of image processing and computer vision, and particularly relates to an instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement. Background Art

[0002] Nowadays, with the rapid development of computer hardware and artificial intelligence, machines can imitate or even replace some human behaviors through artificial intelligence algorithms such as machine learning and deep learning, and perform repetitive, single and intelligent tasks. However, there is still a certain gap in imitating more advanced and complex human behaviors. As an important part of the field of artificial intelligence, computer vision tasks recognize external information by imitating the human brain and sensory organs, thus helping human work in all aspects of life. Instance segmentation is a classic and challenging task in the field of computer vision, aiming to use image masks to perform pixel-level landmarking on different categories and different individuals (instances) existing in the image, so as to clearly distinguish the boundaries and locations of foreground objects. Instance segmentation is widely used in fields such as autonomous driving, medical image analysis, security prevention and control, and industrial sorting. Due to its wide applicability and industrial value, it has also attracted many domestic and foreign scholars to conduct research in this field. At the same time, as an upstream task of some visual processing tasks, in order to better provide accurate object masks for downstream visual tasks (such as 3D reconstruction, etc.), the importance of designing an instance segmentation model with fast segmentation speed, high segmentation accuracy and good robustness is self-evident.

[0003] With the continuous research on object detection algorithms, the research progress of instance segmentation also benefits from powerful object detectors. According to different actual objects, it is divided into single-stage instance segmentation models and two-stage instance segmentation models. The former aims to achieve high efficiency in real-time segmentation, while the latter hopes to use more complex models to make the segmentation of masks more accurate. Some high-precision instance segmentation models are based on the two-stage paradigm designed by Mask R-CNN. They use an object detector to first detect the bounding boxes of instance objects, and then perform instance segmentation within the obtained bounding boxes. Therefore, in two-stage instance segmentation models, if the bounding boxes predicted by object detection are not accurate enough (for example, unable to completely enclose the object) or the predicted categories are incorrect, the accuracy of instance segmentation will also be affected and decreased accordingly. Some existing advanced two-stage instance segmentation models improve the instance segmentation head based on Mask R-CNN, thereby greatly enhancing the effect of the predicted mask, making the edges more refined and the accuracy higher. However, they are limited to simply improving the segmentation branch to enhance the segmentation effect, which only approaches the upper limit of the bounding boxes obtained by the detection branch, but ignores the problem that the detection branch brings an upper limit to the segmentation branch. Although there are individual instance segmentation models that improve the detection branch to increase the upper limit, they only indirectly enhance the segmentation branch, and the two branches still exist independently, ignoring that the detection branch is also helpful for the segmentation branch. Summary of the Invention

[0004] To solve the problems that the two-stage model ignores the detection branch as the upper limit of the segmentation branch, has an indirect limitation on the segmentation branch, and can be directly enhanced, the present invention proposes an instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement. This method first constructs a two-branch object detection head that decouples classification and regression, and uses the detection to indirectly enhance the upper limit of instance segmentation, thereby indirectly enhancing the segmentation. Then, it proposes an instance segmentation head that multi-stage refines and enhances the spatial hole fusion module, and directly enhances the segmentation branch by periodically adding the bounding box features of the detection branch. Thus, the segmentation is enhanced by using detection in two indirect and direct ways, and a high-quality instance segmentation mask result is obtained.

[0005] This method can enhance the segmentation by using detection in two indirect and direct ways, and further obtain a high-quality instance segmentation mask result.

[0006] The method includes the following steps: performing data preprocessing on the images in the training set, including data augmentation and normalization; constructing a decoupled classification and regression dual-branch object detection head; constructing an instance segmentation head with a multi-stage refinement and enhanced spatial hole fusion module, and enhancing the refinement process of the instance segmentation head in multiple stages using the bounding box features of the object detection head; constructing an instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement; training the instance segmentation network using the images in the training set, generating instance segmentation results and calculating the loss function, and backpropagating to optimize the parameters of the entire network to obtain a trained instance segmentation network; inputting the image to be processed into the trained instance segmentation network to obtain the instance segmentation result. This method can improve the performance of instance segmentation through object detection and generate high-quality instance segmentation masks.

[0007] The technical solution adopted by the present invention to solve its technical problems is:

[0008] An instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement, characterized by including the following steps:

[0009] Step A: Performing data preprocessing on the images in the training set, including data augmentation and normalization;

[0010] Step B: Constructing a decoupled classification and regression dual-branch object detection head;

[0011] Step C: Constructing an instance segmentation head with a multi-stage refinement and enhanced spatial hole fusion module, and enhancing the refinement process of the instance segmentation head in multiple stages using the bounding box features of the object detection head;

[0012] Step D: Constructing an instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement;

[0013] Step E: Training the instance segmentation network using the images in the training set, generating instance segmentation results and calculating the loss function, and backpropagating to optimize the parameters of the entire network to obtain a trained instance segmentation network;

[0014] Step F: Inputting the image to be processed into the trained instance segmentation network to obtain the instance segmentation result.

[0015] Further, step A specifically includes the following steps:

[0016] Step A1: Performing scale transformation on the images in the training set. Without changing the aspect ratio, setting the threshold for the length and width of the image to 2048; that is, performing scale transformation on the image according to the long side of the image and the threshold to ensure that neither the long side nor the short side exceeds the threshold; then randomly flipping all the images after scale transformation to achieve data augmentation;

[0017] Step A2: Normalize the enhanced image. The mean values for normalization are [123.675, 116.28, 103.53], and the variance values are [58.395, 57.12, 57.375]. Finally, pad the image so that its length and width are divisible by 32. Each image has a corresponding label, and the label content is the bounding box and mask of each instance object in the image. Synchronously process the image label while performing image scale transformation, data augmentation, and padding.

[0018] Further, in step B, the implementation method of the decoupled classification and regression dual-branch object detection head is as follows:

[0019] Step B1: Extract features from the input image through the backbone network of the instance segmentation network, and use the Feature Pyramid Network (FPN) to construct multi-scale features P2 - P6, where P2 has the largest resolution and P6 has the smallest resolution. Then send them into the Region Proposal Network (RPN) sub-network for region candidate proposal. In the RPN sub-network, perform binary classification prediction of background and foreground, and map the coordinates of the candidate regions predicted as foreground back to the corresponding regions in the multi-scale features obtained by FPN. Use RoIAlign to pool the features of the corresponding regions into detection head features F of a fixed size (P, C, H, W). Bbox ; where P is the number randomly extracted from the candidate region results predicted as foreground in the RPN sub-network, C represents the number of channels of the features, and H and W represent the height and width of the detection head features. Then input the obtained detection head features F Bbox into the decoupled classification and regression dual-branch object detection head.

[0020] After passing through the decoupled classification and regression dual-branch object detection head designed in steps B2 and B3, obtain the predicted class and bounding box results. Map the predicted bounding box coordinates back to the corresponding regions in the multi-scale features obtained by FPN, and use RoIAlign to pool the features of the corresponding regions into segmentation head features F of a fixed size (P′, C, H′, W′). mask ; where P′ is the number of bounding boxes predicted by the object detection head, and H′ and W′ are the height and width of the segmentation head features. Then input the obtained segmentation head features F mask into the instance segmentation head.

[0021] Step B2: For the detection head features F Bbox obtained in step B1, flatten them into features F Bbox ′ of size (P, C×H×W). Utilize the characteristic that the fully connected layer is more spatially sensitive than the convolutional layer. In the classification branch, calculate the prediction of the classification label only using the fully connected layer. Pass the flattened features F Bbox ′ through three fully connected layers to predict the probability Class_Score of each class, including the background. The specific formula is as follows:

[0022] Class_Score = FC3(ReLU(FC2(ReLU(FC1(F Bbox ′))))),

[0023] where ReLU is the activation function; FC1 is a fully connected layer with an input channel number of C×H×W and an output channel number of 4×C; FC2 is a fully connected layer with both input and output channel numbers of 4×C; FC3 is a fully connected layer with an input channel number of 4×C and an output channel number of K + 1, where K is the total number of classes in the dataset excluding the background; then, the cross-entropy loss is calculated between the probability Class_Score of each predicted class, including the background, and the true label to obtain the loss L of the predicted classification cls ;

[0024] Step B3: Taking advantage of the fact that convolutional layers are more conducive to bounding box regression and localization, in the regression branch, multiple convolutional layers are used to calculate the bounding box regression task; the detection head feature F of size (P, C, H, W) obtained in Step B1 Bbox is passed through a residual block to increase the dimension to obtain the residual feature F res , with a size of (P, 4×C, H, W); the specific formula is as follows:

[0025] F res = ReLU(Conv 1×1 (F Bbox ) + Conv 1×1 (Conv 3×3 (F Bbox ))),

[0026] where ReLU is the activation function; Conv 3×3 is a 3×3 convolution with both input and output channel numbers of C, a stride of 1, and a padding of 1; the two Conv 1×1 are both 1×1 convolutions with an input channel number of C and an output channel number of 4×C; then, the residual feature F res is passed through four bottleneck layers to obtain the bottleneck layer feature F bott (i), and the specific formula is as follows:

[0027]

[0028] where BottleNeck i is the bottleneck layer, i = 1, 2, 3, 4; the specific formula for any bottleneck layer BottleNeck is as follows:

[0029] out = ReLU(in + Conv′ 1×1 (Conv 3×3 (Conv 1×1(in)))),

[0030] Among them, ReLU is the activation function; in and out are the input and output of any bottleneck layer respectively; Conv 1×1 is a 1×1 convolution with an input channel number of 4×C and an output channel number of C; Conv 3×3 is a 3×3 convolution with an input and output channel number of C, a stride of 1, and a padding of 1; Conv′ 1×1 is a 1×1 convolution with an input channel number of C and an output channel number of 4×C; then the features F of the last bottleneck layer bott (4) pass through average pooling and a fully connected layer to predict the four coordinates Bbox_Pred of the bounding box. The specific formula is as follows:

[0031] Bbox_Pred = FC(View(AP(F bottle (4)))),

[0032] Among them, AP is average pooling with a pooling window of H×W, changing the size of the bottleneck layer features from (P, 4×C, H, W) to (P, 4×C, 1, 1); View is a feature flattening operation, changing the features from (P 4×C, 1, 1) to (P, 4×C×1×1); FC is a fully connected layer with an input channel number of 4×C and an output channel number of 4×K, where K is the total number of categories in the dataset except the background; finally, the predicted bounding box Bbox_Pred and the true label are used to calculate the smooth L1 loss to obtain the loss L of the bounding box regression reg 。

[0033] Furthermore, in step C, the implementation method of the instance segmentation head of the multi-stage refinement and enhanced spatial hole fusion module is as follows:

[0034] Step C1: The instance segmentation head of the multi-stage refinement and enhanced spatial hole fusion module consists of three parts: a bounding box feature branch, a fine-grained feature branch, and an instance segmentation main branch;

[0035] Step C2: Design the bounding box feature branch to use the boundary features to enhance the refinement process of the instance segmentation main branch; for the bottleneck layer features F bott (i), i = 1, 2, 3, obtained in step B3 with a size of (P, 4×C, H, W), use RoIAlign pooling and a convolutional layer to obtain the bounding box features F box_feat (i) that match the size of the instance segmentation main branch designed in step C3. The specific formula is as follows:

[0036] F box_feat (i) = ReLU(Conv′ 1×1 (RoIAlign(ReLU(Conv 1×1 (Fbott (i))))))

[0037] Among them, ReLU is the activation function; i = 1, 2, 3, respectively corresponding to the three stages of the instance segmentation main branch; in the i-th stage, Conv 1×1 is a 1×1 convolution with 4×C input channels and output channels; RoIAlign changes the feature size from pooling to where P′ is the number of bounding boxes of the object detection prediction result, C represents the number of channels of the feature, and H′ and W′ are the width and height of the feature, which are the same as the width and height of the segmentation head feature F mask described in step B1; Conv′ 1×1 is a 1×1 convolution with both input and output channels of ; the obtained bounding box feature F box_feat (i) is added to the instance segmentation main branch stage by stage to enhance the refinement process of the instance segmentation head;

[0038] Step C3: Design a fine-grained feature branch to enhance the refinement process of the instance segmentation main branch using fine-grained features; the P2 layer feature in the multi-scale features obtained by the Feature Pyramid Network FPN is used to extract the fine-grained feature F fine through a lightweight semantic segmentation head composed of 4 stacked 3×3 convolutions. The specific formula is as follows:

[0039]

[0040] where, ReLU is the activation function; the four Cony 3×3 are all 3×3 convolutions with both input and output channels of C, a stride of 1, and a padding of 1; the size of the fine-grained feature F fine is where and represent the height and width of the feature, which are one-fourth of the height and width of the input image and are the same as the size of the multi-scale feature ; then, the RoIAlign pooling and convolutional layers are used to obtain the semantic segmentation feature F sema_feat (i) and the semantic segmentation mask F sema_mask (i) that match the size of the instance segmentation main branch designed in step C4. The specific formula is as follows:

[0041] F sema_feat (i) = ReLU(Conv′ 1×1 (RoIAlign(ReLU(Conv 1×1 (F fine )))))

[0042]

[0043] Among them, ReLU and Sigmoid are activation functions; i = 1, 2, 3, corresponding to the three stages of the instance segmentation main branch respectively; in the i-th stage, Conv 1×1 is a 1×1 convolution with an input channel number of C and an output channel number of ; RoIAlign changes the feature size to Conv′ 1×1 is a 1×1 convolution with both input and output channel numbers of ; is a 1×1 convolution with an input channel number of C and an output channel number of 1; RoIAlign mask changes the semantic segmentation mask size to (P′, 1, 2 i-1 ×H′, 2 i-1 ×w′); calculates the cross-entropy loss between the result of and the true label to obtain the loss L sema of the fine-grained features;

[0044] Step C4: Design the instance segmentation main branch; use the detection head feature F mask with a size of (P′, C, H′, W′) obtained in step B1, and extract the instance segmentation feature F ins_feat (1) through a lightweight segmentation head composed of 2 stacked 3×3 convolutions. The specific formula is as follows:

[0045] F ins_feat (1) = ReLU(Conv 3×3 (ReLU(Conv 3×3 (F mask ))),

[0046] where ReLU is the activation function; both Conv 3×3 are 3×3 convolutions with an input and output channel number of C, a stride of 1, and a padding of 1; pass F ins_feat (1) through a 1×1 convolution and a Sigmoid activation function to obtain an instance segmentation mask F ins_mask (1) with a size of (P′, K, H′, W′). The specific formula is as follows:

[0047] F ins_mask (1) = Sigmoid(Conv 1×1 (F ins_feat (1))),

[0048] where Sigmoid is the activation function; Conv 1×1It is a 1×1 convolution with the number of input channels being C and the number of output channels being K, where K is the total number of categories in the dataset excluding the background; in the first stage, the enhanced spatial atrous fusion module ESAFM is used to process the bounding box feature F box_feat (1) obtained in step C2, the semantic segmentation feature F sema_feat (1) obtained in step C3, and the semantic segmentation mask F sema_mask (1), as well as the instance segmentation feature F ins_feat (1) and the instance segmentation mask F ins_mask (1) for fusion to obtain the fused instance segmentation feature F fused_feat (1); the fused instance segmentation feature F fused_feat (1) is upsampled by a factor of 2 using bilinear interpolation to obtain the -sized instance segmentation feature F ins_feat (2) as the input for the second stage;

[0049] In the i-th, i = 2, 3 stages, the instance segmentation feature F ins_feat (i) input in the (i - 1)-th stage is passed through a 1×1 convolution and a Sigmoid activation function to obtain an instance segmentation mask F i-1 ×H′, 2 i-1 ×W′) of size (P′, K, 2 ins_mask (i); then the enhanced spatial atrous fusion module ESAFM is used to fuse the five features to obtain a fused instance segmentation feature F of size fused_feat (i), which is upsampled by a factor of 2 using bilinear interpolation to obtain the instance segmentation feature F ins_feat (i + 1) as the input for the next stage. The specific formula is as follows:

[0050] F ins_mask (i) = Sigmoid(Conv 1×1 (F ins_mask (i))),

[0051] F fused_feat (i) = ESAFM i (F box_feat (i), F sema_feat (i), F sema_mask (i), F ins_feat (i), F ins_mask (i)),

[0052] F ins_feat (i + 1) = 2xUP(F fused_feat (i)),

[0053] where \(i = 2, 3\); Sigmoid is the activation function; in the \(i\)-th stage, Conv 1×1 is a \(1\times1\) convolution with the number of input channels being and the number of output channels being \(K\), and ESAFM i represents the enhanced spatial atrous fusion module for different stages; \(2xUP\) refers to bilinear interpolation upsampling by a factor of 2; finally, F ins_feat (4) passes through a \(1\times1\) convolution and the Sigmoid activation function to obtain an instance segmentation mask \(F ins_mask (4) of size \((P', K, 8\times H', 8\times W')\);

[0054] Step C5: Calculate the cross-entropy loss between the four instance segmentation masks \(F i-1 \) of size \((P', K, 2 i-1 \times H', 2 ins_mask \times W')\) obtained in Step C4 and the ground truth labels, and obtain the loss \(L ins_stage (i)\) of the instance segmentation mask for each stage, where \(i = 1, 2, 3, 4\); the total loss \(L ins of the instance segmentation mask is the sum of the losses of the four stages multiplied by weights, and the specific formula is as follows:

[0055] L ins =\omega_1\times L ins_stage (1)+\omega_2\times L ins_stage (2)+\omega_3\times L ins_stage (3)+\omega_4\times L ins_stage (4).

[0056] Furthermore, in Step C4, the implementation method of the enhanced spatial atrous fusion module is as follows:

[0057] Step C41: In the enhanced spatial atrous fusion module ESAFM i in the \(i\)-th stage, \(i = 1, 2, 3\), first select the mask of the true class from the masks predicting \(K\) classes using the ground truth labels for the instance segmentation mask of size \((P', K, 2 i-1 \times H', 2 i-1 \times W')\), and change the size of the instance segmentation mask to \((P', 1, 2 i-1 \times H', 2 i-1 \times W')\); then, for the bounding box feature \(F \), the semantic segmentation feature \(F box_feat (i)\), the instance segmentation feature \(F sema_feat (i)\), and the semantic segmentation mask \(F ins_feat (i)\) of size \((P', 1, 2 i-1 \times H', 2 i-1 \times W')\), sema_mask(i) and instance segmentation mask F ins_mask (i) After feature concatenation, perform a 1×1 convolution to obtain the preliminarily aggregated feature F aggr_feat_in (i), and the specific formula is as follows:

[0058]

[0059] Among them, Concat is the feature concatenation operation, and Conv 1×1 is a 1×1 convolution with the number of input channels being and the number of output channels being ;

[0060] Step C42: Pass the preliminarily aggregated feature F aggr_feat_in (i) through global average pooling and a 1×1 convolution to extract the detailed feature F detail_feat (i), and the specific formula is as follows:

[0061] F detail_feat (i) = UP(Conv 1×1 (GAP(F aggr_feat_in (i)))), i = 1, 2, 3

[0062] Among them, GAP is global average pooling, which changes the feature from to Conv 1×1 is a 1×1 convolution with both the number of input and output channels being , and UP represents bilinear interpolation upsampling to change the feature back to Then, pass the preliminarily aggregated feature F aggr_feat_in (i) through three parallel dilated convolutions with different dilation rates to extract features with different receptive fields. The specific formula is as follows:

[0063]

[0064]

[0065]

[0066] Among them, ReLU is the activation function; and are dilated convolutions with different dilation rates. The dilation rates are a, b, c respectively, and the kernel size is 3×3; in different stages, the values of the dilation rates a, b, c are different from each other and increase with the increase of the stage, so as to extract features with larger receptive fields; finally, add the detailed feature F detail_feat (i) element-wise to the features with different receptive fields to obtain the multi-receptive field feature F multi_feat (i), and the specific formula is as follows:

[0067] Fmulti_feat (i) = Add(F detail_feat (i), F dila_feat_1 (i), F dila_feat_2 (i), F dila_feat_3 (i)), i = 1, 2, 3

[0068] where Add represents element - by - element addition;

[0069] Step C43: The preliminarily aggregated feature F aggr_feat_in (i) is passed through an Enhanced Spatial Attention (ESA) module to obtain the enhanced spatial attention feature F en_spat_atten_feat (i); In the ESA module, the input F aggr_feat_in (i) is passed through a 3×3 convolution to obtain the initial feature F init_feat (i); Then, max - pooling is used to extract the spatial attention feature while reducing the number of parameters, and then it passes through a convolution group composed of three 3×3 convolutions, and then up - sampled back to the original size to obtain the spatial depth feature F spat_deep (i); Then, the initial feature F init_feat (i) is added element - by - element with the spatial depth feature through a 1×1 convolution to obtain the spatially enhanced feature F en_spat (i); The specific formula is as follows:

[0070] F init_feat (i) = ReLU(Conv 3×3 (F aggr_feat_in (i))),

[0071] F spat_deep (i) = UP(ConvGroup(MP(F init_feat (i)))),

[0072] F en_spat (i) = Add(F spat_deep (i) + Conv 1×1 (F init_feat (i))), i = 1, 2, 3

[0073] where ReLU is the activation function; Conv 3×3 has the same number of input and output channels as , with a stride of 1, padding of 1 for the 3×3 convolution; MP is max - pooling with a pooling window of 7 and a stride of 3; UP is bilinear interpolation up - sampling; Conv 1×1 has the same number of input and output channels as for the 1×1 convolution; Add is element - by - element addition; ConvGroup is a convolution group composed of three 3×3 convolutions, and the specific formula is as follows:

[0074] out = Conv3×3 (ReLU(Conv 3×3 (ReLU(Conv 3×3 (in))))),

[0075] Among them, ReLU is the activation function; in and out are the input and output of the convolutional group ConvGroup respectively; the three Convs 3×3 both have the number of input and output channels as a 3×3 convolution with a stride of 1 and a padding of 1; then the spatially enhanced feature F en_spat (i) passes through a 1×1 convolution and the Sigmoid activation function to obtain the factor Factor for enhancing spatial attention i , and then element-wise multiplies it with the preliminarily aggregated feature F aggr_feat_in (i) to obtain the final spatially enhanced attention feature F en_spat_atten_feat (i), and the specific formula is as follows:

[0076] Factor i = Sigmoid(Conv 1×1 (F en_spat (i))),

[0077] F en_spat_atten_feat (i) = Factor i ×F aggr_feat_in (i), i = 1, 2, 3

[0078] Among them, Sigmoid is the activation function; Conv 1×1 is a 1×1 convolution with the number of input and output channels both being ;

[0079] Step C44: After concatenating the multi-receptive field feature F multi_feat (i) obtained in step C42 and the final spatially enhanced attention feature F en_spat_atten_feat (i) obtained in step C43, pass through a 1×1 convolution to change the number of channels to obtain the preliminarily fused instance segmentation feature F init_fuse_feat (i); then pass F init_fuse_feat (i) through a 1×1 convolution to reduce the number of channels by 2, and then concatenate it with the semantic segmentation mask F sema_mask (i) and the instance segmentation mask F ins_mask (i) to obtain the output of the enhanced spatial hole fusion module ESAFM, and the fused instance segmentation feature F fused_feat (i), and the specific formula is as follows:

[0080] F init_fuse_feat (i) = ReLU(Conv 1×1 (Concat(F multi_feat(i), F en_spat_atten_feat (i))))

[0081] F fused_feat (i) = Concat(Conv′ 1×1 (F init_fuse_feat (i)), F sema_mask (i), F ins_mask (i)), i = 1, 2, 3

[0082] where ReLU is the activation function; Concat is feature concatenation; Conv 1×1 is a 1×1 convolution with an input channel number of and an output channel number of ; Conv′ 1×1 is a 1×1 convolution with an input channel number of and an output channel number of .

[0083] Furthermore, in step D, the implementation method of the instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement is as follows:

[0084] Step D1: Use the ResNet-50 backbone network as the feature extraction module to extract features from the input image, send the extracted feature maps into the Feature Pyramid Network (FPN) to construct multi-scale features, and then send them into the Region Proposal Network (RPN) sub-network for region candidate proposals;

[0085] Step D2: Perform background and foreground binary classification prediction in the RPN sub-network, map the coordinates of the candidate regions predicted as foreground to the multi-scale feature representation, then use RoIAlign to pool into fixed-size features, and send them into the decoupled classification and regression double-branch object detection head; after obtaining the bounding box results through the prediction of the double-branch object detection head, map the bounding box coordinates back to the multi-scale feature representation constructed by FPN, use RoIAlign to pool into fixed-size features and send them into the instance segmentation head of the multi-stage refinement and enhancement spatial hole fusion module;

[0086] Step D3: Use the bounding box features generated by the object detection head in step C2 and the fine-grained features generated by the FPN multi-scale features in step C3, and stagewise add them to the process of the instance segmentation main branch for mask segmentation to obtain the final instance segmentation result.

[0087] Furthermore, step E specifically includes the following steps:

[0088] Step E1: Input the images in the pre-processed training set into the instance segmentation network. After the multi-scale features extracted by the backbone network and the feature pyramid network are sent into the RPN sub-network to generate a given number of candidate regions, and the positive samples, i.e., foreground objects, and negative samples, i.e., background regions, are classified. Then, map the coordinates of the candidate regions predicted as foreground back to the multi-scale features, and use RoIAlign to pool them into the size of H×W and send them into the object detection head. Among the bounding box results predicted by the object detection head, map the coordinates back to the multi-scale features, and use RoIAlign to pool them into the sizes of H' and W' and send them into the instance segmentation head;

[0089] Step E2: The fine-grained features generated by the P2 layer of the multi-scale features represented by the FPN, and the bounding box features generated by the object detection head, are added to the instance segmentation head; then use the fine-grained features to perform the segmentation of the semantic mask, and use the object detection head and the instance segmentation head to perform the detection of the bounding box and the segmentation of the mask;

[0090] Step E3: Calculate the loss L of the RPN sub-network rpn , the classification loss L corresponding to Step B2 cls and the regression loss L corresponding to Step B3 reg , the loss L of the fine-grained features corresponding to Step C3 sema , and the total loss L of the instance segmentation mask corresponding to Step C5 ins ; The total loss L all is the sum of the five losses multiplied by the weights, and the specific formula is as follows:

[0091] L all = θ1×L rpn + θ2×L cls + θ3×L reg + θ4×L sema + θ5×L ins ,

[0092] where θ1, θ2, θ3, θ4, and θ5 are hyperparameters; then use the backpropagation method to calculate the gradients of the parameters in the instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement, and use the stochastic gradient descent method to update the parameters of the instance segmentation network.

[0093] Furthermore, Step F specifically includes the following steps:

[0094] Step F1: Input the images without label information into the trained instance segmentation network for processing;

[0095] Step F2: Use the object detection head to predict the bounding boxes of the foreground objects in the image, and use the instance segmentation head to predict the mask results of each instance in the image; the results obtained by the instance segmentation head in the network are the final instance segmentation results.

[0096] Also, an instance segmentation device based on detection enhancement and multi-stage bounding box feature refinement, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement as described above.

[0097] Also, a non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement as described above.

[0098] Compared with the prior art, the present invention and its preferred solutions construct a decoupled classification and regression dual-branch object detection head, use detection to increase the upper limit of instance segmentation, thereby indirectly enhancing instance segmentation; at the same time, construct an instance segmentation head that multi-stages refines and enhances the spatial hole fusion module, and directly enhance the refinement process of the instance segmentation head in multiple stages using the bounding box features obtained from the detection branch. The present invention can not only increase the accuracy of object detection bounding boxes and categories, indirectly improve the upper limit of segmentation accuracy, but also strengthen the segmentation effects of several categories with long-tail effects, making the model more generalizable, with strong practicability and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] The following further details the present invention in conjunction with the drawings and specific embodiments:

[0100] Figure 1 It is a flowchart of the method implementation of the embodiment of the present invention.

[0101] Figure 2 It is a schematic structural diagram of the entire instance segmentation network in the embodiment of the present invention.

[0102] Figure 3 It is a schematic structural diagram of the decoupled classification and regression dual-branch object detection head and the multi-stage refined instance segmentation head in the embodiment of the present invention.

[0103] Figure 4 It is a schematic structural diagram of the enhanced spatial hole fusion module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0104] To make the features and advantages of this patent more obvious and understandable, the following specific embodiments are given for detailed description as follows:

[0105] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0106] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0107] As Figures 1 - 4 shown, the present invention provides an overall design process and implementation process of an instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement, including the following steps:

[0108] As Figure 1 shown, this embodiment provides an instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement, including the following steps:

[0109] Step A: Perform data preprocessing on the images in the training set, including data augmentation and normalization processing;

[0110] Step B: Construct a decoupled classification and regression double-branch object detection head;

[0111] Step C: Construct an instance segmentation head with a multi-stage refinement and enhanced spatial hole fusion module, and use the bounding box features of the object detection head to enhance the refinement process of the instance segmentation head in multiple stages;

[0112] Step D: Construct an instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement;

[0113] Step E: Use the images in the training set to train the instance segmentation network, generate instance segmentation results and calculate the loss function, and backpropagate to optimize the parameters of the entire network to obtain a trained instance segmentation network;

[0114] Step F: Input the image to be processed into the trained instance segmentation network to obtain the instance segmentation result.

[0115] In this embodiment, Step A specifically includes the following steps:

[0116] Step A1: Perform scale transformation on the images in the training set. While keeping the aspect ratio unchanged, set the threshold for the length and width of the image to 2048; that is, perform scale transformation on the image according to the long side of the image and the threshold to ensure that neither the long side nor the short side exceeds the threshold; then randomly flip all the images after scale transformation to achieve data augmentation.

[0117] Step A2: Perform normalization processing on the augmented images. The mean values for normalization are [123.675, 116.28, 103.53], and the variance values are [58.395, 57.12, 57.375]; finally, pad the images so that the length and width are divisible by 32; each image has a corresponding label, and the label content is the bounding box and mask of each instance object in the image. While performing image scale transformation, data augmentation, and padding, synchronously process the image labels.

[0118] Figure 3 This is a schematic diagram of the structure of the decoupled classification and regression dual-branch object detection head and the multi-stage refinement instance segmentation head in this embodiment. As Figure 3 shown, it includes a decoupled classification and regression dual-branch object detection head and a multi-stage refinement instance segmentation head. Among them, the implementation method of the decoupled classification and regression dual-branch object detection head is as follows:

[0119] Step B1: Extract features from the input image through the backbone network of the instance segmentation network, and use the Feature Pyramid Network (FPN) to construct multi-scale features P2 - P6 (P2 has the largest resolution, and P6 has the smallest resolution), and then send them into the Region Proposal Network (RPN) sub-network for region candidate proposals; perform background and foreground binary classification prediction in the RPN sub-network, and map the coordinates of the candidate regions predicted as foreground back to the corresponding regions in the multi-scale features obtained by FPN. Use RoIAlign to pool the features of the corresponding regions into detection head features F with a fixed size of (P, C, H, W). Bbox . Among them, P is the number randomly extracted from the candidate region results predicted as foreground in the RPN sub-network (P = 512), C represents the number of channels of the features, and H and W represent the height and width of the detection head features. Then input the obtained detection head features F Bbox into the decoupled classification and regression dual-branch object detection head.

[0120] After passing through the decoupled classification and regression dual-branch object detection head designed in Steps B2 and B3, the predicted category and bounding box results are obtained. Map the predicted bounding box coordinates back to the corresponding regions in the multi-scale features obtained by FPN, and use RoIAlign to pool the features of the corresponding regions into segmentation head features F with a fixed size of (P′, C, H′, W′). maskAmong them, P′ is the number of bounding boxes predicted by the target detection head, and H′ and W′ are the height and width of the segmentation head features. Then, the obtained segmentation head feature F mask is input into the instance segmentation head.

[0121] Step B2: For the detection head feature F obtained in Step B1 Bbox , it is flattened into a feature F Bbox ′ of size (P, C×H×W). Utilizing the characteristic that the fully connected layer is more spatially sensitive than the convolutional layer, in the classification branch, the prediction of the classification label is calculated only using the fully connected layer. The flattened feature F Bbox ′ passes through three fully connected layers to predict the probability Class_Score of each class, including the background. The specific formula is as follows:

[0122] Class_Score = FC3(ReLU(FC2(ReLU(FC1(F Bbox ′)))))

[0123] where ReLU is the activation function; FC1 is a fully connected layer with an input channel number of C×H×W and an output channel number of 4×C; FC2 is a fully connected layer with both input and output channel numbers of 4×C; FC3 is a fully connected layer with an input channel number of 4×C and an output channel number of K + 1, where K is the total number of classes in the dataset excluding the background. Then, the cross-entropy loss is calculated between the predicted probability Class_Score of each class (including the background) and the true label to obtain the loss L cls of the predicted classification.

[0124] Step B3: Taking advantage of the characteristic that the convolutional layer is more conducive to bounding box regression and localization, in the regression branch, the regression task of the bounding box is calculated using multiple convolutional layers. The detection head feature F of size (P, C, H, W) obtained in Step B1 Bbox passes through a residual block to increase the dimension to obtain a residual feature F res of size (P, 4×C, H, W). The specific formula is as follows:

[0125] F res = ReLU(Conv 1×1 (F Bbox ) + Conv 1×1 (Conv 3×3 (F Bbox )))

[0126] where ReLU is the activation function; Conv 3×3 is a 3×3 convolution with both input and output channel numbers of C, a stride of 1, and a padding of 1; the two Conv 1×1They are all 1×1 convolutions with an input channel number of C and an output channel number of 4×C. Then the residual feature F res passes through four bottleneck layers to obtain the bottleneck layer feature F bott (i) (i = 1, 2, 3, 4), and the specific formula is as follows:

[0127]

[0128] Among them, BottleNeck i is the bottleneck layer, and i = 1, 2, 3, 4. The specific formula for any bottleneck layer BottleNeck is as follows:

[0129] out = ReLU(in + Conv′ 1×1 (Conv 3×3 (Conv 1×1 (in)))).

[0130] Among them, ReLU is the activation function; in and out are the input and output of any bottleneck layer respectively; Conv 1×1 is a 1×1 convolution with an input channel number of 4×C and an output channel number of C; Conv 3×3 is a 3×3 convolution with an input and output channel number of C, a stride of 1, and a padding of 1; Conv′ 1×1 is a 1×1 convolution with an input channel number of C and an output channel number of 4×C. Then the feature F bott (4) of the last bottleneck layer passes through average pooling and a fully connected layer to predict the four coordinates Bbox_Pred of the bounding box, and the specific formula is as follows:

[0131] Bbox_Pred = FC(View(AP(F bottle (4)))).

[0132] Among them, AP is average pooling with a pooling window of H×W, which changes the size of the bottleneck layer feature from (P, 4×C, H, W) to (P, 4×C, 1, 1); View is the feature flattening operation, which changes the feature from (P, 4×C, 1, 1) to (P, 4×C×1×1); FC is a fully connected layer with an input channel number of 4×C and an output channel number of 4×K, where K is the total number of categories except the background in the dataset. Finally, the predicted bounding box Bbox_Pred and the ground truth label are calculated with smooth L1 loss to obtain the loss L reg .

[0133] As Figure 3 shown, it includes a dual-branch object detection head that decouples classification and regression and an instance segmentation head with multi-stage refinement. Among them, the implementation method of the instance segmentation head with multi-stage refinement and enhanced spatial hole fusion module is as follows:

[0134] Step C1: The instance segmentation head of the multi-stage refinement and enhanced spatial hole fusion module consists of three parts: a bounding box feature branch, a fine-grained feature branch, and an instance segmentation main branch.

[0135] Step C2: Design the bounding box feature branch to enhance the refinement process of the instance segmentation main branch using boundary features. For the bottleneck layer feature F of size (P, 4×C, H, W) obtained in step B3 bott (i), where i = 1, 2, 3, use RoIAlign pooling and convolutional layers to obtain the bounding box feature F box_feat (i) that matches the size of the instance segmentation main branch designed in step C3. The specific formula is as follows:

[0136] F box_feat (i) = ReLU(Conv′ 1×1 (RoIAlign(ReLU(Conv 1×1 (F bott (i)))))),

[0137] where ReLU is the activation function; i = 1, 2, 3, corresponding to the three stages of the instance segmentation main branch respectively. In the i-th stage, Conv 1×1 is a 1×1 convolution with an input channel number of 4×C and an output channel number of ; RoIAlign changes the feature size from to where P′ is the number of bounding boxes in the object detection prediction result, C represents the number of channels of the feature, and H′ and W′ are the width and height of the feature, which are the same as the width and height of the segmentation head feature F mask described in step B1; Conv′ 1×1 is a 1×1 convolution with both input and output channel numbers of . The obtained bounding box feature F box_feat (i) is added to the instance segmentation main branch stage by stage to enhance the refinement process of the instance segmentation head.

[0138] Step C3: Design the fine-grained feature branch to enhance the refinement process of the instance segmentation main branch using fine-grained features. The P2 layer feature in the multi-scale features obtained by the feature pyramid network FPN is used to extract the fine-grained feature F fins through a lightweight semantic segmentation head composed of 4 stacked 3×3 convolutions. The specific formula is as follows:

[0139]

[0140] where ReLU is the activation function; the four Conv 3×3All are 3×3 convolutions with the number of input and output channels both being C, a stride of 1, and a padding of 1. The fine-grained feature F fine has a size of where and represent the height and width of the feature, which are one-fourth of the height and width of the input image, and are the same size as the multi-scale feature . Then, the RoIAlign pooling and convolutional layer are used to obtain the semantic segmentation feature F sema_feat (i) and the semantic segmentation mask F sema_mask (i), and the specific formula is as follows:

[0141] F sema_feat (i) = ReLU(Conv′ 1×1 (RoIAlign(ReLU(Conv 1×1 (F fine )))))

[0142]

[0143] where ReLU and Sigmoid are activation functions; i = 1, 2, 3, corresponding to the three stages of the instance segmentation main branch respectively. In the i-th stage, Conv 1×1 is a 1×1 convolution with the number of input channels being C and the number of output channels being ; RoIAlign changes the feature size to , and the meaning is the same as described in step C2; Conv′ 1×1 is a 1×1 convolution with the number of input and output channels both being ; is a 1×1 convolution with the number of input channels being C and the number of output channels being 1; RoIAlign mask changes the size of the semantic segmentation mask to (P′, 1, 2 i-1 ×H′, 2 i-1 ×W′). Calculate the cross-entropy loss between the result of and the true label to obtain the loss L sema of the fine-grained feature.

[0144] Step C4: Design the instance segmentation main branch. Use the detection head feature F mask with a size of (P′, C, H′, W′) obtained in step B1, and extract the instance segmentation feature F ins_feat (1) through a lightweight segmentation head composed of 2 stacked 3×3 convolutions. The specific formula is as follows:

[0145] F ins_feat (1) = ReLU(Conv 3×3 (ReLU(Conv3×3 (F mask )))),

[0146] Among them, ReLU is the activation function; two Convs 3×3 are both 3×3 convolutions with the number of input and output channels being C, the stride being 1, and the padding being 1. Pass F ins_feat (1) through a 1×1 convolution and a Sigmoid activation function to obtain an instance segmentation mask F ins_mask (1) with a size of (P′, K, H′, W′), and the specific formula is as follows:

[0147] F ins_mask (1) = Sigmoid(Conv 1×1 (F ins_feat (1))),

[0148] Among them, Sigmoid is the activation function; Conv 1×1 is a 1×1 convolution with the number of input channels being C and the number of output channels being K, and K is the total number of categories except the background in the dataset. In the first stage, use the Enhanced Spatial Atrous Fusion Module (ESAFM) to fuse the bounding box feature F box_feat (1) obtained in step C2, the semantic segmentation feature F sema_feat (1) obtained in step C3, the semantic segmentation mask F sema_mask (1), as well as the instance segmentation feature F ins_feat (1) and the instance segmentation mask F ins_mask (1) to obtain the fused instance segmentation feature F fused_feat (1). Upsample the fused instance segmentation feature F fused_feat (1) by a factor of 2 using bilinear interpolation to obtain the instance segmentation feature F ins_feat (2) of the size as the input of the second stage.

[0149] In the i-th stage, where i = 2, 3, pass the instance segmentation feature F ins_feat (i) input in the (i - 1)-th stage through a 1×1 convolution and a Sigmoid activation function to obtain an instance segmentation mask F i-1 ×H′, 2 i-1 ×W′) with a size of (P′, k, 2 ins_mask (i). Then use the Enhanced Spatial Atrous Fusion Module (ESAFM) to fuse the five features to obtain a fused instance segmentation feature F H′, 2 i-1 ×W′) with a size of fused_feat (i), and upsample it by a factor of 2 using bilinear interpolation to obtain the instance segmentation feature F ins_feat (i + 1) input in the next stage. The specific formula is as follows:

[0150] F ins_mask (i) = Sigmoid(Conv 1×1 (F ins_mask (i))),

[0151] F fused_feat (i) = ESAFM i (F box_feat (i), F sema_feat (i), F sema_mask (i), F ins_feat (i), F ins_mask (i)),

[0152] F ins_feat (i + 1) = 2xUP(F fused_feat (i)),

[0153] where i = 2, 3; Sigmoid is the activation function; at the i-th stage, Conv 1×1 is a 1×1 convolution with the number of input channels being and the number of output channels being K (K is the total number of classes in the dataset excluding the background), and ESAFM i represents the enhanced spatial hole fusion module at different stages; 2xUP refers to bilinear interpolation upsampling by a factor of 2. Finally, F ins_feat (4) passes through a 1×1 convolution and the Sigmoid activation function to obtain an instance segmentation mask F ins_mask (4) of size (P′, K, 8×H′, 8×W′).

[0154] Step C5: Calculate the cross-entropy loss between the four instance segmentation masks F i-1 of size (P′, K, 2 i-1 ×H′, 2 ins_mask (i)), i = 1, 2, 3, 4 obtained in Step C4 and the ground truth labels to obtain the loss L ins_stage (i), i = 1, 2, 3, 4 of the instance segmentation mask at each stage; the total loss L ins of the instance segmentation mask is the sum of the losses of the four stages multiplied by weights, and the specific formula is as follows:

[0155] L ins = ω1×L ins_stage (1) + ω2×L ins_stage (2) + ω3×L ins_stage (3) + ω4×L ins_stage (4).

[0156] Figure 4 is the structural schematic diagram of the enhanced spatial hole fusion module in this embodiment. As Figure 4As shown in the figure, the implementation method of the enhanced spatial hole fusion module is as follows:

[0157] Step C41: Enhanced Spatial Hole Fusion Module ESAFM i In the i-th (i = 1, 2, 3) stage, first, select the mask of the true class from the masks predicting K classes using the ground truth label for the instance segmentation mask of size (P′, K, 2 i-1 ×H′, 2 i-1 ×W′), and change the size of the instance segmentation mask to (P′, 1, 2 i-1 ×H′, 2 i-1 ×W′). Then, after concatenating the bounding box feature F (f), semantic segmentation feature F box_feat (i), instance segmentation feature F sema_feat (i), and the semantic segmentation mask F ins_feat (i) and instance segmentation mask F i-1 ×H′, 2 i-1 ×W′) of size (P′, 1, 2 sema_mask (i), perform a 1×1 convolution on the concatenated features to obtain the preliminarily aggregated feature F ins_mask (i). The specific formula is as follows: aggr_feat_in Where Concat is the feature concatenation operation, and Conv

[0158]

[0159] is a 1×1 convolution with the number of input channels being 1×1 and the number of output channels being and respectively.

[0160] Step C42: Extract the detailed feature F aggr_feat_in (i) from the preliminarily aggregated feature F detail_feat (i) through global average pooling and 1×1 convolution. The specific formula is as follows:

[0161] F detail_feat (i) = UP(Conv 1×1 (GAP(F aggr_feat_in (i)))), i = 1, 2, 3

[0162] Where GAP is global average pooling, which changes the feature from to Conv 1×1 is a 1×1 convolution with both the number of input and output channels being UP represents bilinear interpolation upsampling, which changes the feature back to Then, the preliminarily aggregated feature F aggr_feat_in(i) Extract features with different receptive fields through three parallel dilated convolutions with different dilation rates. The specific formula is as follows:

[0163]

[0164]

[0165]

[0166] where ReLU is the activation function; and are dilated convolutions with different dilation rates, and the dilation rates are a, b, c respectively. The kernel size of both is 3×3. In different stages, the values of dilation rates a, b, c are different from each other and increase with the increase of stages, so as to extract features with larger receptive fields. Finally, add the detailed feature F detail_feat (i) element-wise to the features with different receptive fields to obtain the multi-receptive field feature F multi_feat (i). The specific formula is as follows:

[0167] F multi_feat (i) = Add(F detail_feat (i), F dila_feat_1 (i), F dila_feat_2 (i), F dila_feat_3 (i)), i = 1, 2, 3

[0168] where Add represents element-wise addition.

[0169] Step C43: Pass the preliminarily aggregated feature F aggr_feat_in (i) through the Enhanced Spatial Attention (ESA) module to obtain the enhanced spatial attention feature F en_spat_atten_feat (i). In the Enhanced Spatial Attention (ESA) module, pass the input F aggr_feat_in (i) through a 3×3 convolution to obtain the initial feature F init_feat (i). Then use max pooling to extract the spatial attention feature while reducing the number of parameters, and then pass through a convolution group composed of three 3×3 convolutions, and then upsample back to the original size to obtain the spatial depth feature F spat_deep (i). Then pass the initial feature F init_feat (i) through a 1×1 convolution and add it element-wise to the spatial depth feature to obtain the spatially enhanced feature F en_spat (i). The specific formula is as follows:

[0170] F init_feat (i) = ReLU(Conv 3×3 (F aggr_feat_in (i))),

[0171] F spat_deep(i) = UP(ConvGroup(MP(F init_feat (i)))),

[0172] F en_spat (i) = Add(F spat_deep (i) + Conv 1×1 (F init_feat (i))), i = 1, 2, 3

[0173] where ReLU is the activation function; Conv 3×3 is a 3×3 convolution with both the number of input and output channels being and a stride of 1 and a padding of 1; MP is max pooling with a pooling window of 7 and a stride of 3; UP is bilinear interpolation upsampling; Conv 1×1 is a 1×1 convolution with both the number of input and output channels being ; Add is element-wise addition; ConvGroup is a convolution group composed of three 3×3 convolutions, and the specific formula is as follows:

[0174] out = Conv 3×3 (ReLU(Conv 3×3 (ReLU(Conv 3×3 (in)))),

[0175] where ReLU is the activation function; in and out are the input and output of the convolution group ConvGroup respectively; the three Conv 3×3 are all 3×3 convolutions with both the number of input and output channels being and a stride of 1 and a padding of 1. Then, the spatially enhanced feature F en_spat (i) passes through a 1×1 convolution and the Sigmoid activation function to obtain the factor Factor i for enhancing spatial attention, and then multiplies it element-wise with the preliminarily aggregated feature F aggr_feat_in (i) to obtain the final spatially enhanced attention feature F en_spat_atten_feat (i), and the specific formula is as follows:

[0176] Factor i = Sigmoid(Conv 1×1 (F en_spat (i))),

[0177] F en_spat_atten_feat (i) = Factor i × F aggr_feat_in (i), i = 1, 2, 3

[0178] where Sigmoid is the activation function; Conv 1×1 has both the number of input and output channels being 1×1 convolution

[0179] Step C44: Multiscale receptive field feature F obtained in step C42 multi_feat (i) and the final enhanced spatial attention feature F obtained in step C43 en_spat_atten_feat (i) are concatenated, and then passed through a 1×1 convolution to change the number of channels to obtain the preliminarily fused instance segmentation feature F init_fuse_feat (i). Then, F init_fuse_feat (i) is passed through a 1×1 convolution to reduce the number of channels by 2, and then concatenated with the semantic segmentation mask F sema_mask (i) and the instance segmentation mask F ins_mask (i) to obtain the output of the enhanced spatial atrous fusion module ESAFM, the fused instance segmentation feature F fused_feat (i), and the specific formula is as follows:

[0180] F init_fuse_feat (i) = ReLU(Conv 1×1 (Concat(F multi_feat (i), F en_spat_atten_feat (i)))),

[0181] F fused_feat (i) = Concat(Conv′ 1×1 (F init_fuse_feat (i)), F sema_mask (i), F ins_mask (i)), i = 1, 2, 3

[0182] where ReLU is the activation function; Concat is feature concatenation; Conv 1×1 is a 1×1 convolution with the input number of channels being and the output number of channels being ; Conv′ 1×1 is a 1×1 convolution with the input number of channels being and the output number of channels being .

[0183] Figure 2 This is the structural diagram of the instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement in this embodiment. As Figure 2 shown, the implementation method of the instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement is as follows:

[0184] Step D1: Use the ResNet-50 backbone network as the feature extraction module to extract features from the input image, send the extracted feature map into the Feature Pyramid Network FPN to construct multi-scale features, and then send them into the RPN sub-network for region candidate proposal;

[0185] Step D2: Perform foreground and background binary classification prediction in the RPN sub-network, map the coordinates of the candidate regions predicted as foreground to the multi-scale feature representation, then use RoIAlign to pool them into features of a fixed size, and send them to the decoupled classification and regression dual-branch object detection head. After obtaining the bounding box results through the prediction of the dual-branch object detection head, map the bounding box coordinates back to the multi-scale feature representation constructed by FPN, use RoIAlign to pool them into features of a fixed size, and send them to the instance segmentation head of the multi-stage refinement and enhanced spatial hole fusion module;

[0186] Step D3: Use the bounding box features generated by the object detection head in Step C2 and the fine-grained features generated by the FPN multi-scale features in Step C3, and periodically add them to the process of the main branch of instance segmentation for mask segmentation to obtain the final instance segmentation result.

[0187] In this embodiment, training the instance segmentation network specifically includes the following steps:

[0188] Step E1: Input the images in the preprocessed training set into the instance segmentation network. After the multi-scale features extracted by the backbone network and the feature pyramid network are sent to the RPN sub-network to generate a certain number of candidate regions, classify the positive samples, i.e., foreground objects, and negative samples, i.e., background regions. Then map the coordinates of the candidate regions predicted as foreground back to the multi-scale features, and use RoIAlign to pool them into a size of H×W and send them to the object detection head; in the bounding box results predicted by the object detection head, map the coordinates back to the multi-scale features, and use RoIAlign to pool them into sizes of H′ and W′ and send them to the instance segmentation head.

[0189] Step E2: The fine-grained features generated by the P2 layer of the multi-scale features represented by FPN and the bounding box features generated by the object detection head are added to the instance segmentation head. Then use the fine-grained features to perform semantic mask segmentation, and use the object detection head and the instance segmentation head to perform bounding box detection and mask segmentation;

[0190] Step E3: Calculate the loss L of the RPN sub-network rpn , the classification loss L corresponding to Step B2 cls and the regression loss L corresponding to Step B3 reg , the loss L of the fine-grained features corresponding to Step C3 sema , and the total instance segmentation mask loss L corresponding to Step C5 ins . The total loss L all is the sum of the five losses multiplied by weights. The specific formula is as follows:

[0191] L all = λ1×L rpn + θ2×Lcls + θ3 × L reg + θ4 × L sema + λ5 × L ins , where θ1, θ2, λ3, θ4, and θ5 are hyperparameters. Then, the backpropagation method is used to calculate the gradients of the parameters in the instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement, and the parameters of the instance segmentation network are updated using the stochastic gradient descent method.

[0192] In this embodiment, the image to be processed is processed as follows:

[0193] Step F1: Input the image without label information into the trained instance segmentation network for processing;

[0194] Step F2: Use the object detection head to predict the bounding boxes of the foreground objects in the image, and use the instance segmentation head to predict the mask results of each instance in the image; the result obtained by the instance segmentation head in the network is the final instance segmentation result.

[0195] This embodiment also provides an instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the above method steps can be implemented.

[0196] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0197] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks

[0198] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0199] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0200] This patent is not limited to the above-described best mode. Anyone inspired by this patent can derive various other forms of instance segmentation methods based on detection enhancement and multi-stage bounding box feature refinement. All equivalent changes and modifications made in accordance with the scope of the patent application of the present invention shall fall within the scope of this patent.

Claims

1. An instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement, characterized in that It includes the following steps: Step A: Perform data preprocessing on the images in the training set, including data augmentation and normalization; Step B: Construct a dual-branch object detection head that decouples classification and regression; Step C: Construct an instance segmentation head with a multi-stage refinement and enhanced spatial hole fusion module, and use the bounding box features of the object detection head to enhance the refinement process of the instance segmentation head in multiple stages; Step D: Construct an instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement; Step E: Use the images in the training set to train the instance segmentation network, generate instance segmentation results and calculate the loss function, and backpropagate to optimize the parameters of the entire network to obtain a trained instance segmentation network; Step F: Input the image to be processed into the trained instance segmentation network to obtain the instance segmentation result.

2. The instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to claim 1, characterized in that: Step A specifically includes the following steps: Step A1: Perform scale transformation on the images in the training set. Without changing the aspect ratio, set the threshold of the image length and width to 2048; that is, perform scale transformation on the image according to the long side of the image and the threshold to ensure that neither the long side nor the short side exceeds the threshold; then randomly flip all the images after scale transformation to achieve data augmentation; Step A2: Perform normalization processing on the enhanced images. The mean of normalization is [123.675, 116.28, 103.53], and the variance is [58.395, 57.12, 57.375]; finally, pad the images so that the length and width are divisible by 32; each image has a corresponding label, and the label content is the bounding box and mask of each instance object in the image. The image labels are also processed synchronously during image scale transformation, data augmentation, and padding.

3. The instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to claim 1, characterized in that: In step B, the implementation method of the dual-branch object detection head that decouples classification and regression is as follows: Step B1: Extract features from the input image through the backbone network of the instance segmentation network, and use the Feature Pyramid Network (FPN) to construct multi-scale features P2 - P6, where P2 has the largest resolution and P6 has the smallest resolution. Then, feed them into the Region Proposal Network (RPN) sub-network for region candidate proposals; perform background and foreground binary classification predictions in the RPN sub-network, map the coordinates of the candidate regions predicted as foreground back to the corresponding regions in the multi-scale features obtained by FPN, and use RoIAlign to pool the features of the corresponding regions into detection head features F of a fixed size of (P, C, H, W). Bbox Among them, P is the number randomly extracted from the candidate region results predicted as foreground in the RPN sub-network, C represents the number of channels of the features, and H and W represent the height and width of the detection head features; then the obtained detection head features F Bbox are input into the decoupled classification and regression dual-branch object detection head. Step B2: For the detection head feature F obtained in Step B1 Bbox , flatten it into a feature F' of size (P, C×H×W); taking advantage of the fact that fully connected layers are more spatially sensitive than convolutional layers, in the classification branch, the prediction of classification labels is calculated only using fully connected layers; pass the flattened feature F Bbox through three fully connected layers to predict the probability Class_Score of each class, including the background. The specific formula is as follows: Bbox ′ ​ Class_Score = FC3(ReLU(FC2(ReLU(FC1(F Bbox ′))))), Among them, ReLU is the activation function; FC1 is a fully connected layer with an input channel number of C×H×W and an output channel number of 4×C; FC2 is a fully connected layer with both input and output channel numbers of 4×C; FC3 is a fully connected layer with an input channel number of 4×C and an output channel number of K+1, where K is the total number of categories except the background in the dataset; then, for each predicted category, including the probability Class_Score of the background, cross-entropy loss calculation is performed with the true label to obtain the loss L of the predicted classification cls ; Step B3: Taking advantage of the feature that the convolutional layer is more conducive to bounding box regression and localization, in the regression branch, use multiple convolutional layers to calculate the regression task of the bounding box; the detection head feature F of size (P, C, H, W) obtained in Step B1 Bbox is passed through a residual block to increase the dimension to obtain a residual feature F res , with a size of (P, 4×C, H, W); the specific formula is as follows: F res = ReLU(Conv 1×1 (F Bbox ) + Conv 1×1 (Conv 3×3 (F Bbox ))), Among them, ReLU is the activation function; Conv 3×3 is a 3×3 convolution with the number of input and output channels both being C, a stride of 1, and a padding of 1; two Convs 1×1 are both 1×1 convolutions with the number of input channels being C and the number of output channels being 4×C; then the residual feature F res passes through four bottleneck layers to obtain the bottleneck layer feature F bott (i), and the specific formula is as follows: Among them, BottleNeck i is the bottleneck layer, where i = 1, 2, 3, 4; the specific formula for any bottleneck layer BottleNeck is as follows: out = ReLU(in + Conv 1×1 (Conv 3×3 (Conv 1×1 (in)))), Among them, ReLU is the activation function; in and out are the input and output of any bottleneck layer respectively; Conv 1×1 is a 1×1 convolution with an input channel number of 4×C and an output channel number of C; Conv 3×3 is a 3×3 convolution with an input and output channel number of C, a stride of 1, and a padding of 1; Conv′ 1×1 is a 1×1 convolution with an input channel number of C and an output channel number of 4×C; then the features F of the last bottleneck layer bott (4) pass through average pooling and a fully connected layer to predict the four coordinates of the bounding box Bbox_Pred, and the specific formula is as follows: Bbox_Pred = FC(View(AP(F bottle (4)))), Among them, AP is average pooling with a pooling window of H×W, which changes the bottleneck layer feature size from (P, 4×C, H, W) to (P, 4×C, 1, 1); View is a feature flattening operation that changes the feature from (P, 4×C, 1, 1) to (P, 4×C×1×1); FC is a fully connected layer with 4×C input channels and 4×K output channels, where K is the total number of classes in the dataset excluding the background; finally, a smooth L1 loss calculation is performed between the predicted bounding box Bbox_Pred and the ground truth label to obtain the loss L of bounding box regression reg ; After the decoupled classification and regression double-branch object detection head designed in steps B2 and B3, the predicted class and bounding box results are obtained; the predicted bounding box coordinates are mapped back to the corresponding regions in the multi-scale features obtained by FPN, and the features of the corresponding regions are pooled into the segmentation head feature F with a fixed size of (P', C, H', W′) using RoIAlign mask ; where P′ is the number of bounding boxes predicted by the object detection head, and H′ and W′ are the height and width of the segmentation head feature; then the obtained segmentation head feature F mask is input into the instance segmentation head.

4. The instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to claim 3, characterized in that: In step C, the implementation method of the instance segmentation head with a multi-stage refinement and enhanced spatial hole fusion module is as follows: Step C1: The instance segmentation head of the multi-stage refinement and enhanced spatial hole fusion module consists of three parts: a bounding box feature branch, a fine-grained feature branch, and an instance segmentation main branch; Step C2: Design the bounding box feature branch and use the boundary features to enhance the refinement process of the instance segmentation main branch; for the bottleneck layer feature F of size (P, 4×C, H, W) obtained in step B3 bott (i), where i = 1, 2, 3, use RoIAlign pooling and convolutional layers to obtain the bounding box feature F box_feat (i) that matches the size of the instance segmentation main branch designed in step C3. The specific formula is as follows: F box_feat (i) = ReLU(Conv′ 1×1 (RoIAlign(ReLU(Conv 1×1 (F bott (i)))))), Among them, ReLU is the activation function; i = 1, 2, 3, corresponding to the three stages of the instance segmentation main branch respectively; in the i-th stage, Conv 1×1 is a 1×1 convolution with 4×C input channels and output channels; RoIAlign changes the feature size from pooling to where P' is the number of bounding boxes of the object detection prediction result, represents the number of channels of the feature, H' and W' are the width and height of the feature, which are consistent with the width and height of the segmentation head feature F mask described in step B1; Conv' 1×1 is a 1×1 convolution with both input and output channels being ; the obtained bounding box feature F box_feat (i) is added to the instance segmentation main branch stage by stage to enhance the refinement process of the instance segmentation head; Step C3: Design a fine-grained feature branch to enhance the refinement process of the instance segmentation main branch using fine-grained features; the feature F of layer P2 in the multi-scale features obtained by the Feature Pyramid Network (FPN) P2 , and the fine-grained feature F fine is extracted through a lightweight semantic segmentation head composed of 4 stacked 3×3 convolutions. The specific formula is as follows: Among them, ReLU is the activation function; the four Convs 3×3 are all 3×3 convolutions with the number of input and output channels both being C, a stride of 1, and a padding of 1; the fine-grained feature F fine has a size of where and represent the height and width of the feature, which are one-fourth of the height and width of the input image, and are the same size as the multi-scale feature F P2 ; then, the RoIAlign pooling and the convolutional layer are used to obtain the semantic segmentation feature F sema_feat (i) and the semantic segmentation mask F sema_mask (i), and the specific formula is as follows: F sema_feat (i) = ReLU(Conv′ 1×1 (RoIAlign(ReLU(Conv 1×1 (F fine ))))), Among them, ReLU and Sigmoid are activation functions; i = 1, 2, 3, corresponding to the three stages of the instance segmentation main branch respectively; in the i-th stage, Conv 1×1 is a 1×1 convolution with an input channel number of C and an output channel number of ; RoIAlign changes the feature size to Conv′ 1×1 is a 1×1 convolution with both input and output channel numbers of ; is a 1×1 convolution with an input channel number of C and an output channel number of 1; RoIAlign mask changes the semantic segmentation mask size to (P', 1, 2 i-1 ×H′, 2 i-1 ×W′); calculates the cross-entropy loss between the result of and the ground truth label to obtain the loss L sema ; Step C4: Design the main branch of instance segmentation; use the detection head feature F of size (P′, C, H', W′) obtained in Step B1 mask , and extract the instance segmentation feature F through a lightweight segmentation head composed of 2 stacked 3×3 convolutions ins_feat (1), and the specific formula is as follows: F ins_feat (1) = ReLU(Conv 3×3 (ReLU(Conv 3×3 (F mask )))), where ReLU is the activation function; two Convs 3×3 are both 3×3 convolutions with the number of input and output channels being C, a stride of 1, and a padding of 1; F ins_feat (1) obtains the instance segmentation mask F ins_mask (1) of size (P', K, H', W′) through a 1×1 convolution and the Sigmoid activation function. The specific formula is as follows: F ins_mask (1) = Sigmoid(Conv 1×1 (F ins_feat (1))), Among them, Sigmoid is the activation function; Conv 1×1 is a 1×1 convolution with C input channels and K output channels, where K is the total number of categories in the dataset except the background; in the first stage, the enhanced spatial atrous fusion module ESAFM is used to fuse the bounding box feature F box_feat (1) obtained in step C2, the semantic segmentation feature F sema_feat (1) obtained in step C3, and the semantic segmentation mask F sema_mask (1), as well as the instance segmentation feature F ins_feat (1) and the instance segmentation mask F ins_mask (1) to obtain the fused instance segmentation feature F fused_feat (1); the fused instance segmentation feature F fused_feat (1) is upsampled by a factor of 2 using bilinear interpolation to obtain the instance segmentation feature F ins_feat (2) as the input for the second stage; In the i-th, where i = 2, 3, stage, the instance segmentation feature F input in the (i - 1)-th stage is ins_feat (i) obtained through a 1×1 convolution and a Sigmoid activation function to get an instance segmentation mask F i-1 of size (P′, K, 2 i-1 ×H′, 2 ins_mask ×W′); then the enhanced spatial atrous fusion module ESAFM is used to fuse the five features to obtain a fused instance segmentation feature F of size fused_feat , which is upsampled by a factor of 2 using bilinear interpolation to obtain the instance segmentation feature F ins_feat (i + 1) input in the next stage. The specific formula is as follows: F ins_mask (i) = Sigmoid(Conv 1×1 (F ins_mask (i))), F fused_feat (i) = ESAFM i (F box_feat (i), F sema_feat (i), F sema_mask (i), F ins_feat (i), F ins_mask (i)), F ins_feat (i + 1)= 2xUP(F fused_feat (i)), where \(i = 2, 3\); Sigmoid is the activation function; in the \(i\)-th stage, Conv 1×1 is a \(1\times1\) convolution with the number of input channels being and the number of output channels being \(K\), and ESAFM i represents the enhanced spatial atrous fusion module at different stages; \(2xUP\) refers to bilinear interpolation upsampling by a factor of 2; finally, \(F\) ins_feat (4) passes through a \(1\times1\) convolution and the Sigmoid activation function to obtain an instance segmentation mask \(F\) ins_mask (4); Step C5: Cross-entropy loss calculation is performed on the four instance segmentation masks F i-1 (i) of size (P′, K, 2 i-1 ×H′, 2 ins_mask ×W′) obtained in Step C4 and the ground truth labels, and the loss L ins_stage (i) of the instance segmentation mask at each stage is obtained, where i = 1, 2, 3, 4; the total loss L ins of the instance segmentation mask is the sum of the losses of the four stages multiplied by weights, and the specific formula is as follows: L ins = ω1 × L ins_stage (1) + ω2 × L ins_stage (2) + ω3 × L ins_stage (3) + ω4 × L ins_stage (4).

5. The instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to claim 4, wherein: In step C4, the implementation method of the enhanced spatial hole fusion module is as follows: Step C41: Enhanced spatial void fusion module ESAFM i In the i,i=1,2,3th stage, firstly, the size of (P′,K,2 i-1 ×H′,2 i-1 The instance segmentation mask of (P′,1,2 i-1 ×H',2 i-1 ×W′); then the size is The bounding box feature F box_feat (i), semantic segmentation feature F sema_feat (i), instance segmentation feature F ins_feat (i), and the size is (P′,1,2 i-1 ×H′,2 i-1 ×W′) semantic segmentation mask F sema_mask (i) and instance segmentation mask F ins_mask (i) After feature concatenation, a 1×1 convolution is performed to obtain the initial aggregated feature F aggr_feat_in (i), the specific formula is as follows: Among them, Concat is the feature concatenation operation, and Conv 1×1 is a 1×1 convolution with an input channel number of and an output channel number of . Step C42: The preliminarily aggregated feature F aggr_feat_in (i) Extract the detailed feature F through global average pooling and 1×1 convolution detail_feat (i), and the specific formula is as follows: F detail_feat (i) = UP(Conv 1×1 (GAP(F aggr_feat_in (i))), i = 1, 2, 3 Among them, GAP is global average pooling, which transforms the feature from to Conv 1×1 is a 1×1 convolution with the number of input and output channels both being UP represents bilinear interpolation upsampling, which transforms the feature back to Then, the preliminarily aggregated feature F aggr_feat_in (i) extracts features with different receptive fields through three parallel dilated convolutions with different dilation rates. The specific formula is as follows: where ReLU is the activation function; and are dilated convolutions with different dilation rates, and the dilation rates are a, b, c respectively, and the kernel size is 3×3; in different stages, the values of the dilation rates a, b, c are different from each other and increase with the increase of the stage, so as to extract features with a larger receptive field; finally, the detailed feature F detail_feat (i) is added element by element to the features with different receptive fields to obtain the multi-receptive field feature F multi_feat (i), and the specific formula is as follows: F multi_feat (i) = Add(F detail_feat (i), F dila_feat_1 (i), F dila_feat_2 (i), F dila_feat_3 (i)), i = 1, 2, 3 Among them, Add represents element-wise addition; Step C43: The preliminarily aggregated feature F aggr_feat_in (i) is passed through an Enhanced Spatial Attention (ESA) module to obtain an enhanced spatial attention feature F en_spat_atten_feat (i); in the ESA module, the input F aggr_feat_in (i) is convolved with a 3×3 convolutional layer to obtain an initial feature F init_feat (i); then, max pooling is used to extract the spatial attention feature while reducing the number of parameters, followed by a convolutional group composed of three 3×3 convolutional layers, and then upsampled back to the original size to obtain a spatial depth feature F spat_deep (i); then, the initial feature F init_feat (i) is added element-wise to the spatial depth feature through a 1×1 convolutional layer to obtain a spatially enhanced feature F en_spat (i); the specific formula is as follows: F init_feat (i) = ReLU(Conv 3×3 (F aggr_feat_in (i))), F spat_deep (i) = UP(ConvGroup(MP(F init_feat (i)))), F en_spat (i) = Add(F spat_deep (i) + Conv 1×1 (F init_feat (i))), i = 1, 2, 3 where ReLU is the activation function; Conv 3×3 has the same number of input and output channels as a 3×3 convolution with a stride of 1 and a padding of 1; MP is a max pooling with a pooling window of 7 and a stride of 3; UP is a bilinear interpolation upsampling; Conv 1×1 has the same number of input and output channels as a 1×1 convolution; Add is an element-wise addition; ConvGroup is a convolution group composed of three 3×3 convolutions, and the specific formula is as follows: out = Conv 3×3 (ReLU(Conv 3×3 (ReLU(Conv 3×3 (in))))), where ReLU is the activation function; in and out are the input and output of the convolutional group ConvGroup respectively; the three Convs 3×3 both have an input and output channel number of 3×3 convolutions with a stride of 1 and a padding of 1; then the spatially enhanced feature F en_spat (i) passes through a 1×1 convolution and the Sigmoid activation function to obtain the factor Factor for enhancing spatial attention i , and then multiplies it element-wise with the preliminarily aggregated feature F aggr_feat_in (i) to obtain the final spatially enhanced attention feature F en_spat_atten_feat (i), and the specific formula is as follows: Factor i = Sigmoid(Conv 1×1 (F en_spat (i))), F en_spat_atten_feat (i) = Factor i ×F aggr_feat_in (i), i = 1, 2, 3 Among them, Sigmoid is the activation function; Conv 1×1 is a 1×1 convolution with the number of input and output channels both being ; Step C44: The multi-receptive field feature F obtained in step C42 multi_feat (i) and the final enhanced spatial attention feature F obtained in step C43 en_spat_atten__feat (i) are concatenated, and then the number of channels is changed through a 1×1 convolution to obtain the preliminarily fused instance segmentation feature F init_fuse_feat (i); then F init_fuse_feat (i) is passed through a 1×1 convolution to reduce the number of channels by 2, and then concatenated with the semantic segmentation mask F sema_mask (i) and the instance segmentation mask F ins_mask (i) to obtain the output of the enhanced spatial hole fusion module ESAFM, and the fused instance segmentation feature F fused_feat (i), and the specific formula is as follows: F init_fuse_feat (i) = ReLU(Conv 1×1 (Concat(F multi_feat (i), F en_spat_atten__feat (i)))), F fused_feat (i) = Concat(Conv′ 1×1 (F init_fuse_feat (i)), F sema_mask (i), F ins_mask (i)), i = 1, 2, 3 Among them, ReLU is the activation function; Concat is the feature concatenation; Conv 1×1 is a 1×1 convolution with an input channel number of and an output channel number of ; Conv′ 1×1 is a 1×1 convolution with an input channel number of and an output channel number of .

6. The instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to claim 5, characterized in that: In step D, the implementation method of the instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement is: Step D1: Use the ResNet-50 backbone network as the feature extraction module to extract features from the input image, send the extracted feature map into the Feature Pyramid Network (FPN) to construct multi-scale features, and then send them into the Region Proposal Network (RPN) sub-network for region candidate proposal; Step D2: Perform background and foreground binary classification prediction in the RPN sub-network, map the coordinates of the candidate regions predicted as foreground to the multi-scale feature representation, then use RoIAlign to pool into features of a fixed size, and send them to the decoupled classification and regression dual-branch object detection head; after the bounding box results are predicted by the dual-branch object detection head, map the bounding box coordinates back to the multi-scale feature representation constructed by FPN, use RoIAlign to pool into features of a fixed size and send them to the instance segmentation head of the multi-stage refinement and enhanced spatial hole fusion module; Step D3: Use the bounding box features generated by the object detection head in Step C2 and the fine-grained features generated by the FPN multi-scale features in Step C3, and add them to the process of the instance segmentation main branch stage by stage for mask segmentation to obtain the final instance segmentation result.

7. The instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to claim 6, wherein Step E specifically includes the following steps: Step E1: Input the images in the preprocessed training set into the instance segmentation network, send the multi-scale features extracted by the backbone network and the feature pyramid network to the RPN sub-network to generate a given number of candidate regions, classify the positive samples, i.e., foreground objects, and negative samples, i.e., background regions, then map the coordinates of the candidate regions predicted as foreground back to the multi-scale features, and then use RoIAlign to pool into a size of H×W and send them to the object detection head; in the bounding box results predicted by the object detection head, map the coordinates back to the multi-scale features, and then use RoIAlign to pool into sizes of H′ and W′ and send them to the instance segmentation head; Step E2: Add the fine-grained features generated by the P2 layer of the multi-scale features represented by FPN and the bounding box features generated by the object detection head to the instance segmentation head; then use the fine-grained features to perform semantic mask segmentation, and use the object detection head and the instance segmentation head to perform bounding box detection and mask segmentation; Step E3: Calculate the loss L of the RPN sub-network rpn , the classification loss L corresponding to step B2 cls and the regression loss L corresponding to step B3 reg , the loss L of the fine-grained features corresponding to step C3 sema , and the total loss L of the instance segmentation mask corresponding to step C5 ins ; The total loss L all is the sum of the five losses multiplied by their weights. The specific formula is as follows: L all = θ1 × L rpn + θ2 × L cls + θ3 × L reg + θ4 × L sema + θ5 × L ins , where θ1, θ2, θ3, θ4, and θ5 are hyperparameters; then use the backpropagation method to calculate the gradients of the parameters in the instance segmentation network based on detection enhancement and multi-stage bounding box feature refinement, and use the stochastic gradient descent method to update the parameters of the instance segmentation network.

8. The instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to claim 1, wherein Step F specifically includes the following steps: Step F1: Input the image without label information into the trained instance segmentation network for processing; Step F2: Use the object detection head to predict the bounding boxes of the foreground objects in the image, and use the instance segmentation head to predict the mask results of each instance in the image; the results obtained by the instance segmentation head in the network are the final instance segmentation results.

9. An instance segmentation device based on detection enhancement and multi-stage bounding box feature refinement, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that when the processor executes the computer program, it implements the instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to any one of claims 1-8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement according to any one of claims 1-8.

Citation Information

Patent Citations

  • Instance segmentation method and system based on multi-scale features and context attention

    CN114693930A

  • Semantic segmentation-based image composite defect detection method and system

    CN114820579A