A Real-time Instance Segmentation Method Integrating Sparse Framework and Spatial Attention

Through the real-time instance segmentation method that integrates sparse framework and spatial attention, the problem of long inference time in the prior art is solved, and fast and accurate instance segmentation is achieved, which is suitable for real-time video processing.

CN115100410BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210803057.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2025-06-13
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

Existing instance segmentation methods rely on intensive suggestion boxes and complex post-processing operations, resulting in long and inefficient inference.

Method used

Using a real-time instance segmentation method that integrates sparse framework and spatial attention, the instance segmentation network is built, including feature extraction network, feature enhancement network, mask branch and instance branch, to avoid complex operations such as prior box design and non-maximum suppression, and improve the network's attention to instance features.

Benefits of technology

It realizes fast and accurate instance segmentation, reduces inference time, avoids complex post-processing operations, and is suitable for real-time video frame processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100410B_ABST
    Figure CN115100410B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time instance segmentation method integrating a sparse framework and spatial attention. First, an image to be processed is obtained and input into a feature extraction network to extract multi-scale feature maps. The multi-scale feature maps are input into a feature enhancement network to obtain enhanced feature maps. Then, the enhanced feature maps are input into an instance branch to obtain object bounding boxes and object classifications. After concatenating the enhanced feature maps with the object bounding boxes output by the instance branch, they are input into a mask branch. First, a convolution operation is performed, and then they pass through a spatial attention module and a mask kernel generation module respectively to obtain spatial attention features and mask kernels. The spatial attention features and the mask kernels are multiplied to obtain a segmentation mask. Finally, the segmentation mask, the object bounding boxes, and the object classifications are mapped onto the image to be processed to obtain an instance segmentation result. The present invention improves the speed and accuracy of the instance segmentation task. When the input is a continuous video frame, real-time and accurate segmentation results can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of instance segmentation, and particularly relates to a real-time instance segmentation method integrating a sparse framework and spatial attention. Background Art

[0002] Instance segmentation has become one of the more important, complex, and challenging fields in machine vision research. It is helpful for many downstream tasks, including crowd detection, autonomous driving, video surveillance, and so on.

[0003] The purpose of object detection is to detect the category of image objects and give the position of the classified image objects in the form of a bounding box or center; the purpose of semantic segmentation is to obtain accurate inference results by predicting the labels of each pixel point in the image, and each pixel is classified and marked according to the object or region it belongs to. Instance segmentation segments the objects in the image (generates masks) and provides different labels for different object instances belonging to the same object class. Therefore, instance segmentation can be defined as solving the task of object detection while solving semantic segmentation, and decomposing each segmented object into its respective sub-components.

[0004] A classic and effective idea for instance segmentation in the prior art is based on the proposal generation algorithm. A proposal is used to detect an object, and then the proposal is segmented. Currently, the three common proposal detectors are dense detectors, dense-sparse detectors, and sparse detectors. Common instance segmentation algorithm models include Mask RCNN, Polar Mask, and SparseRCNN.

[0005] Mask RCNN is a typical top-down detection method. A new branch is added at the end of the dense-sparse detector Faster RCNN to segment the features after ROI Align (a feature alignment method). This method follows the paradigm of detecting first and then segmenting. The model generates dense prior object boxes through RPN, and then uses a mask segmentation head to segment the foreground and background of the detected object boxes. This approach can effectively eliminate the influence of the background on segmentation, but this approach is sensitive to the performance of the detector and requires good detection performance in the first step.

[0006] Polar Mask is an improvement on the dense detector FCOS, unifying instance segmentation into a form of a fully convolutional network, that is, without using non-convolutional operations such as ROI Align. Polar Mask models the mask from a binary image as a graph composed of several rays in the polar coordinate system, thus realizing a single-stage instance segmentation framework.

[0007] The biggest feature of Sparse RCNN is the sparsity throughout the object detection process. The input is a set of sparse proposal boxes, proposal features, and one-to-one interaction. There are no dense proposal regions and dense (global) features in the entire detection process, avoiding a large number of prior box designs and the many-to-one mapping between prior boxes and ground truth boxes. The effect is better than that of dense proposal box generation algorithms, and no post-processing operations such as non-maximum suppression are required, realizing a completely end-to-end object detector.

[0008] Most of the previous object detectors are dense detectors, such as Mask RCNN, which is based on dense proposals, pre-set on the image grid or feature map grid in advance, and then predicts the scores and offsets of these proposals, judges through IOU, and then filters through NMS; a small part are dense-sparse detectors, such as Polar Mask, which first extracts relatively few (sparse) foreground boxes, i.e., region candidate boxes, from the dense proposal regions, and then classifies and regresses the positions for each region candidate box, eliminating from thousands of candidates to a few foregrounds. For the above two methods, since each candidate box has to go through the convolutional neural network alone, and there are also cumbersome prior box designs and post-processing operations, this makes the time spent very long. Summary of the Invention

[0009] The purpose of this application is to provide a real-time instance segmentation method that fuses a sparse framework and spatial attention, avoiding a large number of complex post-processing operations such as prior box design and non-maximum suppression, and realizing fast inference.

[0010] To achieve the above purpose, the technical solution of this application is as follows:

[0011] A real-time instance segmentation method that fuses a sparse framework and spatial attention uses a constructed instance segmentation network to perform instance segmentation on an image. The instance segmentation network includes a feature extraction network, a feature enhancement network, a mask branch, and an instance branch. The real-time instance segmentation method that fuses a sparse framework and spatial attention includes:

[0012] Obtain the image to be processed and input it into the feature extraction network to extract multi-scale feature maps;

[0013] Input the multi-scale feature maps into the feature enhancement network to obtain enhanced feature maps;

[0014] Input the enhanced feature maps into the instance branch to obtain object boxes and object classifications;

[0015] The enhanced feature map is concatenated with the target boxes output by the instance branch and then input into the mask branch. First, it undergoes a convolution operation, and then passes through a spatial attention module and a mask kernel generation module respectively to obtain the spatial attention feature and the mask kernel. The spatial attention feature and the mask kernel are multiplied to obtain the segmentation mask;

[0016] The segmentation mask, target boxes, and target classifications are mapped onto the image to be processed to obtain the instance segmentation result.

[0017] Further, the feature extraction network adopts ResNet50, and the outputs of residual modules 3, 4, and 5 in the ResNet50 are used as the extracted third-scale feature map, second-scale feature, and first-scale feature map;

[0018] The input of the multi-scale feature maps into the feature enhancement network to obtain the enhanced feature map includes:

[0019] The first-scale feature map is input into the pyramid pooling module, and the first feature map is output;

[0020] The second-scale feature map is element-wise added to the first feature map to obtain the second feature map;

[0021] The third-scale feature map is element-wise added to the second feature map to obtain the third feature map;

[0022] The first feature map, second feature map, and third feature map are respectively subjected to convolution operations and then concatenated to obtain the enhanced feature map.

[0023] Further, the input of the enhanced feature map into the instance branch to obtain the target boxes and target classifications includes:

[0024] Initialize the proposal boxes and proposal features;

[0025] The enhanced feature map is used to extract the region of interest features through region feature focusing;

[0026] The region of interest features and the proposal features are input into the dynamic convolution head to generate the final target boxes and target classifications.

[0027] Further, the spatial attention module performs the following operations:

[0028] The input feature map undergoes max pooling and average pooling in the channel dimension to obtain the pooled feature P max and P avg , and then P max and P avg are dot-producted, passed through a 3×3 convolution, and then sigmoid-operated and multiplied with the original input feature map to obtain the spatial attention feature.

[0029] Further, the mask kernel generation module performs the following operations:

[0030] Adjust the dimension of the input feature map through a linear layer to generate a 128-dimensional mask kernel.

[0031] Further, when training the instance segmentation network, inputting the enhanced feature map into the instance branch to obtain the target box and target classification further includes:

[0032] Introduce the intersection over union (IoU) to adjust the target classification probability. The formula is as follows:

[0033]

[0034] Where, is the adjusted target classification probability, S i represents the intersection over union (IoU), and P i represents the target classification probability corresponding to the target box i.

[0035] Further, when training the instance segmentation network, the overall loss function of the network is as follows:

[0036]

[0037] Where, represents the target classification loss, represents the mask loss, represents the bounding box loss, and λ cls is the weight coefficient;

[0038]

[0039] Where, and are the dice loss and pixel-level binary cross-entropy loss, and λ dice and λ pix are the corresponding weight coefficients.

[0040] This application proposes a real-time instance segmentation method that fuses a sparse framework and spatial attention, improves the existing feature extraction and enhancement methods, and enhances the efficiency of this stage. Based on a pure sparse object detection algorithm, it avoids a large number of prior box designs and the many-to-one mapping between prior boxes and ground truth boxes, and does not require complex post-processing operations such as non-maximum suppression. By integrating a spatial attention module, it enhances the network's attention to instance features. This application improves the speed and accuracy of the instance segmentation task and can obtain real-time and accurate segmentation results when the input is continuous video frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the flowchart of the instance segmentation method of this application;

[0042] Figure 2 Schematic diagram of the instance segmentation network structure for this application

[0043] Figure 3 Schematic diagram of the pyramid pooling module

[0044] Figure 4 Schematic diagram of the dynamic convolution head

[0045] Figure 5 Schematic diagram of the spatial attention module Detailed implementation manners

[0046] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0047] In one embodiment, as Figure 1 shown, a real-time instance segmentation method integrating a sparse framework and spatial attention is proposed. An instance segmentation network is used to perform instance segmentation on an image. It is characterized in that the instance segmentation network includes a feature extraction network, a feature enhancement network, a mask branch and an instance branch. The real-time instance segmentation method integrating the sparse framework and spatial attention includes:

[0048] Step S1: Obtain the image to be processed and input it into the feature extraction network to extract multi-scale feature maps.

[0049] In this embodiment, given an image to be processed Image∈R 3×H×W , where 3 represents the RGB channels, and H and W represent the height and width of the image, and preprocessing is performed.

[0050] In the training of the network model and actual applications, the preprocessing methods for images are different. During the training process, in order to enhance the generalization performance of the model, data augmentation of the image is required. First, the image is flipped horizontally with a probability of 0.5, so that the data volume of the training set is doubled. Then, the image is randomly cropped and then scaled to a certain size. Since the pixel range of the input image is from 0 to 255, training within this range is unstable, and the pixel values of the image need to be scaled proportionally to 0 to 1. Therefore, the preprocessing needs to further normalize the pixel values of the image.

[0051] However, in actual applications after the network model has been trained, there is no need to perform data augmentation on the image, and only the inference input consistent with that during training needs to be maintained. Specifically, only the image to be processed needs to be scaled and normalized (no random flipping and cropping are required).

[0052] In this embodiment, ResNet50 is used as the feature extraction network (also known as the backbone), and the preprocessed image is input into ResNet50 to extract image features. The feature extraction network extracts multi-scale feature maps (res3, res4, res5), and its output is a series of feature maps of different sizes. Among them, are {512, 1024, 2048}, corresponding to 1 / 8, 1 / 16, 1 / 32 of the original image height. Similarly, it corresponds to the width of the original image, and the ratio is the same as the height. Here, the obtained different feature maps are used as the input for the next part.

[0053] Step S2: Input the multi-scale feature maps into the feature enhancement network to obtain enhanced feature maps.

[0054] In this step, the multi-scale feature maps are input into the feature enhancement network. The feature enhancement network is also called the neck network to obtain enhanced feature maps.

[0055] In a specific embodiment, the feature enhancement network of this application performs the following operations:

[0056] Input the first-scale feature map into the pyramid pooling module to output the first feature map;

[0057] Add the second-scale feature map and the first feature map element-wise to obtain the second feature map;

[0058] Add the third-scale feature map and the second feature map element-wise to obtain the third feature map;

[0059] The first feature map, the second feature map, and the third feature map are respectively convolved and then concatenated to obtain enhanced feature maps.

[0060] As Figure 2 shown, the input image to be processed passes through ResNet50. The outputs res3, res4, and res5 of the residual modules 3, 4, and 5 in ResNet50 are respectively called the third-scale feature map, the second-scale feature map, and the first-scale feature map; then they are input into the feature enhancement network (neck). In the feature enhancement network, res5 is input into the pyramid pooling module (PPM) to obtain the first feature map, res4 is added to the first feature map to obtain the second feature map, and res3 is added to the second feature map to obtain the third feature map; then the first feature map, the second feature map, and the third feature map are respectively convolved and then concatenated to obtain enhanced feature maps.

[0061] Among them, res5 is input into the Pyramid Pooling Module (PPM) to obtain the first feature map, which can enlarge the receptive field. The Pyramid Pooling Module is as Figure 3 shown. The input features of the Pyramid Pooling Module pass through 4 layers of pyramid pooling modules (pool) from the input side to the output side. The sizes of each layer are 1×1, 2×2, 3×3, and 6×6 respectively. The feature maps are pooled to the target size respectively, and then 1×1 convolution (conv) is performed on the pooled results to reduce the number of channels to 1 / N of the original, where N is 4. Then, for each feature map in the previous step, bilinear interpolation upsampling (upsample) is used to obtain the same size as the original feature map, and then the original feature map and the feature map obtained by upsampling are concatenated (concat) along the channel dimension. The number of channels obtained is twice the number of channels of the original feature map, and finally 1×1 convolution is used to reduce the channels to the original number of channels. The final feature map has the same size and channels as the original feature map.

[0062] Step S3: Input the enhanced feature map into the instance branch to obtain the target bounding box and target classification.

[0063] The instance branch of this embodiment includes the following specific operations:

[0064] Step S31: Initialize the proposal boxes and proposal features.

[0065] Before the enhanced feature map is input into the Instance Branch, a set of learnable parameters N×4 is initialized to represent the proposal boxes init_bboxes, where N is a hyperparameter representing the number of initial proposal boxes. At the same time, a set of learnable proposal features corresponding to the proposal boxes is initialized, with a size of N×d. That is, N proposal boxes and their corresponding N proposal features are initialized.

[0066] Step S32: Extract the region of interest features from the enhanced feature map through region feature focusing.

[0067] In this embodiment, the enhanced feature map P is used to extract the region of interest features roi_features through Region of Interest Align (ROI Align). Since N proposal boxes are initialized, N region of interest features can be obtained. ROI Align is a relatively mature technology in this field and will not be elaborated here.

[0068] Step S33: Input the region of interest features and the proposal features into the dynamic convolution head to generate the final target bounding box and target classification.

[0069] The dynamic convolution head (Dynamic instance interactive head) of this embodiment is as Figure 4As shown, k successive iterations of the dynamic convolutional layers are performed, where k is a hyperparameter, i.e., the preset number of layers of dynamic convolutional iterations.

[0070] The input to the dynamic convolution head is the region of interest features (roi_features) and the proposal features. The proposal features can be regarded as a kind of attention mechanism. The proposal features generate the parameters of the convolutional kernel, and then act on the region of interest features to obtain the final prediction result.

[0071] The output features and output proposal boxes of the previous dynamic convolutional layer serve as the proposal features and proposal boxes of the next dynamic convolutional layer. Through multiple iterations, the final target boxes and target classifications are generated.

[0072] It should be noted that in this embodiment, dynamic convolution is sequentially performed on each region of interest feature and its corresponding proposal feature, so as to obtain all target boxes and target classifications.

[0073] In this embodiment, a set of learnable parameters N×4 is initialized to represent the proposal boxes. This set of learned proposal boxes can be understood as the statistical values of the positions where objects may appear in the image. At the same time, a set of learnable proposal features corresponding to the proposal boxes, with a size of N×d, is initialized to represent the target features. The number of proposal features corresponds to the proposal boxes. ROI Align proposes the region of interest features. This method can avoid losing the information of the original feature map during the process, and does not perform quantization throughout the intermediate process to ensure the maximum information integrity. Then, the obtained features are sent into the dynamic convolution head. The input of this detection head is the region of interest features and the proposal features. For each region of interest feature, it corresponds to the proposal feature, and then the final output features are obtained. The proposal features can be regarded as a kind of attention mechanism. The proposal features generate the parameters (params) of the convolutional kernel, and then act on the region of interest features to obtain the final prediction result. The output features and output proposal boxes of the previous dynamic convolution serve as the proposal features and proposal boxes of the next dynamic convolution. Through multiple iterations, the final target boxes and target classifications are generated. The dynamic convolution head is a relatively mature technology in this field and will not be elaborated here.

[0074] Step S4: The enhanced feature map is concatenated with the target boxes output by the instance branch and then input into the mask branch. First, a convolution operation is performed, and then through a spatial attention module and a mask kernel generation module respectively, the spatial attention features and the mask kernel are obtained. The spatial attention features and the mask kernel are multiplied to obtain the segmentation mask.

[0075] In this embodiment, the enhanced feature map is concatenated with the target boxes output by the instance branch and then input into the mask branch. First, a convolution operation is performed. That is, before the enhanced feature map is input into the mask branch (Mask Branch), the spatiality of the feature map is increased. The normalized coordinate information is added in the first convolution, which can be represented by the following formula:

[0076] mask_features = (mask_convs(cat(coord_features, P)))

[0077] where mask_convs represents four 3×3 convolutions, cat represents the concatenate process, coord_features represents the normalized coordinate information, P represents the output enhanced feature map, and mask_features represents the feature map obtained through four 3×3 convolutions.

[0078] It should be noted that the normalized coordinate information in this embodiment comes from the target box output by the instance branch.

[0079] Next, the obtained mask_features feature map is respectively input into the SAM module (spatial attention feature, spatial attention module) and the mask kernel generation module (mask kernel).

[0080] As Figure 5 shown, for the input feature map mask_features ∈ R C×H×W , the input feature map is subjected to max pooling (Max Pool) and average pooling (Avg Pool) in the channel dimension to obtain the pooled feature P max , P avg ∈ R 1×H×W . Then, P max and P avg are dot - producted, passed through a 3×3 convolution (to restore the channel), then sigmoid - operated and multiplied element - wise with the original input to obtain the spatial attention feature.

[0081] Specifically, the operations corresponding to the SAM module are represented by the following formula:

[0082]

[0083] where A sam (mask_features) represents the feature after sigmoid, σ represents the sigmoid function, and · represents the dot product.

[0084]

[0085] where spatial_attention_features sam represents the spatial attention feature finally output by the SAM module. An operation representing multiplication.

[0086] In this embodiment, the mask kernel module adjusts the dimension of the input feature map mask_features through a linear layer to generate a 128-dimensional mask kernel. The mask kernel module is represented by the following formula:

[0087] w = Linear 256->128 (mask_features)

[0088] where w ∈ R 1×C represents the mask kernel.

[0089] Finally, multiply the spatial attention features by the mask kernel to obtain the segmentation mask output by the mask branch, which is represented by the following formula:

[0090]

[0091] where pred_mask is the segmentation mask output by the mask branch, w ∈ R 1×C represents the mask kernel, and spatial_attention_features sam ∈ R C×H×W are the spatial attention features.

[0092] Step S5: Map the segmentation mask, target box, and target classification to the image to be processed to obtain the instance segmentation result.

[0093] Through the previous steps, the segmentation mask, target box, and target classification are obtained respectively. Finally, only by upsampling the segmentation mask to the resolution of the original image to be processed, the segmentation mask, target box, and target classification can be mapped to the image to be processed to achieve instance segmentation.

[0094] In a specific embodiment, since one-to-one assignment will force most predictions to be background, which will reduce the confidence of the category, this embodiment introduces the intersection over union (IoU) to adjust the target classification probability when training the instance segmentation network. That is, the enhanced feature map is input into the instance branch to obtain the target box and target classification, where the target classification probability corresponding to target box i is P i , and IoU is introduced to adjust the target classification probability, and the formula is as follows:

[0095]

[0096] where, is the adjusted target classification probability, S i represents the intersection over union (IoU), and P i represents the target classification probability corresponding to target box i. Si The calculation formula is as follows:

[0097]

[0098] Where a x,y and b x,y respectively represent the pixels of the predicted target box a and the true target box b at (x, y).

[0099] In a specific embodiment, when training the instance segmentation network of the present application, the overall loss function of the network is also calculated, backpropagation is performed, and the network parameters are updated.

[0100] Among them, the overall loss function of the network is composed of the binary cross-entropy loss function of target classification mask loss function target box loss function linear fusion, which can be expressed as:

[0101]

[0102] Where represents the target classification loss, represents the mask loss, represents the bounding box loss, and λ cls is the weight coefficient.

[0103] To solve the end-to-end training, the label assignment is represented as bipartite graph matching. First, a matching score based on pairwise dice is proposed:

[0104]

[0105] For the i-th prediction and the k-th ground truth object in the equation, it is determined by the classification score of the segmentation mask and the dice coefficient, where α is a hyperparameter set to 0.8 to balance the classification and segmentation of the influence, represents the probability that the class of the i-th instance prediction is the class of the k-th ground truth object, m i and t k are respectively the mask of the i-th predicted object and the mask of the k-th ground truth object. The DICE dice coefficient is defined as:

[0106]

[0107] Where and respectively represent the pixels of the i-th predicted mask m and the k-th ground truth object mask t at (x, y). Then, the Hungarian algorithm is used to find the optimal matching between K ground truth objects and N predictions. The segmentation mask loss function by combining the dice loss and the pixel-level binary cross-entropy loss:

[0108]

[0109] Among them, and are the dice loss and the pixel - level binary cross - entropy loss, and λ dice and λ pix are the corresponding coefficients.

[0110] The above - described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A real-time instance segmentation method that fuses a sparse framework and spatial attention, which uses a constructed instance segmentation network to perform instance segmentation on an image, Characterized in that, The instance segmentation network includes a feature extraction network, a feature enhancement network, a mask branch, and an instance branch. The real-time instance segmentation method that fuses a sparse framework and spatial attention includes: Obtain the image to be processed and input it into the feature extraction network to extract multi-scale feature maps; Input the multi-scale feature maps into the feature enhancement network to obtain enhanced feature maps; Input the enhanced feature maps into the instance branch to obtain bounding boxes and object classifications; Concatenate the enhanced feature maps with the bounding boxes output by the instance branch and then input them into the mask branch. First, perform a convolution operation, and then pass through a spatial attention module and a mask kernel generation module respectively to obtain spatial attention features and mask kernels. Multiply the spatial attention features and mask kernels to obtain a segmentation mask; Map the segmentation mask, bounding boxes, and object classifications onto the image to be processed to obtain the instance segmentation result; Among them, the feature extraction network uses ResNet50, and the outputs of residual modules 3, 4, and 5 in the ResNet50 are used as the third-scale feature map, the second-scale feature, and the first-scale feature map extracted; The step of inputting the multi-scale feature maps into the feature enhancement network to obtain enhanced feature maps includes: Input the first-scale feature map into the pyramid pooling module and output the first feature map; Element-wise add the second-scale feature map and the first feature map to obtain the second feature map; Element-wise add the third-scale feature map and the second feature map to obtain the third feature map; Respectively perform convolution operations on the first feature map, the second feature map, and the third feature map and then concatenate them to obtain enhanced feature maps.

2. The real-time instance segmentation method that fuses a sparse framework and spatial attention according to claim 1, Characterized in that, The step of inputting the enhanced feature maps into the instance branch to obtain bounding boxes and object classifications includes: Initialize the proposal boxes and proposal features; Extract the region of interest features from the enhanced feature maps through region feature focusing; Input the region of interest features and the proposal features into the dynamic convolution head to generate the final bounding boxes and object classifications.

3. The real-time instance segmentation method that fuses a sparse framework and spatial attention according to claim 1, Characterized in that, The spatial attention module performs the following operations: Perform max pooling and average pooling on the input feature map in the channel dimension to obtain the pooled feature P max and P avg , and then perform a dot product on P max and P avg , followed by a 3×3 convolution, then a sigmoid operation, and finally multiply the result with the original input feature map to obtain the spatial attention feature.

4. The real-time instance segmentation method that fuses a sparse framework and spatial attention according to claim 1, Characterized in that, The mask kernel generation module performs the following operations: Adjust the dimension of the input feature map through a linear layer to generate a 128-dimensional mask kernel.

5. The real-time instance segmentation method that fuses a sparse framework and spatial attention according to claim 1, Characterized in that, When training the instance segmentation network, the step of inputting the enhanced feature maps into the instance branch to obtain bounding boxes and object classifications further includes: Introduce the intersection over union IoU to adjust the object classification probability, and the formula is as follows: Among them, is the adjusted target classification probability, S i represents the intersection over union IoU, P i represents the target classification probability corresponding to the target bounding box i.

6. The real-time instance segmentation method that fuses a sparse framework and spatial attention according to claim 1, Characterized in that, When training the instance segmentation network, the overall loss function of the network is as follows: Among them, represents the target classification loss, represents the mask loss, represents the bounding box loss, and λ cls is the weight coefficient; Among them, and are the dice loss and the pixel-level binary cross-entropy loss, and λ dice and λ pix are the corresponding weight coefficients.

Citation Information

Patent Citations

  • Panoptic segmentation

    US20210326656A1

  • Similarity propagation for one-shot and few-shot image segmentation

    US20210397876A1