A method for image instance segmentation based on component hints

By building an image instance segmentation model based on component prompts, and using a multi-level component decoder and a multi-level mask decoder to perform hierarchical decoding and fusion of feature tensors, the problem of insufficient accuracy in processing complex scenes in the prior art is solved, and a higher image instance segmentation accuracy is achieved.

CN118982672BActive Publication Date: 2025-05-13TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411069948.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2025-05-13
Estimated Expiration
2044-08-06

AI Technical Summary

Technical Problem

The existing instance segmentation algorithms are insufficient in handling partial occlusion, angular changes, and complex contour edges, and improvements are needed to improve the image instance segmentation accuracy.

Method used

Using an image instance segmentation method based on component prompts, a model including a basic feature encoder, a multi-stage component decoder and a multi-stage mask decoder is constructed. The model performs hierarchical decoding of feature tensors through a multi-stage component decoder, and fuses coarse component feature tensors and fine component feature tensors through a multi-stage mask decoder to generate mask prediction tensors to improve the accuracy of image instance segmentation.

Benefits of technology

By hierarchically decoding and fusion of component feature tensors at different levels, the expression of local features of the image is enhanced and the accuracy of image instance segmentation is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118982672B_ABST
    Figure CN118982672B_ABST
Patent Text Reader

Abstract

The invention discloses a method for image instance segmentation based on component prompts, and constructs an image instance segmentation model based on component prompts, including a basic feature encoder, a multi-level component decoder and a multi-level mask decoder; the basic feature encoder extracts a single-scale or multi-scale feature map of an input image, and performs fusion encoding on the feature map to obtain a feature tensor; the multi-level component decoder decodes the feature tensor to obtain a category prediction tensor, a coarse component feature tensor and a fine component feature tensor; the multi-level mask decoder fuses the coarse component feature tensor and the fine component feature tensor, and uses the fused tensor as a query to decode the feature tensor to obtain an object-level mask prediction tensor and a component-level mask prediction tensor. The advantages are: hierarchical decoding of component feature tensors of different levels, hierarchical decoding of mask prediction tensors through query quantities containing component feature semantics, enhanced expression of local features of the image, and improved image instance segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of instance segmentation, and in particular to an image instance segmentation method based on component prompts. Background Art

[0002] Instance segmentation is an important technology in the field of computer vision. It aims to solve the problem of pixel-level classification and localization of different target instances in an image. The development of this technology stems from the need for more accurate and detailed image understanding, especially on the basis of target detection and semantic segmentation, it is necessary to further distinguish different instances of the same category. Instance segmentation is of great significance in many practical applications, such as autonomous driving, medical image analysis, image editing, etc.

[0003] With the rise of deep learning, especially the development of convolutional neural networks (CNN), instance segmentation technology has made significant progress. Using the powerful feature extraction capabilities of deep learning models, researchers have designed various complex network structures to solve instance segmentation problems, such as Mask R-CNN and YOLACT. These methods have achieved significant improvements in accuracy and efficiency, laying the foundation for the popularization of instance segmentation in practical applications. However, existing instance segmentation algorithms still have shortcomings when dealing with partial occlusion of objects, angle changes, and complex contour edges. Therefore, it is necessary to improve and innovate existing instance segmentation algorithms to improve the accuracy of image instance segmentation. Summary of the invention

[0004] The purpose of the present invention is to provide an image instance segmentation method based on component hints, so as to solve the above-mentioned problems existing in the prior art.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A component-cue-based image instance segmentation method is provided, wherein a component-cue-based image instance segmentation model is constructed, wherein the model comprises a basic feature encoder, a multi-level component decoder and a multi-level mask decoder;

[0007] The basic feature encoder extracts a single-scale or multi-scale feature map from the input image, and performs fusion encoding on the feature map to obtain a feature tensor;

[0008] The multi-level component decoder decodes the feature tensor generated by the basic feature encoder to obtain a first decoded tensor; obtains a category prediction tensor by performing a linear transformation on the first decoded tensor; obtains a coarse component feature tensor by performing a linear transformation and a nonlinear transformation on the first decoded tensor; obtains a fine component feature tensor by performing a linear transformation and a nonlinear transformation on the first decoded tensor and superimposing the coarse component feature tensor with the transformed first decoded tensor;

[0009] The multi-level mask decoder fuses the coarse component feature tensor and the fine component feature tensor to obtain a fused component feature tensor; decodes the feature tensor generated by the basic feature encoder using the fused component feature tensor as a query quantity to obtain a second decoded tensor; performs linear and nonlinear transformations on the second decoded tensor to obtain an object-level mask embedding tensor, multiplies the feature tensor generated by the basic feature encoder with the object-level mask embedding tensor to obtain an object-level mask prediction tensor; performs linear and nonlinear transformations on the second decoded tensor, and adds the object-level mask embedding tensor to the transformed second decoded tensor to obtain a component-level mask embedding tensor, multiplies the feature tensor generated by the basic feature encoder with the component-level mask embedding tensor to obtain a component-level mask prediction tensor.

[0010] Preferably, the basic feature encoder has two structural forms:

[0011] (1) comprising at least one convolution layer for performing convolution operations to extract feature maps; at least one pooling layer for reducing the spatial size of the feature maps; at least one encoding layer, each encoding layer comprising at least one self-attention module for performing spatial dimension conversion on the feature maps obtained after downsampling, and adding level and position sequence information to obtain an initial feature tensor; at least one feedforward neural network for performing linear and nonlinear transformations on the initial feature tensor to obtain a final feature tensor;

[0012] (2) It includes at least one attention mechanism module and a feedforward neural network for downsampling and extracting feature maps; at least one encoding layer, each encoding layer includes at least one self-attention module and a feedforward neural network for fusion encoding the feature maps to obtain the final feature tensor.

[0013] Preferably, the multi-stage component decoder comprises,

[0014] At least one decoding layer, each decoding layer includes at least one self-attention module, a cross-attention module and a feed-forward neural network, used for decoding the feature tensor generated by the basic feature encoder to generate a first decoding tensor;

[0015] At least one prediction head network is used to perform a linear transformation on the first decoding tensor to obtain a category prediction tensor; perform a linear transformation and a nonlinear transformation on the first decoding tensor to obtain a coarse component feature tensor; perform a linear transformation and a nonlinear transformation on the first decoding tensor, and superimpose the coarse component feature tensor with the transformed first decoding tensor to obtain a fine component feature tensor.

[0016] Preferably, the process of the decoding layer of the multi-level component decoder decoding the feature tensor generated by the basic feature encoder to generate the first decoding tensor is specifically as follows:

[0017] When the self-attention module calculates the feature tensor, the query is a learnable target instance query, the key and value are the feature tensors output by the basic feature encoder, and each element in the input feature tensor is re-weighted according to the self-attention mechanism during the calculation process to obtain an attention-weighted feature tensor;

[0018] The cross attention module calculates the attention weighted feature tensor to obtain a cross attention feature tensor;

[0019] The feedforward neural network performs linear transformation and nonlinear transformation on the cross-attention feature tensor to obtain a first decoding tensor.

[0020] Preferably, the prediction head network of the multi-level component decoder performs multiple linear transformations on the first decoding tensor, and performs at least one nonlinear transformation in the middle of each linear transformation to obtain a coarse component feature tensor.

[0021] Preferably, the prediction head network of the multi-level component decoder performs multiple linear transformations on the first decoding tensor, and performs at least one nonlinear transformation in the middle of each linear transformation, and then adds the coarse component feature tensor to the transformed first decoding tensor to obtain the fine component feature tensor.

[0022] Preferably, the multi-stage mask decoder comprises:

[0023] At least one decoding layer, each decoding layer includes at least one self-attention module, a cross-attention module and a feedforward neural network; the coarse component feature tensor and the fine component feature tensor are fused to obtain a fused component feature tensor, the decoding layer uses the fused component feature tensor as a query quantity, and uses the feature tensor generated by the basic feature encoder as a key and a value for decoding to obtain a second decoding tensor;

[0024] At least one prediction head network is used to perform linear transformation and nonlinear transformation on the second decoded tensor to obtain an object-level mask embedding tensor, and multiply the feature tensor generated by the basic feature encoder with the object-level mask embedding tensor to obtain an object-level mask prediction tensor; perform linear transformation and nonlinear transformation on the second decoded tensor, and add the object-level mask embedding tensor and the transformed second decoded tensor to obtain a component-level mask embedding tensor, and multiply the feature tensor generated by the basic feature encoder with the component-level mask embedding tensor to obtain a component-level mask prediction tensor.

[0025] Preferably, the process in which the decoding layer of the multi-stage mask decoder decodes the feature tensor generated by the basic feature encoder to generate the second decoding tensor is specifically as follows:

[0026] The self-attention module uses the fusion component feature tensor as a query and the feature tensor generated by the basic feature encoder as a key and value for calculation to obtain an attention-weighted feature tensor;

[0027] The cross attention module calculates the attention weighted feature tensor to obtain a cross attention feature tensor;

[0028] The feedforward neural network performs linear transformation and nonlinear transformation on the cross-attention feature tensor to obtain a final second decoding tensor.

[0029] Preferably, the prediction head network of the multi-level mask decoder performs multiple linear transformations on the second decoded tensor, and performs at least one nonlinear transformation in the middle of each linear transformation to obtain an object-level mask embedding tensor.

[0030] Preferably, the prediction head network of the multi-level mask decoder performs multiple linear transformations on the second decoded tensor, performs at least one nonlinear transformation in the middle of each linear transformation, and then adds the object-level mask embedding tensor to the transformed second decoded tensor to obtain a component-level mask embedding tensor.

[0031] The beneficial effects of the present invention are as follows: the method of the present invention hierarchically decodes component feature tensors of different levels, and hierarchically decodes mask prediction tensors through query quantities containing component feature semantics, thereby enhancing the expression of local features of the image and improving the accuracy of image instance segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a principle structure diagram of the image instance segmentation method in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] like Figure 1 As shown, in this embodiment, a component-cue-based image instance segmentation method is provided, a component-cue-based image instance segmentation model is constructed, and the image instance segmentation is implemented using the model. The image instance segmentation model specifically includes a basic feature encoder, a multi-level component decoder, and a multi-level mask decoder.

[0035] 1. Basic Feature Encoder

[0036] The basic feature encoder extracts a single-scale or multi-scale feature map from the input image, and performs fusion encoding on the feature map to obtain a feature tensor. The basic feature encoder has two structural forms:

[0037] (1) comprising at least one convolutional layer for performing convolution operations to extract feature maps; at least one pooling layer for reducing the spatial size of the feature maps; at least one encoding layer, each encoding layer comprising at least one self-attention module for performing spatial dimension conversion on the feature maps obtained after downsampling, and adding level and position sequence information to obtain an initial feature tensor; and at least one feedforward neural network for performing linear and nonlinear transformations on the initial feature tensor to obtain a final feature tensor.

[0038] (2) It includes at least one attention mechanism module and one feedforward neural network for downsampling and extracting feature maps; at least one encoding layer, each encoding layer includes at least one self-attention module and one feedforward neural network for fusion encoding the feature maps to obtain the final feature tensor.

[0039] Specifically, the basic feature encoder can first extract a single-scale or multi-scale feature map through a convolutional neural network structure, such as ResNet, U-Net, etc., or through a Transformer structure, such as SwinTransformer, etc., and then fuse the feature map through a coding layer to obtain a feature tensor.

[0040] In this embodiment, the basic feature encoder uses a ResNet network structure to extract feature maps, where the network depth is 50, including 49 convolutional layers and 1 pooling layer. The sizes of the convolutional layers are 7×7, 3×3, and 1×1. It is mainly used to perform downsampling on the original image after performing convolution operations, gradually reducing the spatial size of the feature map while retaining important feature information to obtain the feature map. The specific formula is as follows:

[0041]

[0042] Among them, X represents the input image, the shape is HxWxC, H is the height, W is the width, and C is the number of channels of the image; K represents the convolution kernel, the shape is K H xK W xC,K H is the height of the convolution kernel, K W is the width of the convolution kernel; s is the stride; P is the number of pixels filled at the edge of the input image; Y is the output image, with a shape of Y H xY W xC out ,Y H is the output height, Y W is the output width, C out is the number of output channels.

[0043] The size of the pooling layer is 3×3, and the maximum pooling method is adopted, which is mainly used to reduce the spatial size of the feature map, reduce the number of parameters, and save computing resources;

[0044] The feature map is then fused and encoded through the MSDeformAttnPixel encoding layer, which contains a self-attention module and a feed-forward neural network. The number of heads in the MSDeformAttnPixel encoding layer is at least 1.

[0045] The encoding layer transforms the spatial dimension of the feature map through the self-attention module and adds level position sequence information to it. Then, the initial feature tensor generated by the self-attention module is linearly transformed and nonlinearly transformed through the feedforward neural network. The nonlinear transformation is completed through the ReLU activation function to obtain a feature tensor with stronger fitting ability and more robustness. The calculation formula is as follows:

[0046] FFN(X)=ReLu(0,X n W 1 +b 1 )W 2 +b 2

[0047] Among them, FFN() represents two linear transformations and the ReLU activation function calculation in the middle, W 1 , W 2 , b 1 , b 2 They represent the weight matrix and bias vector respectively, and X is the initial feature tensor.

[0048] 2. Multi-level component decoder

[0049] The multi-level component decoder decodes the feature tensor generated by the basic feature encoder to obtain a category prediction tensor, a coarse component feature tensor, and a fine component feature tensor.

[0050] Specifically: the multi-level component decoder includes at least one decoding layer and a prediction head network, and each decoding layer includes at least one self-attention module, a cross-attention module and a feedforward neural network.

[0051] The decoding layer decodes the feature tensor generated by the basic feature encoder to generate a first decoding tensor; the prediction head network performs a linear transformation on the first decoding tensor to obtain a category prediction tensor; performs a linear transformation and a nonlinear transformation on the first decoding tensor to obtain a coarse component feature tensor; performs a linear transformation and a nonlinear transformation on the first decoding tensor, and superimposes the coarse component feature tensor with the transformed first decoding tensor to obtain a fine component feature tensor.

[0052] In this embodiment, the multi-level component decoder is composed of at least one layer of DetrTransformer decoding layer, including at least one self-attention module, a cross-attention module, and a feedforward neural network; wherein the number of heads in the DetrTransformer decoding layer is at least 1.

[0053] The self-attention module first calculates the feature tensor. The query is a learnable target instance query. The key and value are the feature tensors output by the basic feature encoder. During the calculation process, each element in the input tensor is re-weighted according to the self-attention mechanism to obtain the attention-weighted feature tensor. The specific formula is:

[0054]

[0055] Among them, Attention() is the self-attention calculation function; Q is the query quantity that can be learned; K is the key matrix of the feature tensor, and V is the value matrix of the feature tensor; d k is the scaling factor; T represents the matrix transpose.

[0056] Then, the aforementioned attention weighted feature tensor is calculated by the cross-attention module. Different from the self-attention calculation, the cross-attention calculation focuses on the interaction between different sequences and can effectively capture the correlation between them, thereby obtaining the cross-attention feature tensor. The specific formula is:

[0057]

[0058] Among them, Attention() is the self-attention calculation function; Q is the query quantity that can be learned; K is the key matrix of the feature tensor, and V is the value matrix of the feature tensor; d k is the scaling factor; T represents the matrix transpose.

[0059] Next, the aforementioned cross-attention feature tensor is linearly transformed and nonlinearly transformed through the feedforward neural network, wherein the nonlinear transformation can be completed by the ReLu activation function, so as to obtain the first decoding tensor with stronger expression ability.

[0060] The prediction head network performs a linear transformation on the first decoded tensor to obtain a category prediction tensor. The specific formula is:

[0061] C l s_Pred = nn.Li near (X)

[0062] Among them, nn.L i near() represents linear transformation, X is the tensor output by the standardized feedforward neural network; Cl s_Pred is the category prediction tensor.

[0063] The prediction head network performs three (the specific number can be selected according to the actual situation to better meet the actual needs) linear transformations on the first decoding tensor, and performs one (at least one, the specific number can be selected according to the actual situation) nonlinear transformation in the middle of each linear transformation to obtain the coarse component feature tensor. The specific formula is:

[0064] Roughpart_Pred=nn.Li near(ReLu(nn.L i near(ReLU(nn.Li near(X)))))

[0065] Among them, Roughpart_Pred is the rough part feature tensor.

[0066] The prediction head network performs three linear transformations on the first decoding tensor (the specific number can be selected according to the actual situation to better meet the actual needs), and performs one (at least one, the specific number can be selected according to the actual situation) nonlinear transformation in the middle of each linear transformation, and then adds the coarse component feature tensor to the transformed first decoding tensor to obtain the fine component feature tensor. The specific formula is:

[0067] Y=nn.Li near(ReLu(nn.Li near(ReLU(nn.Li near(X))))))

[0068] Detai ledpart_Pred=Roughpart_Pred+Y

[0069] Among them, Y is the first decoded tensor after transformation; Detailedpart_Pred is the detailed part feature tensor.

[0070] 3. Multi-level mask decoder

[0071] The multi-level mask decoder fuses the coarse component feature tensor and the fine component feature tensor, and uses the fused tensor as a query to decode the feature tensor generated by the basic feature encoder to obtain the object-level mask prediction tensor and the component-level mask prediction tensor.

[0072] Specifically: the multi-level mask decoder includes at least one decoding layer and a prediction network, and each decoding layer includes at least one self-attention module, a cross-attention module and a feedforward neural network.

[0073] The multi-level mask compiler fuses the coarse component feature tensor and the fine component feature tensor to obtain a fused component feature tensor; the decoding layer uses the fused component feature tensor as a query quantity and decodes the feature tensor generated by the basic feature encoder as a key and a value to obtain a second decoded tensor; the prediction head network performs linear and nonlinear transformations on the second decoded tensor to obtain an object-level mask embedding tensor, and then multiplies the feature tensor generated by the basic feature encoder with the object-level mask embedding tensor to obtain an object-level mask prediction tensor; the prediction head network performs linear and nonlinear transformations on the second decoded tensor, and adds the object-level mask embedding tensor to the transformed second decoded tensor to obtain a component-level mask embedding tensor, and multiplies the feature tensor generated by the basic feature encoder with the component-level mask embedding tensor to obtain a component-level mask prediction tensor.

[0074] In this embodiment, the multi-level mask decoder is composed of at least one layer of DetrTransformer decoding layer, including at least one self-attention module, a cross-attention module, and a feedforward neural network; wherein the number of heads in the DetrTransformer decoding layer is at least 1.

[0075] The aforementioned coarse component feature tensor and fine component feature tensor are added together to fuse them, and the fused component feature tensor is obtained. The calculation formula is:

[0076] M i xedpart=Detai ledpart_Pred+Roughpart_Pred

[0077] Among them, Mixedpart is the fusion component feature tensor.

[0078] The self-attention module uses the fusion component feature tensor as a query, and uses the feature tensor generated by the basic feature encoder as a key and value for calculation to obtain an attention-weighted feature tensor. The cross-attention module calculates the attention-weighted feature tensor to obtain a cross-attention feature tensor. The feedforward neural network performs linear and nonlinear transformations on the cross-attention feature tensor to obtain a final second decoding tensor.

[0079] The prediction head network performs three linear transformations on the second decoded tensor (the specific number of times can be selected according to actual conditions to better meet actual needs), performs one (at least one, the specific number of times can be selected according to actual conditions) nonlinear transformation in the middle of each linear transformation, and obtains the object-level mask embedding tensor, and then multiplies the feature tensor generated by the basic feature encoder with the object-level mask embedding tensor to obtain the object-level mask prediction tensor. The specific calculation formula is:

[0080] ObjectMask_embed=nn.Li near(ReLu(nn.Li near(ReLU(nn.L i near(X))))))

[0081] ObjectMask_Pred=torch.ei nsum(ObjectMask_embed,X)

[0082] Among them, torch.ei nsum() represents the matrix multiplication operation, ObjectMask_embed is the object-level mask embedding tensor, X is the feature tensor, and ObjectMask_Pred is the object-level mask prediction tensor.

[0083] The prediction head network performs three linear transformations on the second decoded tensor (the specific number of times can be selected according to actual conditions to better meet actual needs), performs one nonlinear transformation in the middle of each linear transformation (at least once, the specific number of times can be selected according to actual conditions), and obtains the transformed second decoded tensor, and then adds the object-level mask embedding tensor to the transformed second decoded tensor to obtain the component-level mask embedding tensor, and multiplies the feature tensor generated by the basic feature encoder with the component-level mask embedding tensor to obtain the component-level mask prediction tensor. The specific calculation formula is:

[0084] Mask_embed=nn.Li near(ReLu(nn.Li near(ReLU(nn.Li near(X))))))

[0085] PartMask_embed=Mask_embed+ObjectMask_embed

[0086] PartMask_Pred=torch.ei nsum(PartMask_embed,X)

[0087] Among them, Mask_embed is the transformed second decoded tensor; PartMask_embed is the component-level mask embedding tensor; PartMask_Pred is the component-level mask prediction tensor.

[0088] By adopting the above technical solution disclosed in the present invention, the following beneficial effects are obtained:

[0089] The present invention provides an image instance segmentation method based on component hints, hierarchically decodes coarse component feature tensors and fine component feature tensors, fuses multi-level component feature tensors, and uses them as a query quantity containing component feature semantics to hierarchically decode object-level mask prediction tensors and component-level mask prediction tensors, thereby enhancing the expression of local semantic features of the image and improving the accuracy of image instance segmentation.

[0090] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.

Claims

1. A method for image instance segmentation based on component hints, characterized in that: Constructing a component-cue-based image instance segmentation model, the model comprising a basic feature encoder, a multi-level component decoder, and a multi-level mask decoder; The basic feature encoder extracts a single-scale or multi-scale feature map from the input image, and performs fusion encoding on the feature map to obtain a feature tensor; The multi-level component decoder decodes the feature tensor generated by the basic feature encoder to obtain a first decoded tensor; obtains a category prediction tensor by performing a linear transformation on the first decoded tensor; obtains a coarse component feature tensor by performing a linear transformation and a nonlinear transformation on the first decoded tensor; obtains a fine component feature tensor by performing a linear transformation and a nonlinear transformation on the first decoded tensor and superimposing the coarse component feature tensor with the transformed first decoded tensor; The multi-level mask decoder fuses the coarse component feature tensor and the fine component feature tensor to obtain a fused component feature tensor; The feature tensor generated by the basic feature encoder is decoded using the fused component feature tensor as the query quantity to obtain a second decoded tensor; the second decoded tensor is linearly transformed and nonlinearly transformed to obtain an object-level mask embedding tensor, and the feature tensor generated by the basic feature encoder is multiplied by the object-level mask embedding tensor to obtain an object-level mask prediction tensor; the second decoded tensor is linearly transformed and nonlinearly transformed, and the object-level mask embedding tensor is added to the transformed second decoded tensor to obtain a component-level mask embedding tensor, and the feature tensor generated by the basic feature encoder is multiplied by the component-level mask embedding tensor to obtain a component-level mask prediction tensor.

2. The method for image instance segmentation based on component hints according to claim 1, characterized in that: The basic feature encoder has two structural forms: (1) comprising at least one convolution layer for performing convolution operations to extract feature maps; at least one pooling layer for reducing the spatial size of the feature maps; at least one encoding layer, each encoding layer comprising at least one self-attention module for performing spatial dimension conversion on the feature maps obtained after downsampling, and adding level and position sequence information to obtain an initial feature tensor; at least one feedforward neural network for performing linear and nonlinear transformations on the initial feature tensor to obtain a final feature tensor; (2) It includes at least one attention mechanism module and a feedforward neural network for downsampling and extracting feature maps; at least one encoding layer, each encoding layer includes at least one self-attention module and a feedforward neural network for fusion encoding the feature maps to obtain the final feature tensor.

3. The method for image instance segmentation based on component hints according to claim 1, characterized in that: The multi-stage component decoder comprises, At least one decoding layer, each decoding layer includes at least one self-attention module, a cross-attention module and a feed-forward neural network, used for decoding the feature tensor generated by the basic feature encoder to generate a first decoding tensor; At least one prediction head network is used to perform a linear transformation on the first decoding tensor to obtain a category prediction tensor; perform a linear transformation and a nonlinear transformation on the first decoding tensor to obtain a coarse component feature tensor; perform a linear transformation and a nonlinear transformation on the first decoding tensor, and superimpose the coarse component feature tensor with the transformed first decoding tensor to obtain a fine component feature tensor.

4. The method for image instance segmentation based on component hints according to claim 3, characterized in that: The process of the decoding layer of the multi-level component decoder decoding the feature tensor generated by the basic feature encoder to generate the first decoding tensor is specifically as follows: When the self-attention module calculates the feature tensor, the query is a learnable target instance query, the key and value are the feature tensors output by the basic feature encoder, and each element in the input feature tensor is re-weighted according to the self-attention mechanism during the calculation process to obtain an attention-weighted feature tensor; The cross attention module calculates the attention weighted feature tensor to obtain a cross attention feature tensor; The feedforward neural network performs linear transformation and nonlinear transformation on the cross-attention feature tensor to obtain a first decoding tensor.

5. The method for image instance segmentation based on component hints according to claim 3, characterized in that: The prediction head network of the multi-level component decoder performs multiple linear transformations on the first decoding tensor, and performs at least one nonlinear transformation in the middle of each linear transformation to obtain a coarse component feature tensor.

6. The method for image instance segmentation based on component hints according to claim 3, characterized in that: The prediction head network of the multi-level component decoder performs multiple linear transformations on the first decoding tensor, and performs at least one nonlinear transformation in the middle of each linear transformation, and then adds the coarse component feature tensor to the transformed first decoding tensor to obtain a fine component feature tensor.

7. The method for image instance segmentation based on component hints according to claim 1, characterized in that: The multi-stage mask decoder comprises: at least one decoding layer, each decoding layer comprising at least one self-attention module, a cross-attention module and a feed-forward neural network; The coarse component feature tensor and the fine component feature tensor are fused to obtain a fused component feature tensor, and the decoding layer uses the fused component feature tensor as a query quantity and the feature tensor generated by the basic feature encoder as a key and a value for decoding to obtain a second decoded tensor; At least one prediction head network is used to perform linear transformation and nonlinear transformation on the second decoded tensor to obtain an object-level mask embedding tensor, and multiply the feature tensor generated by the basic feature encoder with the object-level mask embedding tensor to obtain an object-level mask prediction tensor; perform linear transformation and nonlinear transformation on the second decoded tensor, and add the object-level mask embedding tensor and the transformed second decoded tensor to obtain a component-level mask embedding tensor, and multiply the feature tensor generated by the basic feature encoder with the component-level mask embedding tensor to obtain a component-level mask prediction tensor.

8. The method for image instance segmentation based on component hints according to claim 7, characterized in that: The process of the decoding layer of the multi-stage mask decoder decoding the feature tensor generated by the basic feature encoder to generate the second decoding tensor is specifically as follows: The self-attention module uses the fusion component feature tensor as a query and the feature tensor generated by the basic feature encoder as a key and value for calculation to obtain an attention-weighted feature tensor; The cross attention module calculates the attention weighted feature tensor to obtain a cross attention feature tensor; The feedforward neural network performs linear transformation and nonlinear transformation on the cross-attention feature tensor to obtain a final second decoding tensor.

9. The method for image instance segmentation based on component hints according to claim 7, characterized in that: The prediction head network of the multi-level mask decoder performs multiple linear transformations on the second decoded tensor, and performs at least one nonlinear transformation in the middle of each linear transformation to obtain an object-level mask embedding tensor.

10. The method for image instance segmentation based on component hints according to claim 7, characterized in that: The prediction head network of the multi-level mask decoder performs multiple linear transformations on the second decoded tensor, performs at least one nonlinear transformation in the middle of each linear transformation, and then adds the object-level mask embedding tensor to the transformed second decoded tensor to obtain a component-level mask embedding tensor.

Citation Information

Patent Citations

  • Image semantic segmentation method based on large convolution kernel backbone network

    CN116612283A

  • Biological image instance segmentation method based on cross-scale decoding

    CN117152441A