Camouflage object instance segmentation method, device and system based on multi-scale pooling modeling

By employing a multi-scale pooling modeling approach, utilizing pyramid pooling Transformer and multi-scale complementary feature pooling modules, and combining spatial attention, we achieve high efficiency and accuracy in camouflaged object instance segmentation, thus solving the problem of insufficient performance in camouflaged object segmentation in existing technologies.

CN116433911BActive Publication Date: 2026-03-17HENGYANG NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-21
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively understand the semantic information of camouflaged objects in images at the instance level, and are insufficient at multiple scales in extracting detailed information, resulting in poor performance in camouflaged object segmentation.

Method used

A multi-scale pooling modeling approach is adopted, which extracts multi-scale features through a pyramid pooling Transformer backbone network, and combines a multi-scale complementary feature pooling module and an instance normalization method that integrates spatial attention to perform camouflaged object instance segmentation.

Benefits of technology

It improves the accuracy and efficiency of camouflaged object segmentation, can better explore camouflaged object information in images, preserves global information at different scales, and outputs more accurate camouflaged object instance results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116433911B_ABST
    Figure CN116433911B_ABST
Patent Text Reader

Abstract

The application discloses a camouflage object instance segmentation method, device and system based on multi-scale pooling modeling, wherein the method comprises the following steps: acquiring an image to be segmented; using camouflage object data with instance-level labels to supervise training of a model, and obtaining an extraction result of multi-scale features of the image to be segmented through the supervised training; based on the extraction result of the multi-scale features, processing camouflage objects in the image to be segmented by using a multi-scale pooling mode to obtain high-resolution mask features and instance perception results containing global information; and based on the instance perception results, focusing on the high-resolution mask features by using an instance normalization method with fused spatial attention to output predicted instance results of the model. The technical scheme of the application can solve the problem that the existing instance segmentation technology has poor segmentation performance in camouflage image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image segmentation technology, and in particular to a method, apparatus and system for segmenting camouflaged object instances based on multi-scale pooling modeling. Background Technology

[0002] Research by biologists indicates that camouflaged organisms exhibit a high degree of similarity to their surroundings, making visual tasks involving camouflaged objects more challenging than those involving ordinary objects. In recent years, with advancements in camouflaged object detection, camouflaged object instance segmentation has also gradually progressed.

[0003] Currently, the task of camouflaged object instance segmentation requires the algorithm model to understand the camouflage semantic information in the image at the instance level and segment the region to which the camouflaged object belongs from the image.

[0004] The study also found that by applying multi-scale pooling modeling technology, extracting detailed information of camouflaged objects in images at different scales, including texture, shape, and edges, can further improve the algorithm model's ability to understand information about camouflaged objects, thereby improving the model's segmentation performance. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method, apparatus, and system for segmenting camouflaged object instances based on multi-scale pooling modeling. The aforementioned method for segmenting camouflaged object instances based on multi-scale pooling modeling can be used to detect camouflaged objects in images and segment them into instance-level pixels. The user provides an image, and the algorithm automatically segments the image into instance-level pixel regions containing camouflaged objects.

[0006] In a first aspect, the present invention provides a method for segmenting camouflaged object instances based on multi-scale pooling modeling;

[0007] Obtain the image to be segmented; use camouflaged object data with instance-level labels to perform supervised training on the model, and obtain the multi-scale feature extraction results of the image to be segmented through the supervised training;

[0008] Based on the extraction results of the multi-scale features, the camouflaged objects in the image to be segmented are processed by multi-scale pooling to obtain high-resolution mask features and instance perception results containing global information.

[0009] Based on the instance perception results, the instance normalization method that integrates spatial attention is used to focus on the high-resolution mask features, and the model's predicted instance results are output.

[0010] Secondly, the present invention provides a camouflaged object instance segmentation system based on multi-scale pooling modeling;

[0011] A camouflaged object instance segmentation system based on multi-scale pooling modeling includes:

[0012] The multi-scale feature acquisition module is configured to: acquire the image to be segmented; perform supervised training on the model using camouflaged object data with instance-level labels, and obtain the extraction results of multi-scale features of the image to be segmented through the supervised training;

[0013] The multi-scale feature processing module is configured to: based on the extraction results of the multi-scale features, process the camouflaged objects in the image to be segmented using a multi-scale pooling method to obtain high-resolution mask features and instance perception results containing global information.

[0014] The model output prediction module is configured to: based on the instance perception results, focus the high-resolution mask features using an instance normalization method that integrates spatial attention, and output the model's predicted instance results.

[0015] Thirdly, the present invention provides a device for segmenting camouflaged object instances based on multi-scale pooling modeling, including a processor and a storage medium;

[0016] The storage medium is used to store instructions;

[0017] The processor is configured to operate according to the instructions to perform the steps described in the first aspect.

[0018] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0019] Compared with the prior art, the beneficial effects of the present invention are:

[0020] The aforementioned method for segmenting camouflaged objects based on multi-scale pooling modeling includes the following operations: acquiring an image to be segmented; supervising the model using camouflaged object data with instance-level labels, and obtaining the extraction results of multi-scale features of the image to be segmented through the supervised training; processing the camouflaged objects in the image to be segmented using multi-scale pooling based on the extraction results of the multi-scale features, obtaining high-resolution mask features and instance perception results containing global information; and focusing the high-resolution mask features using an instance normalization method that integrates spatial attention, and outputting the model's predicted instance results based on the instance perception results.

[0021] In the above operations, when mining for camouflaged objects in an image at the instance level, the user first provides an image, and the algorithm segments the pixel regions in the image containing the camouflaged objects. Specifically, the pyramid pooling Transformer is used because multi-scale feature extraction helps the model better explore information about camouflaged objects in the image; the pooling learning Transformer is used because adaptive average pooling learning helps the model retain more global information at different scales.

[0022] Specifically, the reason for adopting the multi-scale complementary feature pooling module is that fusing adjacent features in a complementary and pooling manner is conducive to the full fusion of multi-scale information; subsequently, the instance normalization module with fused spatial attention is adopted because fully exploring the spatial information of high-resolution mask features is conducive to the model predicting disguised object instances more accurately.

[0023] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0024] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0025] Figure 1 This is a flowchart of the method in the first embodiment;

[0026] Figure 2 This is a model structure diagram of the first embodiment;

[0027] Figure 3 The first embodiment uses a pooled learning Transformer structure;

[0028] Figure 4 This is an enlarged view of one of the structures of the pooling learning Transformer in the first embodiment;

[0029] Figure 5 A magnified view of another structure of the pooling learning Transformer in the first embodiment;

[0030] Figure 6 The first embodiment is the HS-FFN structure;

[0031] Figure 7 This is the multi-scale complementary feature pooling module of the first embodiment;

[0032] Figure 8 This is the instance normalization module for the fusion of spatial attention in the first embodiment;

[0033] Figure 9 The input image for the first embodiment;

[0034] Figure 10 For instance-level tags of the first embodiment;

[0035] Figure 11 This is a rendering of the MSPNet of the present invention in its first embodiment;

[0036] Figure 12 This is a rendering of the OSFormer from the first embodiment;

[0037] Figure 13 This is a rendering of SOLOv2 from the first embodiment;

[0038] Figure 14 This is a rendering of SOTR in the first embodiment. Detailed Implementation

[0039] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0040] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0041] Example 1

[0042] like Figure 1 As shown, this embodiment provides a method for segmenting camouflaged object instances based on multi-scale pooling modeling, including:

[0043] S101: Obtain the image to be segmented (the image to be segmented refers to the camouflaged image to be segmented); use camouflaged object data with instance-level labels to perform supervised training on the model (the model specifically refers to the image segmentation processing model), and obtain the multi-scale feature extraction results of the image to be segmented through the supervised training;

[0044] S102: Based on the extraction results of the multi-scale features, the disguised objects in the image to be segmented are processed by multi-scale pooling to obtain high-resolution mask features and instance perception results containing global information.

[0045] S103: Based on the instance perception results, use the instance normalization method that integrates spatial attention to focus on the high-resolution mask features and output the predicted instance results of the model (specifically, the image segmentation processing model).

[0046] like Figure 2As shown, the technical solution of this invention divides the task of segmenting camouflaged object instances into three processing stages: feature extraction, multi-scale feature processing, and model output prediction.

[0047] For the feature extraction stage, multi-scale feature extraction of the input image is performed using a pyramid pooling Transformer backbone network (P2TBackbone).

[0048] For the multi-scale feature processing stage, firstly, a pooling learning Transformer structure (PLT Head) is used to enhance the features of the three-layer feature map. Then, a multi-scale complementary feature pooling module (MCFP) is used to fuse the features of different scales of the disguised object instance. Finally, high-quality mask features and instance perception results with global information are obtained.

[0049] For the model output prediction stage, this embodiment uses an instance normalization method fused with spatial attention (FSA-IN) to normalize high-resolution mask features using instance-aware results, thereby improving the quality of the model output prediction map. Figure 2 In this context, Input represents the model input, and Output represents the model output.

[0050] As one or more embodiments, during the execution of step S101 above, the model is trained under supervision using camouflaged object data with instance-level labels. The multi-scale feature extraction results of the image to be segmented are obtained through the supervised training, specifically including the following operation steps:

[0051] S1011: Based on the pyramid pooling Transformer backbone network, feature extraction is performed on the image to be segmented;

[0052] The pyramid pooled Transformer backbone network includes: a first embedded block, a second embedded block, a third embedded block, and a fourth embedded block connected in sequence;

[0053] S1012: The first embedding block performs a 7*7 convolutional layer and a P2T base block on the image to be segmented, outputting a first feature map; a pyramid pooling transformer. This model is used as the backbone network of this invention, and its main function is to extract multi-scale features.

[0054] The first, second, third, and fourth embedding blocks are components of the internal structure of the pyramid pooling transformer. The first embedding block mainly operates on the image to be segmented input into the image segmentation processing model, and the operations performed sequentially include a 7*7 convolutional layer and a P2T base block.

[0055] S1013: The second embedding block performs double downsampling and two P2T base blocks on the first feature map, and outputs the second feature map;

[0056] S1014: The third embedding block performs double downsampling and two P2T base blocks on the second feature map, and outputs the third feature map;

[0057] S1015: The fourth embedding block performs double downsampling and 6 P2T base blocks on the third feature map, and outputs the fourth feature map.

[0058] The P2T basic block comprises, in sequence, a first multi-head self-attention module based on pooling, a first adder, a first normalization layer, a feedforward network, a second adder, and a second normalization layer. The Pyramid poolingtransformer is an existing model in the prior art; therefore, the specific operations of each module within the P2T basic block will not be elaborated in this embodiment.

[0059] Table 1 shows the structure of each embedded block in the pyramid pooled Transformer.

[0060] Table 1. Pyramid Pooling Transformer Structure (assuming the input image size is 3×H×W).

[0061]

[0062] As one or more embodiments, during the execution of step S102 above, based on the extraction results of the multi-scale features, the camouflaged objects in the image to be segmented are processed using a multi-scale pooling method to obtain high-resolution mask features and instance-aware results containing global information. Specifically, the following operation steps are included:

[0063] S1021: Based on the second feature map, the third feature map, and the fourth feature map, the concatenation and flattening operations are input into the pooling learning Transformer structure to obtain the second high-quality mask map, the third high-quality mask map, and the fourth high-quality mask map of the corresponding size, as well as coarse instance perception parameters with global information.

[0064] The coarse instance perception parameters use a fully connected layer to obtain location labels, use the location labels to determine the position of the instance perception parameters and calculate the confidence of the current location label, and use the confidence of the location label to filter out incorrect parameters to obtain the filtered accurate instance perception parameters.

[0065] See Figures 3-5 ; Figure 3 The first embodiment uses a pooled learning Transformer structure; Figure 4 This is an enlarged view of one of the structures of the pooling learning Transformer in the first embodiment; Figure 5 A magnified view of another structure of the pooling learning Transformer in the first embodiment;

[0066] Figure 4 In this context, Multi-head deformable self-attention represents the first multi-head deformable self-attention mechanism. Add&Norm represents the fourth adder and the first normalization operation module. HS-FFN represents the first HS-FFN structure, i.e., an improved feedforward network. restore represents the feature recovery operation. F 1,j (Represents high-quality mask characteristics); f m (representing the second, third, and fourth feature maps); P (representing positional encoding information).

[0067] Figure 5 In this context, the terms include: Multi-head deformable self-attention (representing the second multi-head deformable self-attention mechanism); Add&Norm (representing the fifth adder and the second normalization operation module); Pool (representing adaptive average pooling operation); flatten (representing flattening operation); HS-FFN (the second HS-FFN structure, i.e., an improved feedforward network); and F. 2,j (representing grid features); k i, m i (k i Represents the coarse instance perception parameter; m i That is, F 1,j (representing high-quality mask characteristics);

[0068] S1022: Based on the first feature map, the second high-quality mask map, the third high-quality mask map, and the fourth high-quality mask map, they are simultaneously input into the multi-scale complementary feature pooling module, and output a shared high-resolution mask map.

[0069] Furthermore, the pooling learning Transformer structure includes: an encoder, a multi-scale feature pooling operator, and a decoder;

[0070] The encoder, the multi-scale feature pooling operator, and the decoder are connected sequentially.

[0071] The encoder includes: a structure consisting of a third adder, a first multi-head deformable self-attention module, a fourth adder, a first normalized operation module, and a first HS-FFN structure connected in sequence;

[0072] The multi-scale feature pooling operator is a module that performs the following operations in sequence: feature recovery operation, three adaptive average pooling operations, and flattening operation.

[0073] The decoder includes: a structure consisting of a second multi-head deformable self-attention module, a fifth adder, a second normalized operation module, and a second HS-FFN structure connected in sequence;

[0074] The third adder of the encoder connects the second feature map, the third feature map, the fourth feature map, and the position encoding information;

[0075] The first HS-FFN structure is the same as the second HS-FFN structure; the HS-FFN structure includes: a convolutional layer with a 3*3 kernel, a group normalization layer, a Hardswish activation layer and a convolutional layer with a 3*3 kernel connected in sequence.

[0076] Furthermore, the working principle of the pooling learning Transformer (PLT) structure is as follows: Figure 3 As shown, the above pooling learning Transformer structure is divided into an encoder, a multi-scale feature pooling operator, and a decoder.

[0077] For the encoder, relative position encoding information was fused with the position information of the camouflaged object instance, and methods such as... Figure 6 The HS-FFN structure is used to obtain the positional relationship between adjacent tokens, and high-quality mask features are obtained through feature recovery operations. Finally, through a multi-scale feature pooling operator and a decoder connected in sequence, an instance-aware result containing global information is obtained. See also Figure 6 HS-FFN (represents an improved feedforward network); 3x3 Conv (represents a convolutional layer with a 3x3 kernel); Hardswish (represents a Hardswish activation layer); 3x3 Conv (represents a convolutional layer with a 3x3 kernel); GN (represents a group normalization layer).

[0078] Furthermore, the internal structure of the multi-scale complementary feature pooling module includes: a first input terminal, a second input terminal, a third input terminal, and a fourth input terminal, as well as multiple adders;

[0079] The first input terminal is used to input the first feature map;

[0080] The second input terminal is used to input a second high-quality mask image;

[0081] The third input terminal is used to input a third high-quality mask image;

[0082] The fourth input terminal is used to input the fourth high-quality mask image;

[0083] The second high-quality mask image is upsampled by two times, and then subjected to feature subtraction and absolute value operation with the first feature image. After passing through a convolutional layer with a kernel size of 3*3, the first complementary feature is obtained.

[0084] The third high-quality mask image is upsampled twice, and then subjected to feature subtraction and absolute value operation with the second high-quality mask image. After passing through a convolutional layer with a kernel size of 3*3, the second complementary feature is obtained.

[0085] The fourth high-quality mask image is upsampled twice, and then subjected to feature subtraction and absolute value operation with the third high-quality mask image. After passing through a convolutional layer with a kernel size of 3*3, the third complementary feature is obtained.

[0086] The first complementary feature is connected to the first feature map through an adder (the adder here is different from the first, second, third, fourth, fifth, and sixth adders with the same structure, but the adder here has the same structural form as the first adder, etc.), and after passing through a convolutional layer with a convolutional kernel of size 3*3, the first original feature is obtained, and then input into the PPM module to obtain the first enhanced mask image;

[0087] The aforementioned second complementary feature is connected to the second high-quality mask image through an adder, and after passing through a convolutional layer with a kernel size of 3*3, the second original feature is obtained, which is then input into the PPM module to obtain the second enhanced mask image;

[0088] The aforementioned third complementary feature is connected to the third high-quality mask image through an adder, and after passing through a convolutional layer with a kernel size of 3*3, the third original feature is obtained, which is then input into the PPM module to obtain the third enhanced mask image.

[0089] The aforementioned fourth high-quality mask image is passed through a convolutional layer with a kernel size of 3*3 to obtain the fourth original feature, which is then input into the PPM module to obtain the fourth enhanced mask image.

[0090] Furthermore, the PPM module includes: one residual branch and four side branches;

[0091] Among them, the four lateral branches are connected in parallel;

[0092] The four lateral branches include: the first lateral branch, the second lateral branch, the third lateral branch, and the fourth lateral branch;

[0093] The first side branch includes: a convolutional layer with a kernel size of 1*1 connected in sequence, an adaptive average pooling layer with a kernel size of 1*1, and then an upsampling operation to obtain the first pooling feature;

[0094] The second side branch includes: a convolutional layer with a kernel size of 1*1 connected in sequence, an adaptive average pooling layer with a kernel size of 2*2, and then an upsampling operation to obtain the second pooling feature;

[0095] The third side branch includes: a convolutional layer with a kernel size of 1*1 connected in sequence, an adaptive average pooling layer with a kernel size of 3*3, and then an upsampling operation to obtain the third pooling feature;

[0096] The fourth side branch includes: a convolutional layer with a kernel size of 1*1 connected in sequence, an adaptive average pooling layer with a kernel size of 6*6, and then an upsampling operation to obtain the fourth pooling feature.

[0097] The residual branch includes: a splicer and a convolutional layer with a kernel size of 1*1 connected in sequence;

[0098] The splicer's input includes a first pooling feature, a second pooling feature, a third pooling feature, a fourth pooling feature, and either a first original feature, a second original feature, a third original feature, or a fourth original feature.

[0099] Furthermore, the working principle of the multi-scale complementary feature pooling module (MCFP) includes: Figure 7 As shown, by utilizing the complementary properties of information at different scales in the same image, an enhanced feature S is obtained using a complementary fusion method. a Simultaneously, a pyramid pooling method was employed to obtain a further enhanced high-quality mask image (mf). i .

[0100] As one or more embodiments, during the execution of step S103 above: based on the instance perception result, the instance normalization method that integrates spatial attention is used to focus the high-resolution mask features, and the predicted instance result of the model (specifically, the image segmentation processing model) is output, specifically including:

[0101] S1031: Input the high-resolution mask image into the instance normalization module that fuses spatial attention for spatial weighted attention to obtain a spatially weighted mask;

[0102] The instance normalization module for fusion spatial attention includes: two main branches;

[0103] Among them, the two main branches include the first main branch and the second main branch;

[0104] The first main branch includes: two linear layers, a multiplier, and a sixth adder connected in sequence;

[0105] The second main branch includes: a max pooling layer and an average pooling layer connected in sequence, a splicer, a convolutional layer with a kernel size of 3*3, a sigmoid activation layer, and a multiplier;

[0106] Figure 7 This is the multi-scale complementary feature pooling module of the first embodiment; Figure 8 This is the instance normalization module for the fusion of spatial attention in the first embodiment;

[0107] exist Figure 7 In the middle, f l (representing the first feature map); m l (representing the second high-quality mask); m2 (representing the third high-quality mask); m3 (representing the fourth high-quality mask); PPM (representing a pyramid pooling module); UPx2 (representing a 2x upsampling operation); C1 (representing the first complementary feature); C2 (representing the second complementary feature); C3 (representing the third complementary feature); Conv3x3 (representing a convolutional layer with a 3x3 kernel).

[0108] exist Figure 8 In this context, Linear represents a linear layer; Maxpool represents a max-pooling layer; Avgpool represents an average pooling layer; Conv represents a convolutional layer with a 3x3 kernel; Sigmoid represents a sigmoid activation layer; K i (represents the precise instance-aware parameter); mf1 (represents the first enhanced mask image).

[0109] Among them, the two linear layers are connected in parallel;

[0110] Among them, the maximum pooling layer and the average pooling layer are connected in parallel;

[0111] The input of the first main branch is the instance perception result;

[0112] The input of the second main branch is the first enhanced mask image;

[0113] The multiplier of the first main branch is connected to the output of the second main branch.

[0114] S1032: Based on the instance perception results, the spatial weighted mask is normalized for instance in the channel dimension to obtain the output model prediction instance result of the disguised object; the output model prediction instance result is the instance prediction result map.

[0115] Furthermore, the instance normalization module (FSA-IN) for fusion spatial attention operates as follows: Figure 8 As shown, spatial attention is used to focus information on high-resolution mask features to obtain more spatial information of camouflaged object instances. Then, feature weighting is performed through instance normalization to obtain the final camouflaged object instance M.

[0116] Deep neural network training, parameter initialization: For the first to fourth embedding blocks, the network parameters are initialized using the P2T_Tiny weight parameters pre-trained on the ImageNet1K dataset.

[0117] Training optimization details: The technical solution of this embodiment of the invention adopts 90K iterations of training based on the SGD optimizer, with an initial learning rate of 2.5e-4 and a batch size of 2. When the model is trained to 60K and 80K times, it is divided by 10, and the weight decay rate is 10e-4.

[0118] Table 2 presents the quantitative comparison results with six cutting-edge single-stage instance segmentation models.

[0119] Table 2 is a quantitative comparison table with the cutting-edge single-stage instance segmentation model.

[0120]

[0121] Figures 9-14 The figure shows a qualitative comparison of the technical solution of the present invention with three current cutting-edge models.

[0122] Example 2

[0123] This second embodiment provides a camouflaged object instance segmentation system based on multi-scale pooling modeling, including:

[0124] The multi-scale feature acquisition module is configured to: acquire the image to be segmented; perform supervised training on the model using camouflaged object data with instance-level labels, and obtain the extraction results of multi-scale features of the image to be segmented through the supervised training;

[0125] The multi-scale feature processing module is configured to: based on the extraction results of the multi-scale features, process the camouflaged objects in the image to be segmented using a multi-scale pooling method to obtain high-resolution mask features and instance perception results containing global information.

[0126] The model output prediction module is configured to: based on the instance perception results, focus the high-resolution mask features using an instance normalization method that integrates spatial attention, and output the model's predicted instance results.

[0127] It should be noted that the multi-scale feature acquisition module, multi-scale feature processing module, and model output prediction module mentioned above correspond to steps S101 to S103 in Embodiment 1. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0128] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and the division of modules described above is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0129] Example 3

[0130] This embodiment also provides a camouflaged object instance segmentation device based on multi-scale pooling modeling, including a processor and a storage medium; wherein, the storage medium is used to store instructions; the processor is used to perform operations according to the instructions to execute the method described in Embodiment 1 above.

[0131] It should be understood that in this embodiment, the processor can be a central processing unit (CPU) or other general-purpose processors. A general-purpose processor can be a microprocessor or any conventional processor, etc.

[0132] Example 4

[0133] This fourth embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the first embodiment above.

[0134] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A camouflage object instance segmentation method based on multi-scale pooling modeling, characterized in that, The method comprises the following steps: acquiring an image to be segmented; supervised training of the model using camouflage object data with instance-level labels, obtaining extraction results of multi-scale features of the image to be segmented through the supervised training; processing the camouflage object in the image to be segmented based on the extraction results of the multi-scale features using a multi-scale pooling method to obtain high-resolution mask features and instance-aware results containing global information; wherein, based on the instance-aware results, the high-resolution mask features are focused using an instance normalization method with fused spatial attention to output model predicted instance results; specifically comprising: inputting the high-resolution mask image into the instance normalization module with fused spatial attention for spatial weighted attention to obtain a spatial weighted mask; the instance normalization module with fused spatial attention comprises two main branches; wherein, the two main branches comprise a first main branch and a second main branch; the first main branch comprises two linear layers, a multiplier and a sixth adder connected in sequence; the second main branch comprises a maximum pooling layer and an average pooling layer connected in sequence, a splicer, a convolution layer with a convolution kernel size of 3*3, a Sigmoid activation layer and a multiplier; wherein, the two linear layers are in parallel connection; wherein, the maximum pooling layer and the average pooling layer are in parallel connection; the input end of the first main branch is the instance-aware result; the input end of the second main branch is the first enhanced mask image; the multiplier of the first main branch is connected with the output end of the second main branch; combine the instance-aware results, the spatial weighted mask is processed by instance normalization in channel dimension to obtain the output model predicted instance results of the camouflage object; the output model predicted instance results are instance prediction result images.

2. The camouflage object instance segmentation method based on multi-scale pooling modeling according to claim 1, wherein, supervised training of the model using camouflage object data with instance-level labels, obtaining extraction results of multi-scale features of the image to be segmented through the supervised training, specifically comprising: multi-scale feature extraction of the camouflage image to be segmented based on a pyramid pooling Transformer backbone network; wherein, the pyramid pooling Transformer backbone network comprises a first embedding block, a second embedding block, a third embedding block and a fourth embedding block connected in sequence; the first embedding block performs a convolution layer with a convolution kernel size of 7*7 and one P2T basic block on the image to be segmented to output a first feature map; the second embedding block performs two times down-sampling and two P2T basic blocks on the first feature map to output a second feature map; the third embedding block performs two times down-sampling and two P2T basic blocks on the second feature map to output a third feature map; the fourth embedding block performs two times down-sampling and six P2T basic blocks on the third feature map to output a fourth feature map; wherein, the P2T basic block comprises a first multi-head self-attention module based on pooling, a first adder, a first layer normalization layer, a feedforward network, a second adder and a second layer normalization layer connected in sequence.

3. The camouflage object instance segmentation method based on multi-scale pooling modeling according to claim 1, wherein, Based on the extraction result of the multi-scale feature, the camouflaged object in the to-be-segmented image is processed using a multi-scale pooling method to obtain a high-resolution mask feature and an instance perception result containing global information, specifically including: Based on the second feature map, the third feature map, and the fourth feature map, the second high-quality mask map, the third high-quality mask map, and the fourth high-quality mask map of corresponding sizes and the coarse instance perception parameter with global information are obtained by inputting to the pooling learning Transformer structure through splicing and flattening operations; The coarse instance perception parameter uses a fully connected layer to obtain a position label, uses the position label to determine the position of the instance perception parameter and calculate the confidence of the current position label, filters incorrect parameters through the confidence of the position label, and obtains accurate instance perception parameters after filtering; Based on the first feature map and the second high-quality mask map, the third high-quality mask map, and the fourth high-quality mask map, a shared high-resolution mask map is output by simultaneously inputting to a multi-scale complementary feature pooling module.

4. The camouflage object instance segmentation method based on multi-scale pooling modeling according to claim 3, characterized in that, The pooling learning Transformer structure includes an encoder, a multi-scale feature pooling operator, and a decoder. The encoder, the multi-scale feature pooling operator, and the decoder are in a sequentially connected relationship. The encoder includes a structure connected in sequence, which is a third adder, a first multi-head deformable self-attention, a fourth adder, a first normalization operation module, and a first HS-FFN structure. The multi-scale feature pooling operator is a module that sequentially performs feature restoration operations, three adaptive average pooling operations, and flattening operations. The decoder includes a structure connected in sequence, which is a second multi-head deformable self-attention, a fifth adder, a second normalization operation module, and a second HS-FFN structure. The third adder of the encoder connects the second feature map, the third feature map, the fourth feature map, and the position encoding information. The first HS-FFN structure and the second HS-FFN structure are the same. The HS-FFN structure includes convolution layers with a convolution kernel of 3*3, group normalization layers, Hardswish activation layers, and convolution layers with a convolution kernel of 3*3, which are connected in sequence.

5. The camouflage object instance segmentation method based on multi-scale pooling modeling according to claim 4, characterized in that, The internal structure of the multi-scale complementary feature pooling module includes a first input end, a second input end, a third input end, and a fourth input end, and a plurality of adders. The first input end is used to input the first feature map. The second input end is used to input the second high-quality mask map. The third input end is used to input the third high-quality mask map. The fourth input end is used to input the fourth high-quality mask map. The second high-quality mask map is processed by two times of upsampling, feature subtraction, and absolute value operation with the first feature map, and then a convolution layer with a convolution kernel of 3*3 is used to obtain the first complementary feature. The third high-quality mask map is processed by two times of upsampling, feature subtraction, and absolute value operation with the second high-quality mask map, and then a convolution layer with a convolution kernel of 3*3 is used to obtain the second complementary feature. The fourth high-quality mask image is subjected to twofold upsampling, feature subtraction and absolute value operation with the third high-quality mask image, and then is subjected to a convolution layer with a 3*3 convolution kernel to obtain a third complementary feature; The first complementary feature is connected with the first feature map through an adder, subjected to a convolution layer with a 3*3 convolution kernel to obtain a first original feature, and then input into a PPM module to obtain a first enhanced mask image; The second complementary feature is connected with the second high-quality mask image through an adder, subjected to a convolution layer with a 3*3 convolution kernel to obtain a second original feature, and then input into a PPM module to obtain a second enhanced mask image; The third complementary feature is connected with the third high-quality mask image through an adder, subjected to a convolution layer with a 3*3 convolution kernel to obtain a third original feature, and then input into a PPM module to obtain a third enhanced mask image; The fourth high-quality mask image is subjected to a convolution layer with a 3*3 convolution kernel to obtain a fourth original feature, and then input into a PPM module to obtain a fourth enhanced mask image.

6. The camouflage object instance segmentation method based on multi-scale pooling modeling according to claim 5, wherein, The PPM module comprises one residual branch and four side branches; The four side branches are in parallel with each other; The four side branches comprise a first side branch, a second side branch, a third side branch and a fourth side branch. The first side branch comprises, in sequence, a convolution layer with a 1*1 convolution kernel, an adaptive average pooling layer with a 1*1 size, and a first pooling feature obtained through an upsampling operation; The second side branch comprises, in sequence, a convolution layer with a 1*1 convolution kernel, an adaptive average pooling layer with a 2*2 size, and a second pooling feature obtained through an upsampling operation; The third side branch comprises, in sequence, a convolution layer with a 1*1 convolution kernel, an adaptive average pooling layer with a 3*3 size, and a third pooling feature obtained through an upsampling operation; The fourth side branch comprises, in sequence, a convolution layer with a 1*1 convolution kernel, an adaptive average pooling layer with a 6*6 size, and a fourth pooling feature obtained through an upsampling operation; The residual branch comprises, in sequence, a splicer and a convolution layer with a 1*1 convolution kernel. The input end of the splicer is the first pooling feature, the second pooling feature, the third pooling feature, the fourth pooling feature and the first original feature or the second original feature or the third original feature or the fourth original feature.

7. A camouflage object instance segmentation device based on multi-scale pooling modeling, characterized in that, The device comprises a processor and a storage medium; The storage medium is used for storing instructions; The processor can operate according to the instructions to perform the steps of the method of any one of claims 1-6.

8. A camouflage object instance segmentation system based on multi-scale pooling modeling, characterized in that, The method of any one of claims 1-6 comprises: a multi-scale feature acquisition module configured to: acquire an image to be segmented; use a camouflage object data with an instance-level label to supervise training of a model, and obtain an extraction result of multi-scale features of the image to be segmented through the supervised training. The multi-scale feature processing module is configured to: based on the extraction result of the multi-scale feature, process the camouflaged object in the image to be segmented by using a multi-scale pooling manner to obtain a high-resolution mask feature and an instance perception result containing global information; The model output prediction module is configured to: based on the instance perception result, focus on the high-resolution mask feature by using an instance normalization method with fusion spatial attention, and output a predicted instance result of the model.

9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which is executed by a processor to implement the steps of the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-scale aggregation cloud and cloud shadow identification method, system and device and storage medium

    CN115410081A