Camouflage Target Recognition System, Method, Medium and Device Based on Feature Aggregation

Through the feature aggregation camouflage target recognition system, the internal cascading encoder and decoder modules are used to solve the problem of global information loss in the prior art, and the precise segmentation and robust recognition of the camouflage target are achieved.

CN119850933BActive Publication Date: 2025-07-08XIAN ORDNANCE IND TECH IND DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510319158.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

现有的伪装目标识别算法在特征提取过程中损失结构信息,无法充分捕捉全局信息,导致识别效果不佳。

Method used

A camouflage object recognition system based on feature aggregation is adopted, and the feature aggregation is performed through the internal cascading encoder module and the decoder module, including the encoder module for feature extraction and amplification, and the decoder module for the aggregation of foreground and background feature maps to generate refined segmentation prediction maps.

Benefits of technology

The degree of refinement and robustness of camouflage target recognition is improved, the stability of features is enhanced, and the precise segmentation of camouflage targets is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850933B_ABST
    Figure CN119850933B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a camouflaged target recognition system, method, medium and device based on feature aggregation. The system includes an encoder module, which includes a preset number of encoders for obtaining feature vectors of an initial image; a target region amplification module, which includes a preset number of amplification modules for performing secondary convolution and batch normalization operations on the feature vectors to obtain amplified features, and the amplified features include fused amplified features and top-layer amplified features; a decoder module, which includes a preset number of multi-head self-attention modules for obtaining a foreground feature map and a background feature map of the current-level amplified features through matrix multiplication and element-wise multiplication operations, obtaining a segmentation map of the current-level amplified features through independent convolution and depthwise separable convolution operations, aggregating the foreground feature map, the background feature map and the segmentation map, and combining the amplified features output by the previous level to obtain a segmentation prediction map of the camouflaged target, thereby obtaining a more refined prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of camouflaged target recognition, and in particular, to a camouflaged target recognition system, method, medium and device based on feature aggregation. Background Art

[0002] With the development of computer vision, the camouflaged target recognition task, as a more challenging segmentation task, has received increasing attention. Camouflaged target recognition aims to segment targets hidden in the background and is applied to fields such as medical image analysis, agricultural pest detection, and military target detection. Biological research shows that because the camouflaged objects are highly similar to the surrounding environment, human visual perception is easily deceived by various camouflage strategies, which makes traditional target detection methods perform poorly in the process of recognizing camouflaged targets.

[0003] Existing camouflaged target recognition algorithms are camouflage recognition algorithms based on the CNN (Convolutional Neural Network) architecture, including: SINet (Search Identify NETwork) that searches and identifies two-stage strategies, and RefCOD (Referring Camouflaged Object Detection) that uses salient targets as references to assist in camouflaged target recognition, etc. However, recent research shows that the CNN architecture will lose structural information during the feature extraction process and the actual receptive field is smaller than the theoretical receptive field. Therefore, the camouflaged target detection method based on CNN usually cannot fully capture global information. Summary of the Invention

[0004] The embodiments of the present application provide a camouflaged target recognition system, method, medium and device based on feature aggregation to solve the problem that the camouflaged target detection method cannot fully capture global information.

[0005] In a first aspect, the embodiments of the present application provide a camouflaged target recognition system based on feature aggregation. The system includes an encoder module, a target area amplification module, and a decoder module that are cascaded internally;

[0006] The encoder module includes a preset number of encoders for obtaining feature vectors of the initial image.

[0007] The target area amplification module includes a preset number of amplification modules for performing secondary convolution and batch normalization operations on the feature vectors to obtain amplified features, where the amplified features include fused amplified features and top-layer amplified features.

[0008] The decoder module includes a preset number of multi-head self-attention modules, which are used to obtain the foreground feature map and background feature map of the amplified feature at the current level through matrix multiplication and element-wise multiplication operations, obtain the segmentation map of the amplified feature at the current level through independent convolution and depthwise separable convolution operations, aggregate the foreground feature map, the background feature map and the segmentation map, and combine the amplified feature output by the previous level to obtain the segmentation prediction map of the camouflage target.

[0009] Among them, each amplification module includes a double-branch amplification module and a single-branch amplification module;

[0010] The double-branch amplification module is used to use the feature vectors obtained by the encoder at adjacent levels as the double-branch input, respectively obtain the compressed feature vectors of each branch, perform an upsampling operation on the compressed feature vectors, splice the upsampled compressed feature vectors of the two branches, and then perform a convolution and batch normalization operation to obtain the fused amplified feature.

[0011] The single-branch amplification module is used to perform a convolution and batch normalization operation on the feature vector obtained by the encoder at the highest level, and then perform another convolution operation to obtain the highest-level amplified feature.

[0012] Among them, each decoder includes a foreground self-attention head, a background self-attention head and a standard self-attention head;

[0013] The foreground self-attention head and the background self-attention head are used to obtain prediction masks according to the amplified feature at the current level respectively. The prediction masks, the query matrix and key matrix of the amplified feature obtain the channel attention map through element-wise multiplication and matrix multiplication, and then perform matrix multiplication with the value matrix to obtain the foreground feature map and background feature map at the current level.

[0014] The standard self-attention head is used to perform independent convolution and depthwise separable convolution operations on the amplified feature at the current level to determine the multi-convolution head attention, and obtain the segmentation map at the current level according to the multi-convolution head attention.

[0015] Among them, the foreground feature map is determined according to where MA F is the foreground feature map, V F is the foreground value matrix, Q F is the foreground query matrix, T is the transpose quantity, K F is the foreground key matrix, α F is the foreground learnable parameter, and softmax() is the normalization exponential function.

[0016] The background feature map is determined according to Determine, where MA B is the background feature map, V B is the background value matrix, Q B is the background query matrix, T is the transpose quantity, K B is the background key matrix, α B is the background learnable parameter, and softmax() is the normalization exponential function.

[0017] In a second aspect, an embodiment of the present application provides a camouflaged target recognition method based on feature aggregation, and the method includes:

[0018] Obtain the feature vector of the initial image.

[0019] Perform secondary convolution and batch normalization operations on the feature vector to obtain amplified features, and the amplified features include fused amplified features and top-layer amplified features.

[0020] Through matrix multiplication and element multiplication operations, obtain the foreground feature map and background feature map of the amplified feature at the current level. Through independent convolution and depthwise separable convolution operations, obtain the segmentation map of the amplified feature at the current level. Aggregate the foreground feature map, the background feature map, and the segmentation map, and combine the amplified feature output from the previous level to obtain a segmentation prediction map of the camouflaged target.

[0021] Among them, the obtaining of the feature vector of the initial image specifically includes:

[0022] Divide the initial image into several image blocks, perform linear projection on each image block to obtain image blocks with spatial position embeddings.

[0023] Perform a feature size redefinition operation on the image blocks with spatial position embeddings to obtain the feature vector of the first level.

[0024] Perform a feature size redefinition operation on the feature vector of the first level until the preset number of feature vectors is obtained.

[0025] Among them, the performing of secondary convolution and batch normalization operations on the feature vector to obtain amplified features, and the amplified features include fused amplified features and top-layer amplified features, specifically includes:

[0026] Use the feature vectors of adjacent levels as the dual-branch input, respectively obtain the compressed feature vectors of each branch, perform an upsampling operation on the compressed feature vectors, splice the upsampled compressed feature vectors of the two branches, and then perform a convolution and batch normalization operation to obtain the fused amplified feature.

[0027] After performing a convolution and a batch normalization operation on the feature vector at the highest level, perform another convolution operation to obtain the highest-level amplified feature.

[0028] Among them, the foreground feature map and the background feature map of the amplified feature at the current level are obtained through matrix multiplication and element-wise multiplication operations, and the segmentation map of the amplified feature at the current level is obtained through independent convolution and depthwise separable convolution operations. Aggregate the foreground feature map, the background feature map, and the segmentation map, and combine the amplified feature output from the previous level to obtain the segmentation prediction map for the camouflaged target, which specifically includes:

[0029] Perform independent convolution and depthwise separable convolution operations on the amplified feature at the current level to determine the multi-convolution head attention, and obtain the segmentation map at the current level according to the multi-convolution head attention.

[0030] The amplified feature at the current level generates the prediction map at the current level under the supervision of the initial input mask according to the prediction loss function.

[0031] Perform a convolution operation on the prediction map at the current level to obtain the prediction mask.

[0032] After normalizing the amplified feature at the current level, obtain the query matrix, the key matrix, and the value matrix again through independent convolution and depthwise separable convolution operations.

[0033] The prediction mask at the current level is respectively multiplied element-wise with the query matrix and the key matrix to obtain the query matrix mapping and the key matrix mapping.

[0034] After the key matrix mapping is transposed and multiplied with the query matrix mapping, obtain the channel attention map at the current level.

[0035] After the channel attention map at the current level is multiplied with the value matrix, obtain the foreground feature map and the background feature map at the current level.

[0036] Add the segmentation map at the current level to the foreground feature map and the background feature map to obtain the global interaction feature at the current level.

[0037] Aggregate the global interaction feature at the current level with the fused and amplified feature at the previous level to obtain the segmentation prediction map for the camouflaged target.

[0038] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method described above.

[0039] Fourthly, an embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method as described above.

[0040] Adopting the embodiment of the present invention has the following beneficial effects:

[0041] The present invention extracts image features layer by layer through an internally cascaded encoder module to capture multi-scale features from coarse to fine. Further, the feature vectors output by the encoder module are further processed to enhance the feature representation of the target region. Through secondary convolution and batch normalization operations, not only are the features refined, but also the stability and robustness of the features are improved. Finally, through an internally cascaded decoder module with a multi-head self-attention module, the aggregated features of the magnified features of the previous level and the foreground feature map, background feature map, and segmentation map processed by the multi-head self-attention module at each layer are used to generate a segmentation prediction map for the camouflaged target and assist in the foreground region processing and background region processing through the segmentation prediction map, realizing the improvement of the feature prediction results at each level and obtaining a more refined prediction result. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Among them:

[0044] Figure 1 It is a schematic structural diagram of an embodiment of a camouflaged target recognition system based on feature aggregation provided by the present invention;

[0045] Figure 2 It is a schematic structural diagram of another embodiment of a camouflaged target recognition system based on feature aggregation provided by the present application;

[0046] Figure 3 It is a schematic structural diagram of an embodiment of a target region magnification module provided by the present invention;

[0047] Figure 4 It is a schematic structural diagram of an embodiment of a foreground self-attention module provided by the present invention;

[0048] Figure 5 It is a schematic structural diagram of another embodiment of a camouflaged target recognition system based on feature aggregation provided by the present invention;

[0049] Figure 6 Schematic flowchart of an embodiment of a camouflaged target recognition method based on feature aggregation provided by the present invention;

[0050] Figure 7 Schematic flowchart of another embodiment of a camouflaged target recognition method based on feature aggregation provided by the present invention;

[0051] Figure 8 Schematic structural diagram of an embodiment of the device provided by the present invention;

[0052] Figure 9 Schematic structural diagram of an embodiment of the medium provided by the present invention. Detailed implementation manners

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] As Figure 1 shown, Figure 1 Schematic structural diagram of an embodiment of a camouflaged target recognition system based on feature aggregation provided by the present invention. A camouflaged target recognition system based on feature aggregation, the system includes an internally cascaded encoder module 11, a target area amplification module 12, and a decoder module 13.

[0055] The encoder module 11 includes a preset number of encoders for obtaining the feature vectors of the initial image.

[0056] Exemplarily, with reference to Figure 2 , Figure 2 Schematic structural diagram of another embodiment of a camouflaged target recognition system provided by the present application. The encoder module 11 includes four cascaded ViT (Vision Transformer) encoders, which are the first layer, the second layer, the third layer, and the fourth layer from top to bottom, and the fourth layer is the highest layer. Taking the ViT encoder of the first layer as an example, through this ViT encoder for feature extraction, using an initial image with a size of as the input, generating image patches with a size of , linearly projecting the generated image patches to obtain image patches with a size of and a channel number of with spatial position embeddings; further, performing a feature size redefinition operation on the image patches with spatial position embeddings to obtain image patches with a size of The feature vector of the first layer.

[0057] And so on, input the feature vector output by the encoder of the first layer into the encoder of the second layer to perform the feature size redefinition operation, obtain the feature vector of the second layer, input the feature vector of the second layer into the encoder of the third layer to perform the feature size redefinition operation, obtain the feature vector of the third layer, input the feature vector of the third layer into the encoder of the fourth layer to perform the feature size redefinition operation, and obtain the feature vector of the fourth layer. At this time, four feature vectors are obtained.

[0058] The target region amplification module 12 includes a preset number of amplification modules, which are used to perform secondary convolution and batch normalization operations on the feature vectors to obtain amplified features. The amplified features include fused amplified features and top-layer amplified features.

[0059] Exemplarily, each amplification module includes a dual-branch amplification module and a single-branch amplification module. The dual-branch amplification module is used to take the feature vectors obtained by the encoders of adjacent layers as the dual-branch input, respectively obtain the compressed feature vectors of each branch, perform upsampling on the compressed feature vectors, splice the upsampled compressed feature vectors of the two branches, and then perform one convolution and batch normalization operation to obtain the fused amplified features; the single-branch amplification module is used to perform one convolution and batch normalization operation on the feature vector obtained by the encoder of the highest layer, and then perform another convolution operation to obtain the top-layer amplified features.

[0060] Specifically, with reference to Figure 2 and Figure 3 , Figure 3 is a schematic structural diagram of an embodiment of the target region amplification module provided by the present invention. The target region amplification module 12 includes 4 amplification modules arranged in cascade. In the amplification stage, the feature vectors of the third layer and the fourth layer are used as the two branches input to the amplification module at the third layer. First, perform one convolution operation on the feature vectors of each branch, apply batch normalization BN and GeLU activation function to the convolved features to compress the feature vectors and obtain the compressed feature vectors; perform the upsampling operation UpSample of bicubic interpolation on the compressed feature vectors to complete the scale amplification of the feature vectors; splice (element-wise addition) the compressed feature vectors with compressed channels and amplified scales of the two branches to complete the feature fusion of adjacent layers, and then perform one convolution operation on the fused features, apply batch normalization BN and GeLU activation function to the convolved features to further compress the number of channels, complete the fusion and amplification of the feature vectors of the third layer and the fourth layer, and obtain the fused amplified features of the third layer.

[0061] And so on, the same operation is performed on the feature vectors of the first level and the second level as the two branches of the input of the amplification module at the first level, and the feature vectors of the second level and the third level as the two branches of the input of the amplification module at the second level, to complete the fusion and amplification of the feature vectors of adjacent levels of all features, and obtain the fusion and amplification features of the first level and the fusion and amplification features of the second level.

[0062] Meanwhile, perform a single convolution operation on the feature vectors of the fourth level, apply batch normalization and ReLU activation function to the convolved features, and then perform another convolution operation to obtain the highest-level amplified features of the fourth level.

[0063] Decoder module 13, including a preset number of multi-head self-attention modules, is used to obtain the foreground feature map and background feature map of the amplified features at the current level through matrix multiplication and element-wise multiplication operations, obtain the segmentation map of the amplified features at the current level through independent convolution and depthwise separable convolution operations, aggregate the foreground feature map, background feature map and segmentation map, and combine the amplified features output by the previous level to obtain the segmentation prediction map of the camouflage target.

[0064] Exemplarily, each multi-head self-attention module includes a foreground self-attention head, a background self-attention head and a standard self-attention head. The foreground self-attention head and the background self-attention head are used to obtain prediction masks respectively according to the amplified features at the current level. The query matrix, key matrix of the prediction mask and the amplified features obtain the channel attention map through element-wise multiplication and matrix multiplication, and then perform matrix multiplication with the value matrix to obtain the foreground feature map and background feature map at the current level; the standard self-attention head is used to perform independent convolution and depthwise separable convolution operations on the amplified features at the current level to determine the multi-convolution head attention, and obtain the segmentation map at the current level according to the multi-convolution head attention.

[0065] Specifically, with reference to Figure 2 , taking the highest-level amplified features of the fourth level as an example, the highest-level amplified features of the fourth level are used as the input of the multi-head self-attention module at the fourth level. In the standard self-attention head, independent convolution and depthwise separable convolution operations are performed on the highest-level amplified features input to the multi-head self-attention module at the highest level to determine the multi-convolution head attention, and the segmentation map at the current level is obtained according to the multi-convolution head attention.

[0066] Meanwhile, under the supervision of the initial input mask GT, the highest-level amplified features output by the multi-head self-attention module at the highest level (the fourth level) are used with BCE (Binary Cross-Entropy) and IoU (Intersection over Union) as the prediction loss functions to generate the prediction map of the fourth level. Thus, the prediction loss function can be obtained as shown in the following formula:

[0067] ;

[0068] where Loss is the prediction loss function, i is the number of levels, GT is the initial input mask, BCE is the binary cross-entropy, IoU is the intersection over union, and P i is the prediction map of the i-th level, and loss bce () is the binary cross-entropy loss function, and loss iou () is the intersection over union loss function.

[0069] After performing a convolution operation on the prediction map of the fourth level, a prediction mask P4 of the fourth level with a size of is generated and input into the foreground self-attention head.

[0070] It should be noted that since the foreground self-attention head and the background self-attention head have similar structures, the foreground self-attention head will be taken as an example for illustration.

[0071] Specifically, as Figure 4 shown, Figure 4 is a schematic structural diagram of an embodiment of the foreground self-attention module provided by the present invention. After the highest-level amplified features of the fourth level are input into the foreground self-attention head, they are normalized to generate a feature map with a size of . Further, through three independent convolution and depthwise separable convolution operations , a foreground query matrix, a foreground key matrix, and a foreground value matrix are generated; the prediction mask P4 of the fourth level is respectively multiplied element-wise with the foreground query matrix and the foreground key matrix to obtain a foreground query matrix mapping and a foreground key matrix mapping with a scale size of and a channel number of . At this time, the foreground query matrix mapping and the foreground key matrix mapping have been transposed once. After transposing the foreground key matrix mapping, it is restored to the foreground key matrix before transposition and multiplied with the foreground query matrix mapping to obtain a channel attention map with a scale size of . After multiplying the channel attention map with the foreground value matrix, a result with a size of The foreground feature map processed from the foreground region. Similarly, the background self-attention head also takes the magnified feature of the top layer of the fourth level as input and performs similar operations to obtain the background feature map processed from the background region.

[0072] Furthermore, the features output by the three attention heads are added together to obtain the global interaction feature of the fourth level with global interaction for the foreground and background regions. The global interaction feature of the fourth level and the fused magnified feature of the previous level (the third level) are aggregated and then input into the multi-head self-attention module of the previous level. At the same time, the fused magnified feature of the third level generates a prediction map of the third level according to the prediction loss function under the supervision of the initial input mask GT. A convolution operation is performed on the prediction map of the third level to obtain the prediction mask P3 of the third level. This prediction mask serves as the input for the foreground self-attention head and the background self-attention head in the multi-head self-attention module of the third level. The remaining process is the same as the processing process of the above multi-head self-attention module. By analogy, through the cascaded multi-head self-attention module, using the aggregated feature of the magnified feature of each layer and the feature processed by the multi-head self-attention module of the previous layer, a segmentation prediction map of the camouflage target is generated.

[0073] From the above description, it can be seen that the present invention extracts image features layer by layer through an internally cascaded encoder module to capture multi-scale features from coarse to fine. Further, the feature vectors output by the encoder module are further processed to enhance the feature representation of the target region. Through secondary convolution and batch normalization operations, not only are the features refined, but also the stability and robustness of the features are improved. Finally, through the internally cascaded decoder module with a multi-head self-attention module, using the aggregated feature of the magnified feature of the previous level and the foreground feature map, background feature map, and segmentation map processed by the multi-head self-attention module of each layer, a segmentation prediction map of the camouflage target is generated, and the foreground region processing and background region processing are assisted through the segmentation prediction map, realizing the improvement of the feature prediction results of each level and obtaining a more refined prediction result.

[0074] As Figure 5 shown, Figure 5 is a schematic structural diagram of another embodiment of a camouflage target recognition system based on feature aggregation provided by the present invention. The multi-head self-attention module includes a foreground self-attention head, a background self-attention head, and a standard self-attention head.

[0075] The foreground self-attention head and the background self-attention head are used to obtain the prediction mask respectively according to the magnified feature of the current level. The query matrix, key matrix of the prediction mask and the magnified feature obtain the channel attention map through element-wise multiplication and matrix multiplication, and then perform matrix multiplication with the value matrix to obtain the foreground feature map and background feature map of the current level.

[0076] Exemplarily, with reference to Figure 2 and Figure 4 , taking the highest-level magnification feature of the fourth level as an example, the highest-level magnification feature of the fourth level is used as the input of the multi-head self-attention module of the fourth level. After the highest-level magnification feature of the fourth level is input into the foreground self-attention head, it is normalized to generate a feature map with a size of . Further, through three independent convolutions and depthwise separable convolutions , a foreground query matrix, a foreground key matrix, and a foreground value matrix are generated; the highest-level magnification feature generates a prediction map of the fourth level under the supervision of the initial input mask GT according to the prediction loss function; a convolution operation is performed on the prediction map of the fourth level to obtain a prediction mask P4 of the fourth level; the prediction mask P4 of the fourth level is element-wise multiplied with the foreground query matrix and the foreground key matrix respectively to obtain a foreground query matrix mapping and a foreground key matrix mapping with a scale size of and a channel number of . At this time, the foreground query matrix mapping and the foreground key matrix mapping have been transposed once. After transposing the foreground key matrix mapping, it is restored to the foreground key matrix before transposing and multiplied with the foreground query matrix mapping to obtain a channel attention map with a scale size of . After the channel attention map is multiplied with the foreground value matrix, a foreground feature map processed for the foreground region with a size of is obtained. Similarly, the background self-attention head also uses the highest-level magnification feature of the fourth level as the input to perform similar operations, and a background feature map processed for the background region can be obtained.

[0077] Specifically, the foreground feature map is determined according to the following formula:

[0078] ;

[0079] where MA F is the foreground feature map, V F is the foreground value matrix, Q F is the foreground query matrix, T is the transpose amount, K F is the foreground key matrix, α F is the foreground learnable parameter, and softmax() is the normalization exponential function.

[0080] The background feature map is determined according to the following formula:

[0081] ;

[0082] where MA B is the background feature map, V Bis the background value matrix, Q B is the background query matrix, T is the transpose quantity, K B is the background key matrix, α B is the background learnable parameter, and softmax() is the normalization exponential function.

[0083] The standard self-attention head is used to perform independent convolution and depthwise separable convolution operations on the magnified features at the current level to determine the multi-convolution head attention, and obtain the segmentation map of the current level according to the multi-convolution head attention.

[0084] Exemplarily, taking the magnified features of the highest layer of the fourth level as an example, the magnified features of the highest layer of the fourth level are used as the input of the multi-head self-attention module of the fourth level.

[0085] It should be noted that the standard self-attention head is extended to obtain the foreground self-attention head and the background self-attention head, and the structures of the foreground self-attention head and the background self-attention are similar. For the standard self-attention head, it can be defined as:

[0086] ;

[0087] where MA is the multi-convolution head attention, Q is the query matrix, T is the transpose quantity, K is the key matrix, V is the value matrix, and the query matrix, key matrix, and value matrix are generated through independent convolution and depthwise separable convolution operations, and α is the learnable parameter.

[0088] Specifically, in the standard self-attention head, independent convolution and depthwise separable convolution operations are performed on the magnified features of the highest layer input to the decoder at the fourth level to determine the multi-convolution head attention, and the segmentation map of the fourth level is obtained according to the multi-convolution head attention.

[0089] As Figure 6 shown, Figure 6 is a schematic flowchart of an embodiment of a camouflaged target recognition method based on feature aggregation provided by the present invention. A camouflaged target recognition method based on feature aggregation, the method includes:

[0090] S101: Obtain the feature vector of the initial image.

[0091] Exemplarily, the initial image is divided into several image patches, linear projection is performed on each image patch to obtain image patches with spatial position embeddings; feature size redefinition operations are performed on the image patches with spatial position embeddings to obtain the feature vectors of the first level.

[0092] Feature size redefinition operations are performed on the feature vectors of the first level until the preset number of feature vectors is obtained.

[0093] S102: Perform quadratic convolution and batch normalization operations on the feature vectors to obtain amplified features, where the amplified features include fused amplified features and top-layer amplified features.

[0094] Exemplarily, use the feature vectors of adjacent levels as the dual-branch input, respectively obtain the compressed feature vectors of each branch, perform upsampling on the compressed feature vectors, splice the upsampled compressed feature vectors of the two branches, and then perform a convolution and batch normalization operation to obtain the fused amplified features. After performing a convolution and batch normalization operation on the feature vectors of the top level, perform another convolution operation to obtain the top-layer amplified features.

[0095] S103: Through matrix multiplication and element-wise multiplication operations, obtain the foreground feature map and background feature map of the amplified features at the current level. Through independent convolution and depthwise separable convolution operations, obtain the segmentation map of the amplified features at the current level. Aggregate the foreground feature map, background feature map, and segmentation map, and combine with the amplified features output from the previous level to obtain the segmentation prediction map for the camouflaged target.

[0096] Exemplarily, perform independent convolution and depthwise separable convolution operations on the amplified features at the current level to determine the multi-convolution head attention. Obtain the segmentation map at the current level according to the multi-convolution head attention. The amplified features at the current level generate the prediction map at the current level under the supervision of the initial input mask GT according to the prediction loss function. Perform a convolution operation on the prediction map at the current level to obtain the prediction mask. After normalizing the amplified features at the current level, obtain the query matrix, key matrix, and value matrix again through independent convolution and depthwise separable convolution operations. The prediction mask at the current level performs element-wise multiplication with the query matrix and key matrix respectively to obtain the query matrix mapping and key matrix mapping. At this time, the query matrix mapping and key matrix mapping have been transposed once. After the key matrix mapping is transposed, it is restored to the key matrix before transposition and multiplied with the query matrix mapping to obtain the channel attention map at the current level. After multiplying the channel attention map at the current level with the value matrix, obtain the foreground feature map and background feature map at the current level. Add the segmentation map at the current level to the foreground feature map and background feature map to obtain the global interaction feature at the current level. Aggregate the global interaction feature at the current level with the fused amplified features of the previous level to obtain the segmentation prediction map for the camouflaged target.

[0097] As Figure 7 shown, Figure 7 is a schematic flowchart of another embodiment of a camouflaged target recognition method based on feature aggregation provided by the present invention. A camouflaged target recognition method based on feature aggregation, the method includes:

[0098] S201: Divide the initial image into a number of image patches, perform linear projection on each image patch, and obtain image patches with spatial position embeddings.

[0099] Exemplarily, with reference to Figure 2 , the encoder module includes four cascaded ViT (Vision Transformer) encoders. Taking the ViT encoder of the first layer as an example, feature extraction is performed through this ViT encoder. Specifically, using an initial image with a size of as input, image patches with a size of are generated. Further, the generated image patches are linearly projected to obtain image patches with spatial position embeddings having a size of and a channel number of .

[0100] S202: Perform an operation to redefine the feature size on the image patches with spatial position embeddings to obtain feature vectors of the first layer.

[0101] Exemplarily, perform an operation to redefine the feature size on the image patches with spatial position embeddings to obtain feature vectors of the first layer with a size of .

[0102] S203: Perform an operation to redefine the feature size on the feature vectors of the first layer until a preset number of feature vectors are obtained.

[0103] Exemplarily, use the feature vectors of the first layer as the input to the encoder of the second layer and perform an operation to redefine the feature size to obtain feature vectors of the second layer, and so on. Input the feature vectors output by the encoder of the previous layer into the encoder of the next layer to perform an operation to redefine the feature size until four feature vectors are obtained.

[0104] S204: Use the feature vectors of adjacent layers as the double-branch input, respectively obtain the compressed feature vectors of each branch, perform an upsampling operation on the compressed feature vectors, splice the upsampled compressed feature vectors of the two branches, and then perform a convolution and batch normalization operation to obtain the fused and amplified feature.

[0105] Exemplarily, with reference to Figure 2 and Figure 3 , the target region amplification module includes 4 cascaded amplification modules. In the amplification stage, use the feature vectors of the third layer and the fourth layer as the two branches of the input to the amplification module of the third layer, and perform a convolution operation , apply batch normalization BN and GeLU activation function to the convolved features to compress the feature vectors and obtain compressed feature vectors; perform upsampling operation of bicubic interpolation on the compressed feature vectors to complete the scale amplification of the feature vectors; after splicing the compressed feature vectors with compressed channels and amplified scales from the two branch channels, perform another convolution operation , apply batch normalization BN and GeLU activation function to the convolved features to further compress the number of channels, complete the fusion and amplification of the feature vectors at the third level and the fourth level, and obtain the fused and amplified features at the third level.

[0106] And so on, perform the same operations on the two branches where the feature vectors at the first level and the second level are used as the inputs of the amplification module at the first level, and the feature vectors at the second level and the third level are used as the inputs of the amplification module at the second level, to complete the fusion and amplification of all features, and obtain the fused and amplified features at the first level and the second level.

[0107] S205: After performing a convolution and batch normalization operation on the feature vectors at the highest level, perform another convolution operation to obtain the highest-level amplified features.

[0108] Exemplarily, perform a single convolution operation on the feature vectors at the fourth level, apply batch normalization and ReLU activation function to the convolved features, and then perform another convolution operation to obtain the highest-level amplified features at the fourth level.

[0109] S206: Perform independent convolution and depthwise separable convolution operations on the amplified features at the current level to determine multi-convolution head attention, and obtain the segmentation map at the current level according to the multi-convolution head attention.

[0110] Exemplarily, each multi-head self-attention module includes a foreground self-attention head, a background self-attention head, and a standard self-attention head. Taking the highest-level amplified features at the fourth level as an example, use the highest-level amplified features at the fourth level as the input of the multi-head self-attention module at the fourth level. In the standard self-attention head, perform independent convolution and depthwise separable convolution operations on the highest-level amplified features to determine multi-convolution head attention, and obtain the segmentation map at the current level according to the multi-convolution head attention.

[0111] S207: The amplified features at the current level generate the prediction map at the current level under the supervision of the initial input mask according to the prediction loss function.

[0112] Exemplarily, the highest-level amplified feature output by the multi-head self-attention module at the fourth level generates a prediction map for the fourth level under the supervision of the initial input mask GT with BCE (Binary Cross-Entropy) and IoU (Intersection over Union) as the prediction loss functions. Thus, the prediction loss function can be obtained as shown in the following formula:

[0113] ;

[0114] where Loss is the prediction loss function, i is the number of levels, GT is the initial input mask, BCE is Binary Cross-Entropy, IoU is Intersection over Union, and P i is the prediction map for the i-th level.

[0115] S208: Perform a convolution operation on the prediction map of the current level to obtain a prediction mask.

[0116] Exemplarily, after performing a convolution operation on the prediction map of the fourth level, a fourth-level prediction mask P4 with a size of is generated and input into the foreground self-attention head and the background self-attention head.

[0117] S209: Normalize the amplified feature of the current level and then obtain the query matrix, key matrix, and value matrix through independent convolution and depthwise separable convolution operations again.

[0118] Exemplarily, after the highest-level amplified feature of the fourth level is input into the foreground self-attention head, it is normalized to generate a feature map with a size of . Further, through three independent convolutions and depthwise separable convolutions, a foreground query matrix, a foreground key matrix, and a foreground value matrix are generated.

[0119] After the highest-level amplified feature of the fourth level is input into the background self-attention head, it is normalized to generate a feature map with a size of . Further, through three independent convolutions and depthwise separable convolutions, a background query matrix, a background key matrix, and a background value matrix are generated.

[0120] S210: Element-wise multiply the prediction mask of the current level with the query matrix and the key matrix respectively to obtain the query matrix mapping and the key matrix mapping.

[0121] Exemplarily, element-wise multiply the prediction mask P4 of the fourth level with the foreground query matrix and the foreground key matrix respectively to obtain a scale size of , and the number of channels is The foreground query matrix mapping and foreground key matrix mapping, at this time, the foreground query matrix mapping and foreground key matrix mapping have undergone a matrix transpose once.

[0122] Multiply the prediction mask P4 of the fourth level element-wise with the background query matrix and the background key matrix respectively to obtain a scale size of and the number of channels is of the background query matrix mapping and background key matrix mapping. At this time, the foreground query matrix mapping and foreground key matrix mapping have undergone a matrix transpose once.

[0123] S211: After the key matrix mapping undergoes a matrix transpose, perform matrix multiplication with the query matrix mapping to obtain the channel attention map of the current level.

[0124] Exemplarily, after the foreground key matrix mapping undergoes a matrix transpose, restore it to the foreground key matrix before transpose and perform matrix multiplication with the foreground query matrix mapping to obtain a scale size of of the foreground channel attention map. After the background key matrix mapping undergoes a matrix transpose, restore it to the background key matrix before transpose and perform matrix multiplication with the background query matrix mapping to obtain a scale size of of the background channel attention map.

[0125] S212: After the channel attention map of the current level undergoes matrix multiplication with the value matrix, obtain the foreground feature map and background feature map of the current level.

[0126] Exemplarily, after the foreground channel attention map undergoes matrix multiplication with the foreground value matrix, obtain a size of of the foreground feature map processed for the foreground region. After the background channel attention map undergoes matrix multiplication with the background value matrix, obtain a size of of the background feature map processed for the background region.

[0127] S213: Add the segmentation map of the current level to the foreground feature map and background feature map to obtain the global interaction feature of the current level.

[0128] Exemplarily, add the segmentation map of the current level to the foreground feature map and background feature map to obtain the global interaction feature of the fourth level with global interaction processed for the foreground region and background region.

[0129] S214: Aggregate the global interaction feature of the current level with the fusion and amplification feature of the previous level to obtain the segmentation prediction map for the camouflaged target.

[0130] Exemplarily, after aggregating the global interaction features of the fourth level and the fusion and amplification features of the third level, the result is input into the multi-head self-attention module of the third level. Meanwhile, under the supervision of the initial input mask GT, the fusion and amplification features of the third level generate a prediction map of the third level according to the prediction loss function. A convolution operation is performed on the prediction map of the third level to obtain the prediction mask P3 of the third level. This prediction mask serves as the input to the foreground self-attention head and the background self-attention head in the multi-head self-attention module of the third level. The remaining process is the same as the processing process of the above multi-head self-attention module.

[0131] As Figure 8 shown, Figure 8 FIG. is a schematic structural diagram of an embodiment of the device provided by the present invention. The device 20 includes a memory 21 and a processor 22. The memory 21 stores a computer program, and the processor 22 executes the computer program during operation to implement the method as Figure 6 and Figure 7 shown.

[0132] The specific technical details of a method for identifying camouflaged targets based on feature aggregation implemented when the above device 20 executes the computer program have been described in detail in the foregoing method steps, so they will not be elaborated here.

[0133] As Figure 9 shown, Figure 9 FIG. is a schematic structural diagram of an embodiment of the medium provided by the present invention. The medium 30 stores at least one computer program 31, and the computer program 31 is executed by the processor 22 to implement the method as Figure 6 and Figure 7 shown. The detailed method can be referred to the above, and will not be elaborated here. In one embodiment, the medium 30 may be a storage chip, a hard disk, a mobile hard disk, a USB flash drive, an optical disc, or other writable and readable storage tools, or a server, etc.

[0134] In addition, the processes depicted in the drawings do not necessarily have to be in the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0135] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer-readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0136] The apparatus, device, non-volatile computer-readable storage medium, and method provided in the embodiments of this specification are corresponding. Therefore, the apparatus, device, and non-volatile computer storage medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device, and non-volatile computer storage medium will not be elaborated here.

[0137] The systems, apparatuses, modules, or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0138] For convenience of description, when describing the above apparatus, it is divided into various units according to functions for separate description. Of course, when implementing this specification, the functions of each unit may be implemented in the same or multiple software and / or hardware. Those skilled in the art should understand that the embodiments of this specification may be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0139] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate an apparatus for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0140] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction apparatus that implements the functions in the flow Figure 1one or more processes and / or blocks Figure 1 the functions specified in one or more blocks.

[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 or more processes and / or the functions specified in one or more blocks.

[0142] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0143] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0144] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0145] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0146] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. This specification may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.

[0147] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, they are described relatively simply, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0148] The foregoing disclosure is only for the preferred embodiments of the present invention, and of course cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope covered by the present invention.

Claims

1. A camouflage target recognition system based on feature aggregation, characterized in that, The system includes an internally cascaded encoder module, a target region amplification module, and a decoder module; The encoder module includes a preset number of encoders for obtaining the feature vectors of the initial image; The target region amplification module includes a preset number of amplification modules for performing secondary convolution and batch normalization operations on the feature vectors to obtain amplified features, where the amplified features include fused amplified features and top-layer amplified features; each amplification module includes a dual-branch amplification module and a single-branch amplification module; the dual-branch amplification module is used to take the feature vectors obtained by the encoders at adjacent levels as dual-branch inputs, respectively obtain the compressed feature vectors of each branch, perform an upsampling operation on the compressed feature vectors, splice the upsampled compressed feature vectors of the two branches, and then perform a convolution and a batch normalization operation to obtain the fused amplified features; the single-branch amplification module is used to perform a convolution and a batch normalization operation on the feature vectors obtained by the encoder at the highest level, and then perform another convolution operation to obtain the top-layer amplified features; The decoder module includes a preset number of multi-head self-attention modules for obtaining the foreground feature map and the background feature map of the amplified features at the current level through matrix multiplication and element-wise multiplication operations, obtaining the segmentation map of the amplified features at the current level through independent convolution and depthwise separable convolution operations, aggregating the foreground feature map, the background feature map, and the segmentation map, and combining the amplified features output at the previous level to obtain the segmentation prediction map of the camouflage target; each multi-head self-attention module includes a foreground self-attention head, a background self-attention head, and a standard self-attention head; the foreground self-attention head and the background self-attention head are used to respectively obtain prediction masks according to the amplified features at the current level, obtain the channel attention map through element-wise multiplication and matrix multiplication of the prediction masks and the query matrix and the key matrix of the amplified features, and then perform matrix multiplication with the value matrix to obtain the foreground feature map and the background feature map at the current level; The standard self-attention head is used to perform independent convolution and depthwise separable convolution operations on the amplified features at the current level to determine the multi-convolution head attention, and obtain the segmentation map at the current level according to the multi-convolution head attention.

2. The camouflage target recognition system based on feature aggregation according to claim 1, wherein The foreground feature map is determined according to , where MA F is the foreground feature map, V F is the foreground value matrix, Q F is the foreground query matrix, T is the transpose quantity, K F is the foreground key matrix, α F is the foreground learnable parameter, and softmax() is the normalization exponential function; The background feature map is determined according to , where MA B is the background feature map, V B is the background value matrix, Q B is the background query matrix, T is the transpose quantity, K B is the background key matrix, α B is the background learnable parameter, and softmax() is the normalization exponential function.

3. A method for identifying camouflaged targets based on feature aggregation, characterized in that The method includes: Obtaining the feature vectors of the initial image; Performing secondary convolution and batch normalization operations on the feature vectors to obtain amplified features, where the amplified features include fused amplified features and top-layer amplified features, specifically including: taking the feature vectors at adjacent levels as dual-branch inputs, respectively obtaining the compressed feature vectors of each branch, performing an upsampling operation on the compressed feature vectors, splicing the upsampled compressed feature vectors of the two branches, and then performing a convolution and a batch normalization operation to obtain the fused amplified features; performing a convolution and a batch normalization operation on the feature vectors at the highest level, and then performing another convolution operation to obtain the top-layer amplified features; Obtain the foreground feature map and background feature map of the amplified feature at the current level through matrix multiplication and element-wise multiplication operations, obtain the segmentation map of the amplified feature at the current level through independent convolution and depthwise separable convolution operations, aggregate the foreground feature map, the background feature map, and the segmentation map, and combine the amplified feature output from the previous level to obtain the segmentation prediction map for the camouflaged target, including: obtaining a prediction mask according to the amplified feature at the current level, obtaining a channel attention map through element-wise multiplication and matrix multiplication of the prediction mask and the query matrix and key matrix of the amplified feature, and then performing matrix multiplication with the value matrix to obtain the foreground feature map and the background feature map at the current level; performing independent convolution and depthwise separable convolution operations on the amplified feature at the current level to determine multi-convolution head attention, and obtaining the segmentation map at the current level according to the multi-convolution head attention.

4. The method for identifying a camouflaged target based on feature aggregation according to claim 3, wherein The obtaining of the feature vector of the initial image specifically includes: Divide the initial image into a plurality of image patches, perform linear projection on each of the image patches to obtain image patches with spatial position embeddings; Perform a feature size redefinition operation on the image patches with spatial position embeddings to obtain the feature vector at the first level; Perform a feature size redefinition operation on the feature vector at the first level until the preset number of feature vectors is obtained.

5. The method for identifying a camouflaged target based on feature aggregation according to claim 3, wherein The obtaining of the foreground feature map and background feature map of the amplified feature at the current level through matrix multiplication and element-wise multiplication operations, obtaining the segmentation map of the amplified feature at the current level through independent convolution and depthwise separable convolution operations, aggregating the foreground feature map, the background feature map, and the segmentation map, and combining the amplified feature output from the previous level to obtain the segmentation prediction map for the camouflaged target specifically further includes: The amplified feature at the current level generates a prediction map at the current level according to the prediction loss function under the supervision of the initial input mask; Perform a convolution operation on the prediction map at the current level to obtain a prediction mask; Normalize the amplified feature at the current level and then obtain the query matrix, key matrix, and value matrix through independent convolution and depthwise separable convolution operations again; Perform element-wise multiplication of the prediction mask at the current level with the query matrix and the key matrix respectively to obtain a query matrix mapping and a key matrix mapping; After performing matrix transpose on the key matrix mapping, perform matrix multiplication with the query matrix mapping to obtain the channel attention map at the current level; After performing matrix multiplication of the channel attention map at the current level with the value matrix, obtain the foreground feature map and background feature map at the current level; Add the segmentation map at the current level to the foreground feature map and the background feature map to obtain the global interaction feature at the current level; Aggregate the global interaction feature at the current level with the fused amplified feature from the previous level to obtain the segmentation prediction map for the camouflaged target.

6. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a processor, the processor executes the steps of the method according to any one of claims 3 to 5.

7. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, it causes the processor to execute the steps of the method according to any one of claims 3 to 5.

Citation Information

Patent Citations

  • Camouflage target detection method based on multi-scale context and multi-level feature interaction of three-dimensional attention

    CN116740479A

  • Camouflage target detection method based on feature fusion and attention mechanism

    CN117576411A

  • Camouflage object detection method and device

    CN119273981A