Camouflage target detection method and device based on feature enhancement and readable medium

By introducing multi-scale feature extraction modules and feature enhancement modules in the camouflage object detection model, and using context aggregation and feature decoupling modules, the problem of poor camouflage object detection in the prior art is solved, and more efficient camouflage feature extraction and background suppression are achieved.

CN120219716APending Publication Date: 2025-06-27XIAMEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510296486.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing camouflage object detection methods are not effective when dealing with high-similar environments, making it difficult to effectively extract camouflage features, resulting in unsatisfactory detection results.

Method used

Using a camouflage object detection method based on feature enhancement, the camouflage features in the decoder are explicitly decoupled to weaken background interference by introducing a multi-scale feature extraction module and feature enhancement module in the encoder, and using a context aggregation module and feature decoupling module in the decoder.

Benefits of technology

It significantly improves the accuracy and effectiveness of camouflage object detection, and can more effectively extract and separate camouflage features, reduce background interference, and improve detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219716A_ABST
    Figure CN120219716A_ABST
Patent Text Reader

Abstract

The invention discloses a camouflage target detection method and device based on feature enhancement, and a readable medium, and the method comprises the steps: obtaining a to-be-detected camouflage target image, inputting the to-be-detected camouflage target image into a trained camouflage target detection model, obtaining multi-scale features through a multi-scale feature extraction module in an encoder, and carrying out the detection of the camouflage target image; wherein the second-level feature, the third-level feature and the fourth-level feature respectively pass through three feature enhancement modules and a context aggregation module to obtain a second aggregation feature, a third aggregation feature and a fourth aggregation feature; performing up-sampling on the second aggregation feature, performing element-by-element addition on the second aggregation feature and the first-level feature to obtain a first addition result, and processing the first addition result through a first convolution module to obtain a first mask; the first mask, the second aggregation feature, the third aggregation feature and the fourth aggregation feature are input into a feature decoupling module to obtain a second background suppression feature; and obtaining a corresponding prediction detection result based on the second background suppression feature, and solving the problem of high similarity of the foreground and the background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a camouflaged target detection method, device and readable medium based on feature enhancement. Background Art

[0002] Camouflage is a survival strategy that reduces the probability of being detected. By changing their appearance, color or pattern to mimic the surrounding environment, camouflaged objects can blend seamlessly with the environment. On the contrary, the task of camouflaged object detection (COD) is to segment these camouflaged objects, which has many practical applications in fields such as biology, agriculture, marine animal segmentation and medical analysis. Therefore, COD has received increasing research attention. A recent biological study has shown that the human visual perception system is easily deceived by camouflaged objects because of their high intrinsic similarity to the surrounding environment.

[0003] Compared with traditional general object detection, similarity amplifies the complexity of COD. To address this challenge, many recent studies have been proposed, which can be roughly divided into three types: targeted design of feature exploration modules, multi-task joint learning frameworks, and bio-inspired methods. From the perspective of bionics, some works have introduced bionic frameworks, including two modules: one is to globally search for targets, and the other is to identify targets. Then, some other studies have tried to study multi-task strategies, in which using auxiliary tasks helps to improve performance. For example, some works introduce edge detection as an auxiliary task during segmentation. The UGTR method uses an uncertainty perception method to simulate the inherent uncertainty in COD data. In addition, there have also been efforts to solve high similarity from the frequency perspective in other literatures. Recently, some studies inspired by human visual observation behavior have incorporated multi-scale strategies into neural networks, resulting in improved results. The CFNet method emphasizes the key role of sufficient context information in COD. However, these methods ignore the importance of the encoder and cannot explicitly decouple the camouflage features in the decoder, resulting in unsatisfactory results. Summary of the Invention

[0004] The purpose of this application is to propose a camouflaged target detection method, device and readable medium based on feature enhancement for the above-mentioned technical problems.

[0005] In a first aspect, the present invention provides a camouflaged target detection method based on feature enhancement, including the following steps:

[0006] Construct a camouflaged target detection model and train it to obtain a trained camouflaged target detection model; the camouflaged target detection model includes an encoder and a decoder, the encoder includes a multi-scale feature extraction module and three feature enhancement modules, and the decoder includes a context aggregation module, a first convolution module, a feature decoupling module and a second convolution module;

[0007] Obtain the camouflaged target image to be detected and input it into the trained camouflaged target detection model. The camouflaged target image to be detected first passes through the multi-scale feature extraction module in the encoder to obtain multi-scale features, which include first-level features, second-level features, third-level features, and fourth-level features. The second-level features, third-level features, and fourth-level features respectively pass through three feature enhancement modules to obtain second enhanced features, third enhanced features, and fourth enhanced features. The second enhanced feature, third enhanced feature, and fourth enhanced feature respectively pass through the context aggregation module to obtain second aggregated features, third aggregated features, and fourth aggregated features. The second aggregated feature is upsampled and then element-wise added to the first-level feature to obtain a first addition result, and the first addition result passes through the first convolutional module to obtain a first mask. The first mask, second aggregated feature, third aggregated feature, and fourth aggregated feature are input into the feature decoupling module to obtain second background suppression features, third background suppression features, and fourth background suppression features. The second background suppression feature is upsampled and then element-wise added to the first-level feature to obtain a second addition result. The second addition result passes through the second convolutional module to obtain the predicted detection result corresponding to the camouflaged target image to be detected.

[0008] Preferably, the multi-scale feature extraction module is a backbone network. The backbone network includes an initial stage, a first stage, a second stage, a third stage, and a fourth stage, and three first convolutional layers respectively connected to the second stage, the third stage, and the fourth stage. The camouflaged target image to be detected first passes through the initial stage and the first stage in sequence and then passes through the first convolutional layer connected to the first stage to obtain first-level features. The camouflaged target image to be detected first passes through the initial stage, the first stage, and the second stage in sequence and then passes through the first convolutional layer connected to the second stage to obtain second-level features. The camouflaged target image to be detected first passes through the initial stage, the first stage, the second stage, and the third stage in sequence and then passes through the first convolutional layer connected to the third stage to obtain third-level features. The camouflaged target image to be detected first passes through the initial stage, the first stage, the second stage, the third stage, and the fourth stage in sequence and then passes through the first convolutional layer connected to the fourth stage to obtain fourth-level features.

[0009] Preferably, the backbone network adopts ResNet-50 or PVT-v2; the feature enhancement module is a transformer module, and the transformer module includes a multi-head self-attention module and a first fully connected feed-forward network connected in sequence; the context aggregation module is a deformable attention module, and the feature decoupling module includes an embedding generation module, a cross-attention module, and a second fully connected feed-forward network connected in sequence; the first mask and the second aggregated feature pass through the embedding generation module to obtain a second foreground embedding feature and a second background embedding feature; the first mask and the third aggregated feature pass through the embedding generation module to obtain a third foreground embedding feature and a third background embedding feature; the first mask and the fourth aggregated feature pass through the embedding generation module to obtain a fourth foreground embedding feature and a fourth background embedding feature; perform a linear transformation on the features obtained by splicing the second foreground embedding feature, the second background embedding feature, the third foreground embedding feature, the third background embedding feature, the fourth foreground embedding feature, and the fourth background embedding feature to generate a key matrix and a value matrix, and perform a linear transformation on the second aggregated feature, the third aggregated feature, and the fourth aggregated feature respectively to generate a first query matrix, a second query matrix, and a third query matrix; input the first query matrix, the key matrix, and the value matrix into the cross-attention module to obtain a first intermediate feature, and pass the first intermediate feature through the second fully connected feed-forward network to obtain a second background suppression feature; input the second query matrix, the key matrix, and the value matrix into the cross-attention module to obtain a second intermediate feature, and pass the second intermediate feature through the second fully connected feed-forward network to obtain a third background suppression feature; input the third query matrix, the key matrix, and the value matrix into the cross-attention module to obtain a third intermediate feature, and pass the third intermediate feature through the second fully connected feed-forward network to obtain a fourth background suppression feature.

[0010] Preferably, the embedding generation module uses mask average pooling to generate foreground embedding features and background embedding features, as shown in the following formula:

[0011]

[0012] where, when i = 2, 3, or 4, represents the second aggregated feature, the third aggregated feature, or the fourth aggregated feature, represents the second foreground embedding feature, the third foreground embedding feature, or the fourth foreground embedding feature, represents the second background embedding feature, the third background embedding feature, or the fourth background embedding feature, and M represents the first mask.

[0013] Preferably, both the first convolution module and the second convolution module adopt a convolution structure, and the convolution structure includes a second convolution layer, a group normalization layer, a ReLu activation function layer, and a third convolution layer connected in sequence.

[0014] Preferably, during the training process of the camouflage target detection model, training data is first obtained. The training data includes camouflage target images and their corresponding true detection results. A third convolutional module is further provided behind the feature enhancement module. The third convolutional module adopts a convolutional structure. The second enhanced feature, the third enhanced feature, and the fourth enhanced feature corresponding to the camouflage target image respectively pass through the third convolutional module to obtain a second mask, a third mask, and a fourth mask; calculate the sum of the cross-entropy loss and the Dice loss between the predicted detection result corresponding to the first mask and the true detection result to obtain a first loss; calculate the sum of the cross-entropy loss and the Dice loss between the second mask and the true detection result to obtain a second loss; calculate the sum of the cross-entropy loss and the Dice loss between the third mask and the true detection result to obtain a third loss; calculate the sum of the cross-entropy loss and the Dice loss between the fourth mask and the true detection result to obtain a fourth loss; calculate the sum of the cross-entropy loss and the Dice loss between the predicted detection result generated by replacing the first mask input feature with the true detection result and the true detection result to obtain a fifth loss; the total loss function used in the training process of the camouflage target detection model is the sum of the first loss, the second loss, the third loss, the fourth loss, and the fifth loss.

[0015] In a second aspect, the present invention provides a camouflage target detection device based on feature enhancement, including:

[0016] A model construction module configured to construct and train a camouflage target detection model to obtain a trained camouflage target detection model; the camouflage target detection model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and three feature enhancement modules. The decoder includes a context aggregation module, a first convolutional module, a feature decoupling module, and a second convolutional module;

[0017] The detection module is configured to obtain a camouflaged target image to be detected and input it into a trained camouflaged target detection model. The camouflaged target image to be detected first passes through the multi-scale feature extraction module in the encoder to obtain multi-scale features, which include first-level features, second-level features, third-level features, and fourth-level features. The second-level features, third-level features, and fourth-level features respectively pass through three feature enhancement modules to obtain second enhanced features, third enhanced features, and fourth enhanced features. The second enhanced features, third enhanced features, and fourth enhanced features respectively pass through the context aggregation module to obtain second aggregation features, third aggregation features, and fourth aggregation features. The second aggregation feature is upsampled and then element-wise added to the first-level feature to obtain a first addition result. The first addition result passes through the first convolution module to obtain a first mask. The first mask, the second aggregation feature, the third aggregation feature, and the fourth aggregation feature are input into the feature decoupling module to obtain a second background suppression feature, a third background suppression feature, and a fourth background suppression feature. The second background suppression feature is upsampled and then element-wise added to the first-level feature to obtain a second addition result. The second addition result passes through the second convolution module to obtain the predicted detection result corresponding to the camouflaged target image to be detected.

[0018] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.

[0020] In a fifth aspect, the present invention provides a computer program product, including a computer program, which when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] (1) The camouflaged target detection method based on feature enhancement proposed by the present invention highlights the foreground area by performing early supervision on the encoder and suppressing background features in the decoder, weakening the interference of the background, and solving the problems of the traditional model being too complex and the encoder being difficult to extract camouflage features.

[0023] (2) The camouflaged target detection method based on feature enhancement proposed by the present invention performs a self-attention operation on the features extracted by the encoder, and uses the third convolutional module to obtain the predicted mask. Then, the mask is aligned with the true detection result, enabling the network to capture camouflaged features at an early stage.

[0024] (3) In the decoder of the camouflaged target detection method based on feature enhancement proposed by the present invention, the true detection result is used to generate foreground embedding features and background embedding features, and these two embeddings are used to highlight the foreground and suppress the background, aligning the finally generated predicted detection result with the true detection result, and solving the problem of high similarity between the foreground and background in camouflaged target detection. Brief Description of the Drawings

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0026] Figure 1 It is a schematic flow chart of the camouflaged target detection method based on feature enhancement for the embodiments of the present application;

[0027] Figure 2 It is a schematic diagram of the camouflaged target detection model of the camouflaged target detection method based on feature enhancement for the embodiments of the present application;

[0028] Figure 3 It is a schematic diagram of the feature enhancement module of the camouflaged target detection method based on feature enhancement for the embodiments of the present application;

[0029] Figure 4 It is a schematic diagram of the feature decoupling module of the camouflaged target detection method based on feature enhancement for the embodiments of the present application;

[0030] Figure 5 It is a schematic diagram of the camouflaged target detection device based on feature enhancement for the embodiments of the present application;

[0031] Figure 6 It is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present invention. Detailed Embodiments

[0032] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0033] Figure 1 A camouflaged target detection method based on feature enhancement is shown, including the following steps:

[0034] S1, construct a camouflaged target detection model and train it to obtain a trained camouflaged target detection model; the camouflaged target detection model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and three feature enhancement modules, and the decoder includes a context aggregation module, a first convolutional module, a feature decoupling module, and a second convolutional module.

[0035] In a specific embodiment, the multi-scale feature extraction module is a backbone network. The backbone network includes an initial stage, a first stage, a second stage, a third stage, and a fourth stage, and three first convolutional layers respectively connected to the second stage, the third stage, and the fourth stage. The camouflaged target image to be detected first passes through the initial stage and the first stage in sequence, and then passes through the first convolutional layer connected to the first stage to obtain the first-level feature; the camouflaged target image to be detected first passes through the initial stage, the first stage, and the second stage in sequence, and then passes through the first convolutional layer connected to the second stage to obtain the second-level feature; the camouflaged target image to be detected first passes through the initial stage, the first stage, the second stage, and the third stage in sequence, and then passes through the first convolutional layer connected to the third stage to obtain the third-level feature; the camouflaged target image to be detected first passes through the initial stage, the first stage, the second stage, the third stage, and the fourth stage in sequence, and then passes through the first convolutional layer connected to the fourth stage to obtain the fourth-level feature.

[0036] In a specific embodiment, during the training process of the camouflage target detection model, training data is first obtained. The training data includes camouflage target images and their corresponding true detection results. A third convolutional module is further provided behind the feature enhancement module. The third convolutional module adopts a convolutional structure. The second enhanced feature, the third enhanced feature, and the fourth enhanced feature corresponding to the camouflage target image respectively pass through the third convolutional module to obtain a second mask, a third mask, and a fourth mask; calculate the sum of the cross-entropy loss and the Dice loss between the predicted detection result corresponding to the first mask and the true detection result to obtain a first loss; calculate the sum of the cross-entropy loss and the Dice loss between the second mask and the true detection result to obtain a second loss; calculate the sum of the cross-entropy loss and the Dice loss between the third mask and the true detection result to obtain a third loss; calculate the sum of the cross-entropy loss and the Dice loss between the fourth mask and the true detection result to obtain a fourth loss; calculate the sum of the cross-entropy loss and the Dice loss between the predicted detection result generated by replacing the input feature of the first mask with the true detection result and the true detection result to obtain a fifth loss; the total loss function used in the training process of the camouflage target detection model is the sum of the first loss, the second loss, the third loss, the fourth loss, and the fifth loss.

[0037] Specifically, referring to Figure 2 , the embodiment of the present application uses a multi-scale feature extraction module to extract multi-scale features. The multi-scale feature extraction module uses a backbone network. The backbone network includes 5 stages, namely the initial stage stage0, the first stage stage1, the second stage stage2, the third stage stage3, and the fourth stage stage4. A first convolutional layer with a convolutional kernel size of 1×1 connected thereto is provided in the first stage, the second stage, the third stage, and the fourth stage. The feature extracted in the initial stage is the initial feature f0. The features extracted in the first stage, the second stage, the third stage, and the fourth stage respectively pass through the first convolutional layer and the sequentially extracted features are the initial feature f0, the first-level feature f1, the second-level feature f2, the third-level feature f3, and the fourth-level feature f4; the embodiment of the present application selects the high-level first-level feature f1, the second-level feature f2, the third-level feature f3, and the fourth-level feature f4 as multi-scale features, and the multi-scale features are represented as The second-level feature f2, the third-level feature f3, and the fourth-level feature f4 in the extracted multi-scale features are respectively sent to the feature enhancement module for feature enhancement to obtain enhanced features That is, the second enhanced feature The third enhanced feature And the fourth enhanced feature The feature enhancement module adopts a Transformer module. During the training process of the camouflaged object detection model, the third convolutional module is used to obtain the second mask mask2, the third mask mask3, and the fourth mask mask4, and then is aligned with the real detection result (GT, Ground Truth); the alignment method here is to calculate the cross-entropy loss and Dice loss between and the real detection result GT respectively and then sum them up.

[0038] Furthermore, the second mask mask2, the third mask mask3, and the fourth mask mask4 are input into the context aggregation module. The context aggregation module adopts a variability attention module, and uses the variability attention mechanism to aggregate the enhanced features to obtain the second aggregated feature the third aggregated feature and the fourth aggregated feature The second aggregated feature is upsampled and then element-wise added to the first-level feature f1 to obtain the first addition result. The first addition result passes through the first convolutional module to generate the first mask mask g . During the training process of the camouflaged object detection model, the first mask mask g is aligned with the real detection result (GT, Ground Truth). The alignment method adopted here is the same as the above alignment method. And the real detection result, the second aggregated feature the third aggregated feature and the fourth aggregated feature are input into the feature decoupling module to obtain the background-suppressed features, namely the second background-suppressed feature f2′, the third background-suppressed feature f3′, and the fourth background-suppressed feature f4′. The second background-suppressed feature f2′ is upsampled and then element-wise added to the first-level feature f1 to obtain the second addition result. The second addition result passes through the second convolutional module to obtain the corresponding predicted detection result. The predicted detection result is aligned with the real detection result GT. The alignment method adopted here is the same as the above alignment method. The losses obtained in all alignment processes are summed up to obtain the total loss function. The camouflaged object detection model is trained using this total loss function to obtain the trained camouflaged object detection model.

[0039] S2. Obtain the camouflage target image to be detected and input it into the trained camouflage target detection model. The camouflage target image to be detected first passes through the multi-scale feature extraction module in the encoder to obtain multi-scale features, including the first-level feature, the second-level feature, the third-level feature, and the fourth-level feature. The second-level feature, the third-level feature, and the fourth-level feature respectively pass through three feature enhancement modules to obtain the second enhanced feature, the third enhanced feature, and the fourth enhanced feature. The second enhanced feature, the third enhanced feature, and the fourth enhanced feature respectively pass through the context aggregation module to obtain the second aggregated feature, the third aggregated feature, and the fourth aggregated feature. The second aggregated feature is upsampled and then element-wise added to the first-level feature to obtain the first addition result. The first addition result passes through the first convolution module to obtain the first mask. The first mask, the second aggregated feature, the third aggregated feature, and the fourth aggregated feature are input into the feature decoupling module to obtain the second background suppression feature, the third background suppression feature, and the fourth background suppression feature. The second background suppression feature is upsampled and then element-wise added to the first-level feature to obtain the second addition result. The second addition result passes through the second convolution module to obtain the predicted detection result corresponding to the camouflage target image to be detected.

[0040] In a specific embodiment, the backbone network adopts ResNet-50 or PVT-v2; the feature enhancement module is a transformer module, and the transformer module includes a multi-head self-attention module and a first fully-connected feed-forward network connected in sequence; the context aggregation module is a variability attention module, and the feature decoupling module includes an embedding generation module, a cross-attention module, and a second fully-connected feed-forward network connected in sequence; the first mask and the second aggregated feature pass through the embedding generation module to obtain a second foreground embedding feature and a second background embedding feature; the first mask and the third aggregated feature pass through the embedding generation module to obtain a third foreground embedding feature and a third background embedding feature; the first mask and the fourth aggregated feature pass through the embedding generation module to obtain a fourth foreground embedding feature and a fourth background embedding feature; linear transformation is performed on the features obtained by splicing the second foreground embedding feature, the second background embedding feature, the third foreground embedding feature, the third background embedding feature, the fourth foreground embedding feature, and the fourth background embedding feature to generate a key matrix and a value matrix, and linear transformation is respectively performed on the second aggregated feature, the third aggregated feature, and the fourth aggregated feature to generate a first query matrix, a second query matrix, and a third query matrix; the first query matrix, the key matrix, and the value matrix are input into the cross-attention module to obtain a first intermediate feature, and the first intermediate feature passes through the second fully-connected feed-forward network to obtain a second background suppression feature; the second query matrix, the key matrix, and the value matrix are input into the cross-attention module to obtain a second intermediate feature, and the second intermediate feature passes through the second fully-connected feed-forward network to obtain a third background suppression feature; the third query matrix, the key matrix, and the value matrix are input into the cross-attention module to obtain a third intermediate feature, and the third intermediate feature passes through the second fully-connected feed-forward network to obtain a fourth background suppression feature.

[0041] In a specific embodiment, both the first convolution module and the second convolution module adopt a convolution structure, and the convolution structure includes a second convolution layer, a group normalization layer, a ReLu activation function layer, and a third convolution layer connected in sequence.

[0042] Specifically, the trained camouflaged target detection model is deployed. In the inference stage, the camouflaged target image to be detected is input into the trained camouflaged target detection model, and multi-scale features are extracted through the multi-scale feature extraction module in sequence. The second-level feature, third-level feature, and fourth-level feature in the multi-scale features are input into the feature enhancement module for feature enhancement to obtain the second enhanced feature, third enhanced feature, and fourth enhanced feature. The second enhanced feature, third enhanced feature, and fourth enhanced feature respectively pass through the context aggregation module to obtain the second aggregated feature, third aggregated feature, and fourth aggregated feature. The second aggregated feature is upsampled to the same resolution as the first-level feature f1, and then added element-wise to the first-level feature f1 to obtain the first addition result. Then, the first addition result is convolved using the first convolution module to obtain the first mask. The first mask, second aggregated feature, third aggregated feature, and fourth aggregated feature are input into the feature decoupling module to obtain the second background suppression feature, third background suppression feature, and fourth background suppression feature. The second background suppression feature is upsampled to the same resolution as the first-level feature f1, and then added element-wise to the first-level feature f1 to obtain the second addition result. Then, the second addition result is convolved using the second convolution module to obtain the corresponding predicted detection result. Among them, the first convolution module, second convolution module, and third convolution module all adopt a convolution structure, and this convolution structure includes a second convolution layer with a convolution kernel size of 3×3, a group normalization layer, a ReLu activation function layer, and a third convolution layer with a convolution kernel size of 3×3.

[0043] Reference Figure 3 , in the feature enhancement module, one of the second-level feature, third-level feature, and fourth-level feature first passes through the multi-head self-attention module (MHA). The multi-head self-attention module includes three linear layers, which are used to perform linear transformation on one of the second-level feature, third-level feature, and fourth-level feature, and then input it into the self-attention layer. The outputs of multiple self-attention layers pass through a merging layer (not shown) and a linear layer (not shown) to obtain the fourth intermediate feature. The fourth intermediate feature passes through the first fully connected feed-forward network (FFN) to obtain the enhanced feature, and the enhanced feature

[0044] In a specific embodiment, the embedding generation module generates foreground embedding features and background embedding features using masked average pooling, as shown in the following formula:

[0045]

[0046] Among them, when i = 2, 3, or 4, represents the second aggregated feature, third aggregated feature, or fourth aggregated feature, represents the second foreground embedding feature, the third foreground embedding feature, or the fourth foreground embedding feature, represents the second background embedding feature, the third background embedding feature, or the fourth background embedding feature, and M represents the first mask.

[0047] Specifically, referring to Figure 4 , in the feature decoupling module, the first mask is respectively input into the embedding generation module together with the second aggregated feature, the first mask and the third aggregated feature, and the first mask and the fourth aggregated feature, and the corresponding foreground embedding features and background embedding features are calculated. Then, all the foreground embedding features and background embedding features are concatenated. The concatenated features are respectively linearly transformed through the corresponding linear layers to obtain the key matrix and the value matrix. The second aggregated feature, the third aggregated feature, and the fourth aggregated feature are respectively linearly transformed through the linear layer to obtain the first query matrix, the second query matrix, and the third query matrix. The first query matrix, the second query matrix, and the third query matrix are respectively input into the cross-attention module together with the key matrix and the value matrix for cross-attention calculation to obtain the corresponding intermediate features. The corresponding intermediate features are passed through the second fully connected feed-forward network to obtain the second background suppression feature, the third background suppression feature, and the fourth background suppression feature. The second background suppression feature is upsampled and then added pixel by pixel to the first-level feature to obtain the second addition result. The second addition result is passed through the second convolutional module to obtain the corresponding predicted detection result.

[0048] The trained camouflaged object detection model (denoted as SDNet) proposed in the embodiments of this application was implemented using PyTorch\cite{paszke2019pytorch}, and the structures of CNN and ViT were respectively used as the backbones. For the CNN-based backbone network, ResNet-50 pre-trained on ImageNet was used, and the remaining parameters were randomly initialized. The embodiments of this application use Adam as the default optimizer. To adjust the learning rate, the embodiments of this application implemented a cosine strategy with an initial learning rate of 0.0002. During the training and inference processes, the image size was set to 576*576. The model was trained 160 times with a batch size of 4 on 4 Nvidia 2080Ti GPUs, taking about 4 hours. Data augmentation techniques such as random flipping and rotation were used to enrich the training dataset. For the transformer-based backbone network, the pre-trained PVTv2-B4 was used in the embodiments of this application, and 4 Nvidia 3090 GPUs were used for 40 times of training. The initial learning rate was set to 8e-5, and other settings remained unchanged.

[0049] Table 1 shows the results of the present invention on the COD10K, NC4K, CHAMELEON, and CAMO datasets. As can be seen from Table 1, on the four datasets, the present invention significantly outperforms other methods. Compared with the method using resnet-50, our method has been greatly improved. Especially for the baseline method, without any enhancement, the detection performance is not weaker than the existing sota methods.

[0050] Table 1

[0051]

[0052] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present application provides an embodiment of a camouflaged target detection device based on feature enhancement. This device embodiment corresponds to the method embodiment shown in Figure 1 and can be specifically applied to various electronic devices.

[0053] The embodiment of the present application provides a camouflaged target detection device based on feature enhancement, including:

[0054] A model construction module 1, configured to construct and train a camouflaged target detection model to obtain a trained camouflaged target detection model; the camouflaged target detection model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and three feature enhancement modules, and the decoder includes a context aggregation module, a first convolution module, a feature decoupling module, and a second convolution module;

[0055] A detection module 2, configured to obtain a camouflaged target image to be detected and input it into the trained camouflaged target detection model. The camouflaged target image to be detected first passes through the multi-scale feature extraction module in the encoder to obtain multi-scale features, which include first-level features, second-level features, third-level features, and fourth-level features; the second-level features, third-level features, and fourth-level features respectively pass through three feature enhancement modules to obtain second enhanced features, third enhanced features, and fourth enhanced features; the second enhanced features, third enhanced features, and fourth enhanced features respectively pass through the context aggregation module to obtain second aggregation features, third aggregation features, and fourth aggregation features; the second aggregation feature is upsampled and then element-wise added to the first-level feature to obtain a first addition result, and the first addition result passes through the first convolution module to obtain a first mask; the first mask, the second aggregation feature, the third aggregation feature, and the fourth aggregation feature are input into the feature decoupling module to obtain second background suppression features, third background suppression features, and fourth background suppression features; the second background suppression feature is upsampled and then element-wise added to the first-level feature to obtain a second addition result; the second addition result passes through the second convolution module to obtain the predicted detection result corresponding to the camouflaged target image to be detected.

[0056] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. As Figure 6 shown, the electronic device of this embodiment includes: a processor 601 and a memory 602; wherein the memory 602 is used to store computer-executable instructions; the processor 601 is used to execute the computer-executable instructions stored in the memory to implement the respective steps executed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0057] Optionally, the memory 602 can be either independent or integrated with the processor 601.

[0058] When the memory 602 is independently provided, the electronic device further includes a bus 603 for connecting the memory 602 and the processor 601.

[0059] The embodiment of the present invention also provides a computer storage medium, in which computer-executable instructions are stored. When the processor 601 executes the computer-executable instructions, the above method is implemented.

[0060] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 601, the above method is implemented.

[0061] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0062] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0063] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module exists physically alone, or two or more modules are integrated in one unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a hardware plus software functional unit.

[0064] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor 601 to execute some steps of the methods according to the various embodiments of the present application.

[0065] It should be understood that the above-mentioned processor 601 may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor 601 may also be any conventional processor 601, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor 601, or can be implemented by the combination of the hardware and software modules in the processor 601.

[0066] The memory 602 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disc, etc.

[0067] The bus 603 may be an Industry Standard Architecture (ISA for short), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 603 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus 603 in the drawings of the present application does not limit that there is only one bus 603 or one type of bus 603.

[0068] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0069] An exemplary storage medium is coupled to the processor 601, enabling the processor 601 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 601. The processor 601 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 601 and the storage medium can also exist as discrete components in an electronic device or a master control device.

[0070] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A camouflaged target detection method based on feature enhancement, characterized in that: The following steps are involved: Constructing and training a disguised target detection model to obtain a trained disguised target detection model; the disguised target detection model includes an encoder and a decoder, the encoder includes a multi-scale feature extraction module and three feature enhancement modules, and the decoder includes a context aggregation module, a first convolution module, a feature decoupling module and a second convolution module; Acquire a disguised target image to be detected and input it into the trained disguised target detection model, wherein the disguised target image to be detected first passes through a multi-scale feature extraction module in the encoder to obtain multi-scale features, wherein the multi-scale features include first-level features, second-level features, third-level features, and fourth-level features; The second-level features, the third-level features and the fourth-level features are respectively subjected to the three feature enhancement modules to obtain the second enhanced features, the third enhanced features and the fourth enhanced features; The second enhanced feature, the third enhanced feature and the fourth enhanced feature are respectively passed through the context aggregation module to obtain a second aggregated feature, a third aggregated feature and a fourth aggregated feature; The second aggregated feature is upsampled and added element-by-element with the first-level feature to obtain a first addition result, and the first addition result is passed through the first convolution module to obtain a first mask; the first mask, the second aggregated feature, the third aggregated feature and the fourth aggregated feature are input into the feature decoupling module to obtain a second background suppression feature, a third background suppression feature and a fourth background suppression feature; the second background suppression feature is upsampled and added element-by-element with the first-level feature to obtain a second addition result; the second addition result is passed through the second convolution module to obtain a predicted detection result corresponding to the camouflaged target image to be detected.

2. The method for detecting camouflaged targets based on feature enhancement according to claim 1, characterized in that: The multi-scale feature extraction module is a backbone network, which includes an initial stage, a first stage, a second stage, a third stage and a fourth stage, and three first convolutional layers respectively connected to the second stage, the third stage and the fourth stage. The camouflaged target image to be detected first passes through the initial stage and the first stage in sequence, and then passes through the first convolutional layer connected to the first stage to obtain a first-level feature; The camouflaged target image to be detected first passes through the initial stage, the first stage, and the second stage in sequence, and then passes through the first convolution layer connected to the second stage to obtain the second-level features; The disguised target image to be detected first passes through the initial stage, the first stage, the second stage and the third stage in sequence, and then passes through the first convolution layer connected to the third stage to obtain the third-level features; the disguised target image to be detected first passes through the initial stage, the first stage, the second stage, the third stage and the fourth stage in sequence, and then passes through the first convolution layer connected to the fourth stage to obtain the fourth-level features.

3. The method for detecting camouflaged targets based on feature enhancement according to claim 2, characterized in that: The backbone network adopts ResNet-50 or PVT-v2; the feature enhancement module is a transformer module, and the transformer module includes a multi-head self-attention module and a first fully connected feedforward network connected in sequence; the context aggregation module is a variability attention module, and the feature decoupling module includes an embedding generation module, a cross attention module and a second fully connected feedforward network connected in sequence; the first mask and the second aggregated features are passed through the embedding generation module to obtain a second foreground embedding feature and a second background embedding feature; The first mask and the third aggregated feature are passed through the embedding generation module to obtain a third foreground embedding feature and a third background embedding feature; the first mask and the fourth aggregated feature are passed through the embedding generation module to obtain a fourth foreground embedding feature and a fourth background embedding feature; Performing linear transformation on the concatenation of the second foreground embedding feature, the second background embedding feature, the third foreground embedding feature, the third background embedding feature, the fourth foreground embedding feature and the fourth background embedding feature to generate a key matrix and a value matrix, performing linear transformation on the second aggregated feature, the third aggregated feature and the fourth aggregated feature respectively to generate a first query matrix, a second query matrix and a third query matrix; inputting the first query matrix, the key matrix and the value matrix into the cross attention module to obtain a first intermediate feature, and passing the first intermediate feature through the second fully connected feed-forward network to obtain a second background suppression feature; Inputting the second query matrix, the key matrix and the value matrix into the cross attention module to obtain a second intermediate feature, and passing the second intermediate feature through the second fully connected feedforward network to obtain a third background suppression feature; The third query matrix, key matrix and value matrix are input into the cross attention module to obtain a third intermediate feature, and the third intermediate feature is passed through the second fully connected feedforward network to obtain a fourth background suppression feature.

4. The method for detecting camouflaged targets based on feature enhancement according to claim 3, characterized in that: The embedding generation module generates foreground embedding features and background embedding features using mask average pooling, as shown in the following formula: When i=2, 3 or 4, represents the second aggregate feature, the third aggregate feature, or the fourth aggregate feature, represents the second foreground embedding feature, the third foreground embedding feature or the fourth foreground embedding feature, represents the second background embedding feature, the third background embedding feature or the fourth background embedding feature, and M represents the first mask.

5. The method for detecting camouflaged targets based on feature enhancement according to claim 1, characterized in that: The first convolution module and the second convolution module both adopt a convolution structure, and the convolution structure includes a second convolution layer, a group normalization layer, a ReLu activation function layer and a third convolution layer connected in sequence.

6. The method for detecting camouflaged targets based on feature enhancement according to claim 5, characterized in that: In the training process of the disguised target detection model, training data is first obtained, the training data includes a disguised target image and its corresponding real detection result, and a third convolution module is further arranged behind the feature enhancement module. The third convolution module adopts the convolution structure, and the second enhancement feature, the third enhancement feature and the fourth enhancement feature corresponding to the disguised target image are respectively passed through the third convolution module to obtain a second mask, a third mask and a fourth mask; Calculate the sum of the cross entropy loss and the Dice loss between the predicted detection result and the true detection result corresponding to the first mask to obtain a first loss; Calculate the sum of the cross entropy loss and the Dice loss between the second mask and the true detection result to obtain the second loss; calculate the sum of the cross entropy loss and the Dice loss between the third mask and the true detection result to obtain the third loss; calculate the sum of the cross entropy loss and the Dice loss between the fourth mask and the true detection result to obtain the fourth loss; The sum of the cross entropy loss and the Dice loss between the predicted detection result and the true detection result generated by replacing the first mask input feature decoupling module with the true detection result is calculated to obtain a fifth loss; the total loss function used in the training process of the camouflaged target detection model is the sum of the first loss, the second loss, the third loss, the fourth loss and the fifth loss.

7. A camouflaged target detection device based on feature enhancement, characterized in that: include: A model building module is configured to build and train a disguised target detection model to obtain a trained disguised target detection model; The camouflaged target detection model includes an encoder and a decoder, the encoder includes a multi-scale feature extraction module and three feature enhancement modules, and the decoder includes a context aggregation module, a first convolution module, a feature decoupling module and a second convolution module; a detection module configured to obtain a disguised target image to be detected and input it into the trained disguised target detection model, wherein the disguised target image to be detected first passes through a multi-scale feature extraction module in the encoder to obtain multi-scale features, wherein the multi-scale features include first-level features, second-level features, third-level features, and fourth-level features; The second-level features, the third-level features and the fourth-level features are respectively subjected to the three feature enhancement modules to obtain the second enhanced features, the third enhanced features and the fourth enhanced features; The second enhanced feature, the third enhanced feature and the fourth enhanced feature are respectively passed through the context aggregation module to obtain a second aggregated feature, a third aggregated feature and a fourth aggregated feature; The second aggregated feature is upsampled and added element-by-element with the first-level feature to obtain a first addition result, and the first addition result is passed through the first convolution module to obtain a first mask; the first mask, the second aggregated feature, the third aggregated feature and the fourth aggregated feature are input into the feature decoupling module to obtain a second background suppression feature, a third background suppression feature and a fourth background suppression feature; the second background suppression feature is upsampled and added element-by-element with the first-level feature to obtain a second addition result; the second addition result is passed through the second convolution module to obtain a predicted detection result corresponding to the camouflaged target image to be detected.

8. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Camouflage target detection method and device based on structure prior and double-domain decoupling hybrid expert

    CN122049346A

  • Camouflage target detection method and device based on structural prior and dual-domain decoupling hybrid expert

    CN122049346B