Camouflage target detection method based on multi-scale feature fusion and interference suppression
Through the camouflage object detection method of multi-scale feature fusion and interference suppression, the problems of low feature transmission efficiency and background noise interference in the prior art are solved, and efficient and accurate detection of camouflage object is achieved.
Patent Information
- Application Number
- CN202510358800.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing camouflage object detection methods are inefficient in feature delivery and are prone to introduce background noise, and ignore interference caused by the similarity between the target and the background, resulting in low detection accuracy.
A camouflage object detection method based on multi-scale feature fusion and interference suppression is adopted. Through parallel coding and multi-scale feature fusion, combined with edge feature information and interference filtering module, the detection performance is significantly improved.
Effectively integrate multi-scale features, reduce error propagation, improve detection accuracy and robustness of camouflage targets, and significantly reduce false detection and missed detection.
Smart Images

Figure CN120298867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to object detection, and specifically to a camouflaged object detection method based on multi-scale feature fusion and interference suppression, which is mainly applied to camouflaged object detection tasks in scenarios such as autonomous driving, intelligent security, and military reconnaissance. The aim is to improve the recognition accuracy and detection efficiency of the system for concealed or camouflaged objects, and it belongs to the field of intelligent recognition technology. Background Art
[0002] Camouflage is a widespread and effective means for evading biological detection and recognition. In nature, camouflage bodies have evolved a series of hiding strategies to interfere with the perception and cognitive mechanisms of prey or predators, such as background matching, self-shadow hiding, object iterative shadow, disruptive coloration, and distraction markings. Compared with general object detection, these camouflage strategies greatly increase the challenge of camouflaged object detection (COD). COD aims to identify camouflaged objects highly similar to the background, and requires computer vision models to assist the human visual system to enhance the detection ability. Its applications cover multiple fields, including polyp segmentation, lung infection recognition, pest monitoring, and entertainment art.
[0003] With the rise of Transformer in the visual field, self-attention and cross-attention mechanisms provide effective means for capturing long-range dependencies and constructing global content-aware interactions. However, many existing methods simply combine CNN and Transformer or use a shared encoder for feature extraction, as shown in Figure 1 (a-c), where (a) is a basic encoder-decoder architecture; (b) is a skip-layer structure; (c) is global guidance; this mechanism has two significant defects: firstly, as information is passed from bottom to top, high-level spatial features will experience progressive attenuation; secondly, due to the lack of an effective feature fusion mechanism, simple cross-layer connections are difficult to fully integrate multi-level features, which not only restricts the effective transmission of multi-scale information, but may also lead to error accumulation.
[0004] Existing COD methods generally regard the background simply as noise or interference sources. Taking SINetv2 as an example, this network architecture adopts a reverse attention mechanism to strengthen the representation of background information by suppressing the saliency of foreground features, so as to achieve accurate positioning and feature extraction of potential camouflaged areas. However, in the COD task, due to the high similarity of visual features such as texture and color between the target and the background environment, the modeling strategy relying solely on background interference information often fails to achieve effective feature discrimination. Figure 2 It is a detection result graph of existing different COD methods on specific instances. The four pictures on the left, from top to bottom, are objects with artificial camouflage, small objects, occluded objects, and objects with difficult-to-segment boundaries. From Figure 2It can be seen that in the first row, in the artificial camouflage scenario, most current camouflage methods cannot identify the camouflaged object; in the second row, the camouflaged object is highly integrated with the surrounding environment, resulting in false detection and missed detection by the model; in the third row, the camouflaged animal is partially blocked by sand, causing most networks to ignore this detail; in the fourth row, there are two camouflaged organisms of the same category in the scene, but due to their texture being extremely similar to the target, the recognition ability of the network is still limited. Therefore, there are still great challenges in the current methods for identifying these camouflaged objects. Summary of the Invention
[0005] Aiming at the problems that the existing camouflage target detection methods have low feature transfer efficiency and are prone to introducing background noise; at the same time, the existing methods only regard the background as an interference factor and ignore the two types of interference brought by the similarity between the target and the background, thus affecting the detection accuracy, the present invention proposes a camouflage target detection method based on multi-scale feature fusion and interference suppression. The present invention fuses multi-scale features and interference suppression, and through parallel coding and multi-scale feature fusion, combines edge feature information and an interference filtering module, significantly improving the detection performance of camouflage targets.
[0006] The technical solution of the present invention is implemented as follows:
[0007] A camouflage target detection method based on multi-scale feature fusion and interference suppression, the steps are as follows.
[0008] 1) Train a camouflage target detection model; the camouflage target detection model includes a pre-processing network module, a differential feature interaction module, a detail feature fusion module, a boundary perception module, and an interference filtering module.
[0009] 1.1) The pre-processing network module includes a main image processing branch and a sub-image processing branch; the main image processing branch is a pyramid vision Transformer architecture for processing original image data to extract four-stage feature representations f i (i = 1, 2, 3, 4); the sub-image processing branch is a convolutional neural network for processing the image data after the original image is scaled by a factor of two to sequentially generate four groups of multi-scale feature maps y i (i = 1, 2, 3, 4);
[0010] 1.2) Through bilinear interpolation spatial transformation operation, align the feature resolution of the sub-image processing branch with the feature resolution of the main image processing branch, that is, the feature y i and the feature f i+1(i = 1, 2, 3, 4) achieves synchronous matching of spatial resolution at the corresponding levels; among them, feature y1 and f2, and feature y2 and f3 respectively obtain feature Z2 and feature Z3 through the synergistic effect of max pooling and average pooling, so as to enhance the discriminative ability of the feature while retaining the key feature information; feature y3 and f4 are deeply fused through the differential feature interaction module, and the output after fusion is denoted as feature Z4;
[0011] 1.3) Feature Z4 undergoes a convolution operation to obtain feature F4; feature F4 and feature Z3 jointly undergo a convolution operation to obtain feature F3; feature F3 and feature Z2 jointly undergo a convolution operation to obtain feature F2; feature F2 and feature f1 jointly undergo a convolution operation to obtain feature F1;
[0012] 1.4) Transmit features F1, F2, F3, and F4 to the detailed feature fusion module and the boundary awareness module simultaneously; the rough prediction result F is output after fusion by the detailed feature fusion module c ; the boundary awareness module outputs the edge feature F e ;
[0013] 1.5) Transmit the preliminary prediction result F c and the edge feature F e to the interference filtering module together. The interference filtering module introduces a two-branch structure and explicitly models and processes interference factors through a lightweight encoder and an attention mechanism, and outputs the final prediction map F p , and this final prediction map F p is the output of the camouflage target detection model;
[0014] 2) Input the RGB image to be detected into the camouflage target detection model, and the output of the camouflage target detection model is the detected camouflage target map.
[0015] Further, in step 1.2), feature y3 and f4 are deeply fused through the differential feature interaction module to obtain feature Z4, which is specifically carried out according to the following formula:
[0016]
[0017] θ(f4) refers to the process of converting f4 ∈ R C×H×W to f′4 ∈ R H×W×C through flattening and permutation operations, and then obtaining through layer normalization and linear transformation. The process of generating θ(y3) is the same as the process of θ(f4).
[0018] Further, the detailed feature fusion module outputs the rough prediction result F according to the process reflected by the following formula c ,
[0019] F′4 = Conv3(F4) #(3)
[0020]
[0021] F c = Conv3(F′1) #(10)
[0022] Conv3 and Conv1 represent 3×3 and 1×1 convolutions respectively, Up represents the upsampling operation, [;] represents channel concatenation, represents element-wise multiplication, represents element-wise addition.
[0023] Further, in step 1.2), the boundary awareness module outputs the edge feature F according to the following process e ,
[0024] First, the number of channels of the input feature is adjusted to 64 through a 1×1 convolutional layer, then the features F2 - F4 are upsampled to the same size as the feature F1, and then channel concatenation is performed to obtain the preliminary edge feature f e ; The formula is expressed as follows:
[0025] f e = Concat(Conv1(F1), Up(Conv1(F2)), Up(Conv1(F3)), Up(Conv1(F4))) #(11)
[0026] where Up(·) represents the upsampling operation and Concat(·) represents the concatenation operation;
[0027] Finally, channel attention and residual connection are performed on the preliminary edge feature f e to generate clearer and more accurate edge information, which is the edge feature F e , and the formula is expressed as follows:
[0028]
[0029] where BConv(·) represents a convolutional operation with a kernel size of 3×3, followed by a batch normalization layer and a ReLU activation function, and ⊕ represents the element-wise addition operation.
[0030] Further, the interference filtering module outputs the final prediction map F according to the following method p ,
[0031] 1.5.1) The interference filtering module processes the preliminary prediction result F through the dynamic spatial attention mechanism respectively c to obtain the feature f n and the feature f p ;
[0032] 1.5.2) Then, feature f n is concatenated with F c in the channel dimension and fed into the attention mechanism to generate feature f w . Then, feature f w and the edge feature F e are multiplied element-wise, and the original feature information is retained through residual connection to generate the enhanced feature F fn ;
[0033] 1.5.3) f p is concatenated with F fn , and the concatenated feature is fed into two 3×3 convolutional units to obtain the predicted feature that suppresses the interference of f p ;
[0034] 1.5.4) The predicted feature that suppresses the interference of f fn is subtracted from F p , and finally processed by a 3×3 convolutional kernel to obtain the final predicted map F p .
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] Compared with the existing camouflaged object detection COD technology, the present invention adopts a parallel coding strategy, combines CNN and Transformer to extract local details and global context information respectively, and realizes the effective fusion of multi-scale features through the differential feature interaction module. Compared with the design of sharing an encoder in the existing method, this method can better utilize multi-scale information, reduce error propagation, and improve the feature transfer efficiency. Secondly, two types of interference factors in camouflaged object detection are modeled, and an interference filtering module is designed to refine the rough predicted map through supervised learning. This design significantly improves the model's ability to identify interference factors and reduces the cases of false detection and missed detection. In addition, through the detailed feature fusion module, the present invention realizes bottom-up dense feature transfer, effectively suppresses background interference, and retains the original information, thereby improving the accuracy and robustness of feature representation.
[0037] The present invention constructs a two-stream feature extraction framework, and then uses an attention-guided interference filtering module to fuse semantic clues in the spatial and channel dimensions to achieve precise segmentation and localization of the target area.
[0038] The present invention explicitly models the camouflaged object information and camouflaged edge information in the network to retain the boundary of the camouflaged object, thereby improving the detection accuracy.
[0039] The model proposed by the present invention combines multi-scale feature extraction and explicit modeling of interference factors to achieve more accurate segmentation of camouflaged objects.
[0040] The present invention designs a dual-stream feature extraction architecture. By integrating the local feature extraction ability of Res2Net50 and the global modeling advantage of Transformer, this architecture realizes the collaborative encoding of multi-scale features. Furthermore, an attention-guided feature interaction mechanism is designed to achieve the alignment and fusion of cross-modal features through cross-attention, so as to fully exploit the complementary feature expression capabilities of the two. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 - Comparison diagram of the network architecture proposed by the present invention and other existing network frameworks. Among them, (a) basic encoder-decoder architecture; (b) skip layer structure; (c) global guidance; (d) architecture proposed by the present invention.
[0042] Figure 2 - Detection result diagrams of existing different camouflaged object detection methods on specific instances. The four pictures on the left, from top to bottom, are objects with artificial camouflage, small objects, occluded objects, and objects with difficult-to-segment boundaries respectively.
[0043] Figure 3 - Schematic diagram of the model architecture of the camouflaged object detection method of the present invention.
[0044] Figure 4 - Schematic diagram of the architecture of the detailed feature fusion module proposed by the present invention.
[0045] Figure 5 - Schematic diagram of the architecture of the boundary awareness module of the present invention.
[0046] Figure 6 - Schematic diagram of the structure of the interference filtering module of the present invention.
[0047] Figure 7 - Visualization comparison diagram between the present invention and other different camouflaged object detection models.
[0048] Figure 8 - Ablation experiment comparison diagram provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0050] The present invention aims to solve two main problems in the existing camouflaged object detection (COD) technology: one is the effective fusion and transmission of multi-scale features, and the other is the accurate modeling and processing of interference factors. To this end, the present invention proposes a camouflaged object detection method based on multi-scale feature fusion and interference suppression. The basis of this detection method is a camouflaged object detection model of a multi-module system. The overall model framework is as Figure 3As shown in the figure, the model of the present invention includes four core components: a Differential Feature Interaction Module (DFIM for short), a Detail Feature Fusion Module (DFFM for short), a Boundary Perception Module (BPM for short), and an Interference Filtering Module (IFM for short).
[0051] The specific steps of the method for detecting camouflaged targets of the present invention are as follows.
[0052] 1) Train a camouflaged target detection model; save the model with the best training effect as the final detection model. The camouflaged target detection model includes a pre-processing network module, a differential feature interaction module, a detail feature fusion module, a boundary perception module, and an interference filtering module.
[0053] 1.1) The pre-processing network module includes a main image processing branch and a sub-image processing branch; the main image processing branch is a Pyramid Vision Transformer architecture for processing original image data to extract four-stage feature representations f i (i = 1, 2, 3, 4); the sub-image processing branch is a convolutional neural network for processing the image data of the original image after being scaled by a factor of two to sequentially generate four groups of multi-scale feature maps y i (i = 1, 2, 3, 4);
[0054] 1.2) Through bilinear interpolation space transformation operation, align the feature resolution of the sub-image processing branch with that of the main image processing branch, that is, the feature y i and the feature f i+1 (i = 1, 2, 3, 4) achieve synchronous matching of spatial resolution at the corresponding levels; among them, the feature y1 and f2, and the feature y2 and f3 respectively obtain the features Z2 and Z3 through the synergistic effect of max pooling and average pooling to enhance the discriminative ability of the features while retaining key feature information; the feature y3 and f4 are deeply fused through the differential feature interaction module, and the fused output is denoted as the feature Z4;
[0055] 1.3) The feature Z4 undergoes a convolution operation to obtain the feature F4; the feature F4 and the feature Z3 jointly undergo a convolution operation to obtain the feature F3; the feature F3 and the feature Z2 jointly undergo a convolution operation to obtain the feature F2; the feature F2 and the feature f1 jointly undergo a convolution operation to obtain the feature F1;
[0056] 1.4) Transmit features F1, F2, F3, and F4 to the detailed feature fusion module and the boundary awareness module simultaneously; after fusion by the detailed feature fusion module, output the preliminary prediction result F c ; the boundary awareness module outputs the edge feature F e ;
[0057] 1.5) Transmit the preliminary prediction result F c and the edge feature F e to the interference filtering module together. The interference filtering module introduces a dual-branch structure and, through a lightweight encoder and an attention mechanism, explicitly models and processes interference factors, and outputs the final prediction map F p , and this final prediction map F p is the output of the camouflage target detection model;
[0058] 2) Input the RGB image to be detected into the camouflage target detection model, and the output of the camouflage target detection model is the detected camouflage target map.
[0059] The model proposed in the present invention is implemented based on the PyTorch framework on the NVIDIA 4090 GPU platform. In the embodiment, the model of the present invention is trained on four standard datasets, namely CAMO, CHAMELEON, COD10K, and NC4K, and the input images are uniformly adjusted to 352×352 pixels. During the model training and testing process, the momentum factor is set to 0.9, the weight decay coefficient is 0.0005, the initial learning rate is set to 1e-4, and the batch size is configured to 10.
[0060] After the model training is completed, the model is tested and the effectiveness of the model is verified using public evaluation metrics. The following are the evaluation metrics used in the present invention:
[0061] (1) Mean Absolute Error (MAE):
[0062]
[0063] (2) Structural similarity (S m ) between the evaluation region awareness Sr and the object awareness So:
[0064]
[0065] (3) Combine the influence of precision and recall through a weighted sum (F β ):
[0066]
[0067] (4) Combine local pixel values with the image-level average (Em):
[0068]
[0069] To better understand the technical solution of the present invention, the following further details the above four core modules with reference to the accompanying drawings.
[0070] 1. Difference Feature Interaction Module (DFIM)
[0071] The design purpose of the Difference Feature Interaction Module DFIM is to fuse the global features extracted by the Transformer branch and the local features extracted by the CNN branch through the cross-attention mechanism, so as to introduce global context information while retaining local details. This process can significantly enhance the multi-scale feature information crucial for the recognition of camouflaged targets and provide a richer and more accurate data basis for subsequent modules.
[0072] Specifically, in the sub-image processing branch (2× scaling), we use the Res2Net50 architecture as the feature extraction backbone network to sequentially generate four groups of multi-scale feature maps y i (i = 1, 2, 3, 4). In the main image processing branch (1× original scale), the Pyramid Vision Transformer (PVT) architecture is used to extract four-stage feature representations f i (i = 1, 2, 3, 4). Considering that low-resolution features are prone to significant information attenuation during transmission, we do not use the y4 feature at this stage.
[0073] Compared with other simple feature fusion methods, the Difference Feature Interaction Module proposed by the present invention can effectively utilize the global semantic features extracted by the Transformer to guide and combine the rich detail features extracted by the Res2Net50. Specifically, the present invention performs spatial transformation operations such as bilinear interpolation on the feature groups to align the feature resolution of the sub-image processing branch with that of the main image processing branch. The input image resolutions of the sub-branch and the main branch are different. After being downsampled by the network, the sub-branch feature y i and the main-branch feature f i+1(i = 1, 2, 3) achieves synchronous matching of spatial resolution at the corresponding levels. In the first two feature levels, through the synergistic effect of max pooling and average pooling, while retaining key feature information, the discriminative ability of features is enhanced. Deep features often contain richer semantic information and higher-level abstract representations. The simple fusion method of y3 extracted by Res2Net50 and f4 extracted by Transformer is not sufficient to fully explore the potential between these two types of features, and it is difficult to suppress the background noise in the Res2Net50 features. Based on this, the present invention deeply fuses the global context semantic information extracted from the main image branch and the local detail features of the sub-image branch through a cross-attention mechanism to achieve refined feature enhancement for each pixel point. The specific architecture of the differential feature interaction module is as shown in Figure 3 the dotted box in the upper right corner, and the formula is as follows:
[0074]
[0075] θ(f4) refers to converting f4 ∈ R H×W×C to f′4 ∈ R H×W×C through flattening and permutation operations, and then obtaining using layer normalization and linear transformation. At the same time, the process of generating by θ(y3) is the same as the principle of θ(f4), and the process is similar.
[0076] 2. Detail Feature Fusion Module (DFFM)
[0077] Different from the strategy of simply concatenating (Concat) adjacent feature layers and then performing convolution operations in traditional methods, the present invention transforms deep semantic features into dynamic filters and modulates them element by element with the features of the current layer. The specific implementation includes two key steps: First, adaptively weight the current features in the spatial and channel dimensions using deep semantic features to effectively suppress background interference; Second, retain the original feature information through residual connections to ensure the integrity of feature representation. Figure 4 Details show the architecture design and information flow process of this module. The architecture design of the detail feature fusion module proposed by the present invention introduces a bottom-up dense connection mechanism in the feature fusion module, fuses features at different levels to form a semantic filter, and effectively suppresses background interference and retains the original information through element-wise multiplication and residual connections to generate a rough prediction map.
[0078] The feature representation after being processed by the differential feature interaction module is F i(i = 1, 2, 3, 4). Considering F4 as the top - level output, we directly apply a 3×3 convolution on it for feature refinement to obtain the optimized feature representation F′4. For the F3 feature layer, first, a learnable spatial filtering operation is performed on F4 to generate a feature modulation filter Subsequently, the filter is applied to the current feature through element - wise multiplication, and at the same time, a weight adaptive learning module is constructed by combining the channel attention mechanism to achieve the significance evaluation of feature channels and dynamic weight allocation. Different from F4, F3 uses deformable convolution to better capture the features of irregularly camouflaged objects. Similarly and are also formed in a similar way, and the formula is as follows:
[0079] F′4 = Conv3(F4)#(3)
[0080]
[0081] The rough prediction map output in the first stage is represented by F c :
[0082] F c = Conv3(F′1)#(10)
[0083] Conv3 and Conv1 represent 3×3 and 1×1 convolutions respectively, Up represents the up - sampling operation, [;] represents channel concatenation, represents element - wise multiplication, represents element - wise addition.
[0084] The detailed feature fusion module designed in the present invention fuses the features input from Transformer and Res2net50 respectively to obtain a shared mask feature. Since the shape of the camouflage instance is very irregular and difficult to obtain, the present invention uses deformable convolution in DFFM to enhance the ability to capture local features.
[0085] 3. Boundary Perception Module (BPM)
[0086] An effective edge prior plays an important role in the segmentation and localization in object detection. Therefore, it is necessary to combine high - level semantic or location information to more effectively mine the edge features related to camouflaged objects. Inspired by this, the present invention proposes a boundary perception module BPM to enhance the sensitivity to edge features, aiming to effectively extract edge information for more accurate localization of the camouflaged object position. In this module, we combine low - level features and high - level features (F1 - F4) to model the edge information related to the object, and the module architecture is as Figure 5As shown. Specifically, first, the number of channels of the input features is adjusted to 64 through a 1×1 convolutional layer, and then the features (F2 - F4) are upsampled to the same size as F1, and channel concatenation is performed to finally generate edge features. This stage can be expressed by the following formula:
[0087] f e = Concat(Conv1(F1), Up(Conv1(F2)), Up(Conv1(F3)), Up(Conv1(F4))) #(11)
[0088] where Up(·) represents the upsampling operation and Concat(·) represents the concatenation operation.
[0089] Next, we perform channel attention and residual connection on the obtained edge features to generate clearer and more accurate edge information:
[0090]
[0091] where BConv(·) represents a convolutional operation with a kernel size of 3×3, followed by a batch normalization layer and a ReLU activation function, and ⊕ represents the element-wise addition operation.
[0092] BPM outputs the edge feature Fe, which is used to guide the decoding process of the model and highlight the edge details.
[0093] 4. Interference Filtering Module (IFM)
[0094] Based on the analysis of the preliminary prediction results, we believe that there are mainly two sources of interference: one is that the camouflaged object is not detected; the other is that the non-camouflaged object is misdetected. To address the impact of these two types of interference on the detection results, the present invention innovatively proposes an interference filtering module IFM. The interference filtering module (IFM) of the present invention explicitly models and processes two types of interference factors by combining two different types of interference that may occur. The proposed IFM architecture introduces a dual-branch structure, uses a lightweight encoder to extract interference features, and further optimizes the prediction results through an attention mechanism and a refinement unit, significantly improving the model's ability to identify interference factors, reducing the cases of false detection and missed detection, thereby greatly enhancing the detection accuracy of the model for camouflaged targets and improving the final segmentation efficiency.
[0095] As Figure 6 shown, the present invention processes F c ∈R 64×H×W through the designed dynamic spatial attention mechanism (DSA) to obtain f nFeature. The DSA module constructs a spatial attention map through convolutional operations, which depicts the importance borne by each spatial position in the image. Subsequently, through per-pixel multiplication, this attention map is superimposed on the original feature map, thereby effectively suppressing the response intensity of interference objects in the spatial dimension. The DSA module can capture the relationships between spatial positions more quickly and accurately locate camouflaged objects, avoiding the occurrence of the first type of interference. To better utilize the f n feature, we generate the prediction map of f n Then, we concatenate f with F n in the channel dimension and feed it into the attention mechanism to generate the feature f c w . Next, we perform element-wise multiplication on f w and the extracted edge feature F e fn and retain the original feature information through residual connection, finally generating the enhanced feature representation F e fn . This feature enhancement mechanism can effectively strengthen the response in the target boundary region, enabling the network to more accurately identify potential target regions that are misclassified as the background. This process can be described as follows:
[0096]
[0097] f n = DSA(F c ) #(13)
[0097]
[0098] where, represents the binarization operation. GT represents the Ground Truth, the true classification of each pixel.
[0099] Similarly, we use the DSA module to process F c to extract the feature f p , then concatenate f p with F fn and feed it into two 3×3 refinement units to capture richer context information, thereby distinguishing incorrect detection regions. Then, the predicted feature that suppresses the fp interference is subtracted from Ffn, and finally, through processing with a 3×3 convolutional kernel, the prediction map Fp is obtained.
[0100] 5. Loss Function
[0101] Our network uses a multi-level supervision strategy. For the loss of the prediction map Fp, we follow the supervision paradigm of mainstream camouflaged object detection methods and adopt the BCE loss and the IOU loss, as shown in Equation (15). When performing edge detection, we use the dice loss L edgeTo distinguish the importance of different pixels, such as those near fine or distinct edges. This process is achieved by incorporating pixel intensity into the L1 loss, thereby overcoming the severe imbalance between edge information and background information and making the network more focused on key pixel information, as shown in Equation (16). For f p and f n loss, we use weighted BCE loss, denoted by L fp and L fn respectively, as shown in Equation (17). The expression of the loss function is as follows:
[0102]
[0103] Loss2 = βL edgs (P e , G e ) #(16)
[0104]
[0105] Therefore, the total loss can be defined as:
[0106]
[0107] λ1, λ2, and β are all coefficients for calculation; in the experiment, λ1 and λ2 are set to 10, and β is set to 3. N n and N p represent the number of positive pixels and negative pixels respectively. Among them, P e is the prediction map of the camouflaged object, and G e is the ground truth data of the edge of the camouflaged object.
[0108] The detection model of the present invention adopts a progressive feature extraction architecture. In the initial feature extraction stage, the network constructs a basic feature representation through multi-scale convolution operations, generating a primary feature map with global semantic information. In the second stage, edge information is introduced to assist feature extraction, and in the third stage, the feature map is refined through interference filtering by IFM. The present invention adopts a multi-resolution input strategy, different from the single encoder sharing strategy adopted by ZoomNet. We innovatively designed a two-stream feature extraction architecture. This architecture uses the Pyramid Vision Transformer (PVT) as the main-scale feature extractor, and at the same time combines Res2Net50 to construct a sub-scale feature extraction branch, and realizes the collaborative coding of multi-scale features through a two-stream parallel mechanism. Based on this, we designed a differential feature interaction module through an attention-guided feature alignment and fusion strategy to aggregate different features at these two scales, make full use of the advantages of each encoder, and extract more valuable semantic clues. Then, the extracted features are transmitted to the detailed feature fusion module, and efficient information transmission is achieved through bottom-up dense connections. Finally, the camouflaged object is accurately located by combining edge information and supervised by the ground truth, and finally transmitted to the designed two-branch interference filtering module to suppress redundant information and background noise.
[0109] To verify the effectiveness and practicality of the present invention, the present invention was systematically compared and evaluated with 10 current leading camouflaged object detection algorithms on four public COD datasets, namely CAMO, CHAMELEON, COD10K, and NC4K, and a detailed visual analysis of the experimental results was carried out. Table 1 shows the test results of the present invention on 4 public datasets (black bold represents the best, and black bold with a horizontal line represents the second best). The experimental results show that the present invention shows strong competitiveness in the four datasets compared with other methods. The present invention is significantly better than the existing state-of-the-art methods on multiple datasets, especially achieving significant improvements in key indicators such as F-measure and MAE, proving the effectiveness and superiority of this method. This method has achieved leading performance in the key evaluation indicators of the four benchmark datasets. In particular, in the CAMO dataset, compared with the sub-optimal MFCFNet method, this method achieved performance improvements of 0.9% and 2.2% in the structure similarity S m and F β indicators respectively. It is worth noting that even compared with the PopNet method that fuses depth prior information, this method shows significant advantages in all four test sets, fully proving the effectiveness and generalization ability of the proposed method.
[0110] Table 1
[0111]
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the applicant has described the present invention in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention, without departing from the spirit and scope of the technical solutions, shall be covered by the scope of the claims of the present invention.
Claims
1. A camouflaged target detection method based on multi-scale feature fusion and interference suppression, characterized in that: The steps are as follows: 1) Train a camouflage object detection model; the camouflage object detection model includes a pre-processing network module, a differential feature interaction module, a detailed feature fusion module, a boundary awareness module, and an interference filtering module; 1.1) The pre-processing network module includes a main image processing branch and a sub-image processing branch; The main image processing branch is a pyramid vision Transformer architecture, which is used to process the original image data to extract a four-stage feature representation f i (i = 1, 2, 3, 4); The sub-image processing branch is a convolutional neural network, which is used to process the image data after the original image is scaled by a factor of two to sequentially generate four sets of multi-scale feature maps y i (i = 1, 2, 3, 4); 1.2) Align the feature resolution of the sub-image processing branch with that of the main image processing branch through a bilinear interpolation spatial transformation operation, i.e., feature y i and feature f i+1 (i = 1, 2, 3) achieve synchronous matching of the spatial resolution at the corresponding levels; Among them, through the cooperation of max-pooling and average-pooling for feature y1 and f2, and feature y2 and f3 respectively, feature Z2 and feature Z3 are correspondingly obtained to enhance the discriminative ability of the features while retaining the key feature information; Feature y3 and f4 are deeply fused through the differential feature interaction module, and the fused output is denoted as feature Z4; 1.3) Feature Z4 undergoes a convolution operation to obtain feature F4; Feature F4 and feature Z3 jointly undergo a convolution operation to obtain feature F3; Feature F3 and feature Z2 jointly undergo a convolution operation to obtain feature F2; Feature F2 and feature f1 jointly undergo a convolution operation to obtain feature F1; 1.4) Transmit features F1, F2, F3, and F4 to the detailed feature fusion module and the boundary perception module simultaneously; after fusion by the detailed feature fusion module, output the rough prediction result F c ; the boundary perception module outputs the edge feature F e ; 1.5) Transmit the preliminary prediction result F c and the edge feature F e to the interference filtering module together. The interference filtering module introduces a dual-branch structure and explicitly models and processes interference factors through a lightweight encoder and an attention mechanism, and outputs the final prediction map F p . This final prediction map F p is the output of the camouflaged target detection model; 2) Input the RGB image to be detected into the camouflage object detection model, and the output of the camouflage object detection model is the detected camouflage object map.
2. The camouflaged target detection method based on multi-scale feature fusion and interference suppression according to claim 1, characterized in that: In step 1.2), feature y3 and f4 are deeply fused through the differential feature interaction module to obtain feature Z4, which is specifically carried out according to the following formula: θ(f4) refers to the process of transforming f4 ∈ R C×H×W into f′4 ∈ R H×W×C through flattening and permutation operations, and then obtaining θ(y3) by using layer normalization and linear transformation. The process of generating 3. A camouflaged target detection method based on multi-scale feature fusion and interference suppression according to claim 1, characterized in that: The detailed feature fusion module outputs the rough prediction result F according to the process reflected by the following formula c , F′4=Conv3(F4)#(3) F c = Conv3(F′1) #(10) Conv3 and Conv1 represent 3×3 and 1×1 convolutions respectively, Up represents the upsampling operation, [;] represents channel concatenation, represents element-wise multiplication, and ⊕ represents element-wise addition.
4. A camouflaged target detection method based on multi-scale feature fusion and interference suppression according to claim 1, characterized in that: In step 1.2), the boundary awareness module outputs the edge feature F according to the following process e , First, the number of channels of the input features is adjusted to 64 through a 1×1 convolutional layer. Then, the features F2 - F4 are upsampled to the same size as the feature F1, and then channel concatenation is performed to obtain the preliminary edge feature f e ; The formula is expressed as follows: f e = Concat(Conv1(F1), Up(Conv1(F2)), Up(Conv1(F3)), Up(Conv1(F4)))#(11) Where Up(·) represents the upsampling operation, and Concat(·) represents the concatenation operation; Finally, channel attention and residual connection are performed on the preliminary edge feature f e to generate clearer and more accurate edge information, which is the edge feature F e , and the formula is expressed as follows: Where, BConv(·) represents a convolution operation with a kernel size of 3×3, followed by a batch normalization layer and a ReLU activation function, and ⊕ represents the element-wise addition operation.
5. The camouflaged target detection method based on multi-scale feature fusion and interference suppression according to claim 4, wherein: The boundary-aware module outputs the edge feature F e When this is the case, the loss function used is Loss2 = βL edge (P e , G e ); where G e is the prediction map of the camouflaged object, and G e is the ground truth data of the edge of the camouflaged object; β is a coefficient for calculation.
6. A camouflaged target detection method based on multi-scale feature fusion and interference suppression according to claim 1, characterized in that: The interference filtering module outputs the final prediction map F according to the following method p , 1.5.1) The interference filtering module processes the preliminary prediction result F through the dynamic spatial attention mechanism respectively to obtain the feature f c and the feature f n and the feature f p ; 1.5.2) Then the feature f n is concatenated with F c in the channel dimension and fed into the attention mechanism to generate the feature f w . Then, the feature f w and the edge feature F e are multiplied element-wise, and the original feature information is retained through the residual connection to generate the enhanced feature F fn ; 1.5.3) Concatenate f p with F fn , and send the concatenated feature into two 3×3 convolutional units to obtain the predicted feature that suppresses f p interference; 1.5.4) Subtract the predicted feature for suppressing f fn from F p and finally process it through a 3×3 convolution kernel to obtain the final predicted map F p .
7. A camouflaged target detection method based on multi-scale feature fusion and interference suppression according to claim 6, characterized in that: Step 1.5.1) The loss function used when processing the preliminary prediction result F through the dynamic spatial attention mechanism c is where N n and N p represent the numbers of positive pixels and negative pixels, respectively.
8. A method for detecting camouflaged targets based on multi-scale feature fusion and interference suppression according to claim 6, characterized in that: Step 1.5) Output the final prediction map F p The loss function used when
Citation Information
Cited By
Method, device and equipment for detecting port state of optical cable cross connecting cabinet and storage medium
CN121053138A