A method and device for quickly and accurately detecting camouflaged objects from thick to thin
Through the combination of the cross-level attention fusion module and the residual selective kernel module, the problems of blurred boundaries and inaccurate positioning in camouflage object detection are solved, and more accurate camouflage object detection is achieved, which improves detection performance.
Patent Information
- Application Number
- CN202310319154.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-28
AI Technical Summary
The existing camouflage object detection methods have problems such as blurred boundaries and inaccurate positioning in prediction, especially in complex natural scenes, which are difficult to accurately detect camouflage objects.
The detection method from coarse to thin is adopted, and the cross-level attention fusion module is used to integrate feature and boundary information at different scales, combined with the residual selective kernel module to enhance discriminative feature representation, and the weighted hybrid loss supervision network is used to achieve accurate camouflage object detection.
It improves the accuracy and clarity of camouflage object detection, significantly improves detection performance, especially in multi-scale and complex scenarios.
Smart Images

Figure CN116363467B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of camouflaged object detection, and specifically relates to a method and device for quickly and accurately detecting camouflaged objects from coarse to fine. Background Art
[0002] Camouflaged object detection (COD) aims to discover objects that attempt to hide their textures into the surrounding environment. However, existing COD methods still have problems of blurred boundaries and inaccurate localization in prediction.
[0003] The purpose of camouflaged object detection (COD) is to accurately detect and segment objects hidden in the surrounding environment. Camouflaged objects have the ability to make their colors, textures or patterns similar to the environment, so camouflaged objects are usually difficult to detect. The results of COD can be used as the initial step for many visual tasks, such as image segmentation / editing and retrieval, visual tracking and photo synthesis. In addition, the research of COD has direct applications, such as art, polyp segmentation, defective product detection and weapon detection.
[0004] However, compared with the general object detection / segmentation tasks in computer vision, COD is a very challenging task because camouflage strategies work by deceiving and misleading the observer's visual perception system. This is caused by the following reasons: (i) blurred object boundaries, (ii) highly similar textures between the foreground and the background, (iii) different scales and shapes of the objects.
[0005] To address these challenges, early COD methods were mainly based on handcrafted low-level features (such as texture, color and motion), and these features have limited ability to distinguish camouflaged or non-camouflaged objects, so these methods can only be used in a few simple scenarios. Recently, many important efforts have been made to detect / segment camouflaged objects in complex natural scenes. In addition, many COD datasets have been constructed and shared for training deep learning models. Although these deep learning-based efforts have achieved remarkable achievements so far, few have fully considered the above challenges, so there are still inaccurate and incomplete boundaries in the prediction maps. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and device for quickly and accurately detecting camouflaged objects from coarse to fine, so as to solve at least one of the above problems existing in the prior art.
[0007] Based on the above purpose, one or more embodiments in the present application provide a method for quickly and accurately detecting camouflaged objects from coarse to fine, which includes fusing multi-scale context information extracted by a backbone, and obtaining a rough prediction sub-network and a fine prediction sub-network for obtaining accurate camouflaged target detection results;
[0008] The rough prediction sub-network uses a cross-level attention fusion module to integrate features of different scales and fuse boundary information, and then makes an initial rough prediction; the cross-level attention fusion module includes four branches for capturing long-range semantic dependencies, precise position information, global features, and local features respectively;
[0009] The fine prediction sub-network uses a residual selective kernel module to enhance discriminative feature representation, and then uses two cross-level attention fusion modules to further capture multi-scale textures and obtain a fine-grained prediction result; the residual selective kernel module is constructed by stacking selective kernel networks to enhance the representation ability of the network.
[0010] Based on the above technical solutions of the present invention, the following improvements can also be made:
[0011] Optionally, in the rough prediction sub-network, ResNet-50 is used as the backbone to extract multi-level features, and the multi-level features include f1, f2, f3, and f4. The last three-level features are fed into three 3×3 convolutional blocks with a dilation rate of 3 to effectively increase the receptive field and capture more context feature representations. Then, the cross-level attention fusion module is added to each convolutional layer to fuse these multi-scale features to obtain a rough prediction.
[0012] Optionally, in the fine prediction sub-network, the overall attention module is used to merge the rough prediction features and f2 features to obtain refined features, and two new convolutional block networks are fed after the rough prediction sub-network; then three 3×3 convolutional blocks with a dilation rate of 3 and three residual selective modules are applied to enhance the context features; then the cross-level attention fusion module is used to further aggregate the multi-scale context features to obtain a fine prediction.
[0013] Optionally, in the cross-level attention fusion module, the high-level features are first upsampled twice, and added element-wise with the low-level features to obtain the context feature X′; two spatial ranges of the pooling kernel are used to encode each channel along the horizontal and vertical coordinates respectively; then the intermediate features along the spatial dimension are divided into separate tensors, and a 1×1 convolutional layer and an activation function are added respectively; finally, the following long-range semantic dependencies and precise position information are obtained from two directions, denoted as A1:
[0014]
[0015] Where G h and G w are learned attention weights;
[0016] Two branches are further divided from X′ to obtain global features and local features. The features of the two branches are respectively operated by point convolution, ReLU, and another point convolution. The obtained global and local attention weights are described as follows:
[0017]
[0018] where G(X′) and L(X′) represent global and local channel contexts, and σ is the sigmoid function;
[0019] Finally, the cross-level attention fusion feature is obtained:
[0020] F fusion = Conv 3×3 (C(F′, (A l + A r ), F″)) (3)
[0021] where C is the concatenation operation, Conv 3×3 represents a 3×3 convolutional layer, and F′ and F″ are low-level features and high-level features respectively;
[0022] Finally, the rough prediction is obtained through the following operation:
[0023]
[0024] where f4 is the multi-level feature, is the cross-level attention feature after fusion in the rough prediction stage.
[0025] Optionally, in the residual selective kernel module, the selective kernel convolution is integrated into the residual block. The i-th residual block in the j-th residual group is defined as:
[0026] F j,i = F j,i-1 + SK j,i (X j,i )·X j,i (5)
[0027] where F j,i and F j,i-1 are the input and output of the residual selective kernel module respectively, X j,i is the feature to be learned from the input, and SK j,i is the SK convolution;
[0028] Then, the residual feature is obtained through two stacked convolutional layers as follows:
[0029]
[0030] where and It is the weight set of two stacked convolutional layers in the residual selective kernel module.
[0031] Finally, the refined prediction is obtained through the following operations:
[0032]
[0033] where f4′ is the multi-level feature, is the cross-layer attention feature after fusion in the refined prediction stage.
[0034] Optionally, after obtaining the refined prediction result, the weighted consistency enhancement loss and the weighted boundary loss are combined with the weighted binary cross-entropy and the weighted IoU loss to supervise the network to obtain more accurate predictions.
[0035] Optionally, the weighted consistency enhancement loss is represented by the function as follows:
[0036]
[0037] where TP, FP, and FN represent true positive, false positive, and false negative respectively, |·| represents the calculated area, γ is a hyperparameter; weights are assigned to the pixel at position (i, j) in the image, and α ij represents the attention weight at position (i, j). Hard pixels correspond to larger attention weights, and simple pixels will be assigned smaller attention weights.
[0038] Optionally, the weighted boundary loss is defined as:
[0039]
[0040] where P k,i,j and E k,i,j represent the values of the pixels in the predicted boundary map P k and the edge map E k respectively.
[0041] Optionally, it is defined as the weighted mixture loss and is expressed by the following formula:
[0042]
[0043] where, is the weighted consistency enhancement loss, is the weighted boundary loss, is the weighted binary cross-entropy loss, is the weighted IoU loss.
[0044] As a second aspect of the present invention, there is provided a device for quickly and accurately detecting camouflaged objects from coarse to fine, including a storage unit and a processing unit. The storage unit is used to store computer programs, and the processing unit is used to execute the method for quickly and accurately detecting camouflaged objects from coarse to fine as described in any one of the above by means of the computer programs stored in the storage unit.
[0045] The beneficial effects of the present invention are as follows: The present invention provides a method and a device for quickly and accurately detecting camouflaged objects from coarse to fine, and proposes a novel coarse-to-fine network, which includes a coarse prediction sub-network and a fine prediction sub-network. It can fuse multi-scale context information extracted by the backbone and obtain a fine prediction for accurate camouflaged target detection. A cross-level attention fusion module is used to fuse multi-scale features, and a residual selective kernel module is used to further enhance the discriminative feature representation. Both are beneficial to improving the COD performance. A weighted hybrid loss is proposed, and by assigning different weights to different positions, more attention can be paid to hard pixels, thereby promoting the network to learn clear boundaries and the accurate positions of camouflaged objects. A large number of experiments show that the state-of-the-art camouflaged object detection performance is achieved on three benchmark datasets in terms of four evaluation metrics. Qualitative and quantitative results prove the effectiveness of the present invention. Description of the Drawings
[0046] Figure 1 It is a schematic diagram of a visual example of camouflaged object detection for the method and device for quickly and accurately detecting camouflaged objects from coarse to fine according to an embodiment of the present invention.
[0047] Figure 2 It is a schematic diagram of the architecture of the method and device for quickly and accurately detecting camouflaged objects from coarse to fine according to an embodiment of the present invention.
[0048] Figure 3 It is a schematic diagram of a qualitative comparison between the method and device for quickly and accurately detecting camouflaged objects from coarse to fine according to an embodiment of the present invention and five leading models.
[0049] Figure 4 It is a schematic diagram of a quantitative comparison between the method and device for quickly and accurately detecting camouflaged objects from coarse to fine according to an embodiment of the present invention and 27 state-of-the-art COD methods on four benchmark datasets.
[0050] Figure 5 It is a schematic diagram of a quantitative comparison of the ablation study of the method and device for quickly and accurately detecting camouflaged objects from coarse to fine according to an embodiment of the present invention on the CAMO and COD10K datasets.
[0051] Figure 6 It is a schematic diagram of a comparison of the average test speed and complexity between the method and device for quickly and accurately detecting camouflaged objects from coarse to fine according to an embodiment of the present invention and other cutting-edge models. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0053] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second", and similar terms used in one or more embodiments of the present application do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before the term cover the elements or objects listed after the term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0054] Reference Figures 1 - 6 , one or more embodiments of the present application provide a method and apparatus for quickly and accurately detecting camouflaged objects from coarse to fine, which include a rough prediction sub-network and a fine prediction sub-network that fuse multi-scale context information extracted by a backbone and obtain accurate detection results of camouflaged targets;
[0055] The rough prediction sub-network uses a cross-level attention fusion module to integrate features of different scales and fuse boundary information, and then performs an initial rough prediction; the cross-level attention fusion module includes four branches respectively used to capture long-range semantic dependencies, precise position information, global features, and local features;
[0056] The fine prediction sub-network uses a residual selective kernel module to enhance discriminative feature representations, and then uses two cross-level attention fusion modules to further capture multi-scale textures and obtain fine-grained prediction results; the residual selective kernel module is constructed by stacking selective kernel networks to enhance the representation ability of the network.
[0057] It can be understood that in this embodiment, a new COD framework called Coarse-to-Fine Network (CFNet) is proposed, which achieves accurate camouflaged object detection with high-quality boundaries and greatly improves the existing camouflaged object detection performance. CFNet is a boundary refinement architecture for generating fine camouflaged objects, which consists of two sub-networks: a coarse prediction sub-network and a fine prediction sub-network. In the coarse prediction sub-network, an encoder-decoder network with a newly designed cross-level attention fusion module (CAFM) is used to integrate features of different scales, fuse boundary information, and then perform an initial coarse prediction. In the fine prediction sub-network, a new residual selective kernel module (RSKM) is developed to enhance discriminative feature representation, and then two CAFMs are used to further capture multi-scale textures and obtain fine-grained prediction results. Specifically, CAFM consists of four branches, which respectively capture remote semantic dependencies, precise location information, global and local features. RSKM is constructed by stacking selective kernel networks to enhance the representation ability of CNN. In addition, a weighted consistency enhancement (wCE) loss and a weighted boundary (wB) loss are designed, combined with the commonly used weighted binary cross-entropy (wBCE) and weighted IoU (wIoU) losses, called weighted hybrid (wH) loss, which can supervise the network to obtain more accurate predictions.
[0058] The above technical solutions of this embodiment can achieve at least the following effects:
[0059] A novel coarse-to-fine network (CFNet) is proposed, which includes a coarse prediction sub-network and a fine prediction sub-network, can fuse multi-scale context information extracted by the backbone, and obtain fine predictions for accurate camouflaged object detection. We designed a cross-level attention fusion module (CAFM) to fuse multi-scale features, and a residual selective kernel module (RSKM) to further enhance discriminative feature representation. Both are beneficial to improving COD performance.
[0060] A weighted hybrid loss is proposed, which can pay more attention to hard pixels by assigning different weights to different positions, thus promoting the network to learn clear boundaries and the accurate positions of camouflaged objects.
[0061] In a possible embodiment, in the coarse prediction sub-network, ResNet-50 is used as the backbone to extract multi-level features, the multi-level features include f1, f2, f3, and f4, and the last three levels of features are fed into three 3×3 convolutional blocks with a dilation rate of 3 to effectively increase the receptive field and capture more context feature representations, and then the cross-level attention fusion module is added to each convolutional layer to fuse these multi-scale features to obtain a coarse prediction.
[0062] In the fine prediction sub-network, the overall attention module is used to merge the coarse prediction features and f2 features to obtain refined features, which are then fed into two new convolutional block networks after the coarse prediction sub-network; then three 3×3 convolutional blocks with a dilation rate of 3 and three residual selective modules are applied to enhance the context features; finally, the cross-level attention fusion module is used to further aggregate the multi-scale context features to obtain the fine prediction.
[0063] It can be understood that the overall architecture of the proposed CFNet is as Figure 2 shown, which consists of a coarse prediction sub-network and a fine prediction sub-network. Specifically, in the coarse prediction sub-network, we use ResNet-50 as the backbone to extract multi-level features (i.e., f1, f2, f3, f4), and the last three-level features are fed into three 3×3 convolutional blocks with a dilation rate of 3 to effectively increase the receptive field and capture more context feature representations. Then, the proposed cross-level attention fusion module (CAFM) is added to each convolutional layer to fuse these multi-scale features to obtain the coarse prediction. After that, in the fine prediction sub-network, we use the holistic attention (HA) module to merge the coarse prediction map and the feature map f2 to obtain refined features, which are then fed into two new convolutional block networks after the coarse sub-network. Next, three 3×3 convolutional blocks with a dilation rate of 3 and three RSKMs are applied to enhance the context features. Finally, the CAFM is used to further aggregate the multi-scale context features to obtain the fine prediction.
[0064] In a possible implementation, when the camouflaged object has different sizes and shapes, the performance of COD usually decreases significantly. To solve this problem, we designed a cross-level attention fusion module (CAFM) to better fuse inconsistent semantic and multi-scale context features. By aggregating multi-scale information, CAFM can not only capture different direction and position information, but also capture global and local features, thus reducing the high variation of object size.
[0065] Specifically, as Figure 2 shown, the high-level features are upsampled twice and added element-wise with the low-level features to obtain the context feature X′. Two spatial extents of pooling kernels (1, W) and (H, 1) are used to encode each channel along the horizontal and vertical coordinates respectively. Then, we divide the intermediate features along the spatial dimension into separate tensors and add a 1×1 convolutional layer and an activation function respectively. Finally, the following long-range dependencies and precise position information can be obtained from two directions, denoted by A l :
[0066]
[0067] where G h and G wis the learned attention weight;
[0068] Meanwhile, two branches are further divided from X′ to obtain global features and local features, and these branches can handle multi-scale problems. As Figure 2 shown, the features of the two branches are respectively operated through point convolution, ReLU, and another point convolution. The difference is that the global branch has an additional global average pooling to emphasize contextual information. The obtained global and local attention weights can be described as follows:
[0069] A r = X′ × σ(G(X′) + L(X′)) (2)
[0070] where G(X′) and L(X′) represent global and local channel contexts, and σ is the sigmoid function;
[0071] Finally, the cross-level attention fusion feature is obtained:
[0072] F fusion = Conv 3×3 (C(F′, (A l + A r ), F″)) (3)
[0073] where C is the concatenation operation, Conv 3×3 represents a 3×3 convolutional layer, and F′ and F″ are the low-level feature and the high-level feature respectively.
[0074] Finally, the coarse prediction is obtained through the following operation:
[0075]
[0076] where f4 is the multi-level feature, is the cross-level attention feature after fusion in the coarse prediction stage.
[0077] In the coarse prediction network, a cross-level attention fusion module is proposed to fuse these multi-scale context features and obtain the coarse prediction. However, due to the highly similar textures between the foreground and the background, the rough prediction usually has an ambiguous semantic part. To overcome this challenge, a residual selective kernel module based on selective kernel convolution is developed to further enhance the multi-scale discriminative features. Selective kernel convolution is a dynamic selection mechanism in the form of soft attention, allowing each neuron to adaptively adjust its receptive field size according to the multi-scale input information. Therefore, selective kernel convolution can accurately capture objects of different scales.
[0078] However, features at different scales contain rich context information. If these features are treated equally across channels, the representational ability of the CNN will be hindered. Therefore, a residual selective kernel module is proposed to enable the network to focus on more discriminative features. Specifically, we integrate selective kernel convolution into the residual block (RB). As Figure 2 shown, the i-th RB in the j-th residual group (RG) can be expressed as:
[0079] F j,i = F j,i-1 + SK j,i (X j,i )·X j,i (5)
[0080] where F j,i and F j,i-1 are the input and output of the residual selective kernel module respectively, X j,i is the feature to be learned from the input, and SK j,i is the SK convolution;
[0081] Then, the residual features are obtained through two stacked convolutional layers as follows:
[0082]
[0083] where and are the weight sets of the two stacked convolutional layers in the residual selective kernel module.
[0084] Finally, the refined prediction is obtained through the following operation:
[0085]
[0086] where f4′ is the multi-level feature, is the cross-layer attention feature after fusion in the refined prediction stage.
[0087] After obtaining the refined prediction result, the weighted consistency enhancement loss and the weighted boundary loss are combined with the weighted binary cross-entropy and the weighted IoU loss to supervise the network to obtain more accurate predictions.
[0088] The weighted consistency enhancement loss is represented by the function as follows:
[0089]
[0090] Where TP, FP, and FN represent true positive, false positive, and false negative respectively, |·| represents calculating the area, and γ is a hyperparameter; weights are assigned to the pixels at position (i, j) in the image, and α ij represents the attention weight at position (i, j). Hard pixels correspond to larger attention weights, and simple pixels will be assigned smaller attention weights.
[0091] Define the weighted boundary loss as:
[0092]
[0093] where P k,i,j and E k,i,j represent the values of the pixels in the predicted boundary map P k and the edge map E k respectively.
[0094] Define it as the weighted mixture loss, which is expressed by the following formula:
[0095]
[0096] where, is the weighted consistency enhancement loss, is the weighted boundary loss, is the weighted binary cross-entropy loss, is the weighted IoU loss.
[0097] Experiment
[0098] In this embodiment, the implementation details include that the proposed CFNet is implemented using PyTorch. Use the pre-trained ResNet-50 on ImageNet as the base model, and the rest are randomly initialized. The input images and the ground truth are both adjusted to 352×352 for training and testing. During training, use Adam to optimize the network with a learning rate of 5e-5, divide it by 10 every 50 epoches, and apply various augmentation strategies (i.e., horizontal flipping, boundary cropping, and random rotation) to avoid overfitting. Training is carried out on a PC equipped with an Intel(R) Xeon(R) Gold 6,240, 2.60GHz CPU and an NVIDIA Tesla V100 GPU (with 32GB of memory). The whole training takes about 5 hours, and the batch size is 36.
[0099] Datasets. Experiments were conducted on four public COD benchmark datasets. Chameleon contains 76 images collected through the Google search engine with the keyword "camouflage animals". The camouflage covers 8 categories and has 1,250 camouflage images, divided into 1K for training and 0.25K for testing. COD10K is currently the most challenging COD dataset, consisting of 5,066 camouflage images (3,040 for training and 2,026 for testing), covering 5 superclasses and 69 subclasses. We used camouflage and COD10K as the training set (4,040 images), and evaluated the entire Chameleon dataset as well as the test sets of camouflage and COD10K. NC4K is another large COD test dataset, including 4,121 images from the Internet. It contains many more samples than previous evaluation benchmarks and can well reveal the generalization ability of the COD model.
[0100] Evaluation Metrics. Four popular evaluation metrics were used to comprehensively evaluate the performance of the technical solution of the embodiment. Structure-measure (S α ) is used to focus on evaluating the structural information between the prediction and the ground truth. E-measure (E φ ) has been proven to be related to human visual perception and is used to evaluate the overall and local accuracy of COD. Weighted F-measure can provide more reliable evaluation results, considering both weighted precision and weighted recall. Mean Absolute Error (M) is widely used to evaluate the average pixel-level relative error between the prediction map and the ground truth.
[0101] Comparison with the State of the Art:
[0102] The proposed technical solution was compared with 27 state-of-the-art SOD / COD methods, including NLDF, PiCANet, BASNet, CPD, PoolNet, EGNet, F3Net, SCRN, CSNet, SSAL, UCNet, MINet, ITSD, PraNet, ANetSRM, SINet, RMGL, ERRNet, TINet, UGTR, PFANet, SLSR, MirrorNet, BgNet, C2F-Net, BGNet, and SINet-v2. For fair comparison, these results mainly come from PFNet or are generated by models retrained with the code released by the authors of the comparative methods.
[0103] Figure 4Summarizes the quantitative results of CFNet compared to 27 other state-of-the-art methods on four benchmark datasets. As can be seen from the results, the proposed model consistently and significantly outperforms recent methods on all datasets without relying on any post-processing tricks. In particular, we can see that our method outperforms all previous methods in all four evaluation metrics. For example, our method outperforms the second-best method BgNet by 2.2%, 1.2%, 2.9%, and 0.8% on the CAMO, CHAMELEON, COD10K, and NC4K datasets respectively. Compared with the fast camouflaged object detection (ERRNet), our method outperforms it by 8.1%, 2.2%, and 6.3% on the CAMO, CHAMELEON, and COD10K datasets respectively. In addition, we also evaluated the performance of our Res2Net-50-based method. Compared with the state-of-the-art COD method SINet-v2, our method improves by 1.8%, 1.8%, and 2.5% on the CAMO, CHAMELEON, and COD10K datasets respectively. In terms of, it outperforms it by 8.1%, 2.2%, and 6.3% on the CAMO, CHAMELEON, and COD10K datasets respectively. In addition, we also evaluated the performance of our Res2Net-50-based method. Compared with the state-of-the-art COD method SINet-v2, our method improves by
[0104] Qualitative evaluation Figure 3 shows the qualitative results of our method compared with other methods. They intuitively introduce the superior performance of the method from different aspects. For example, we can see that most of the comparison methods have some blurred boundaries (e.g., columns 5 to 8). When facing objects that are extremely similar in the foreground and background (e.g., rows 4 and 5), most methods fail to detect the complete object. When facing objects with different scales, sizes, and shapes, most existing methods detect irrelevant background regions (e.g., rows 6 and 7). On the contrary, our method can continuously produce more accurate and complete visual maps with clear and coherent details.
[0105] Ablation verification
[0106] The effectiveness of the rough-to-fine architecture. We conducted ablation experiments to verify the effectiveness of the coarse-fine structure. In Figure 5 it, we removed the fine prediction sub-network (d) and the rough prediction sub-network (e) and compared it with the complete architecture (f). We can find that adding the fine prediction sub-network to the rough prediction sub-network greatly improves the performance. This shows that the fine prediction sub-network is effective in improving the performance of accurate COD. This is because in the rough prediction stage, we mainly try to use low-level networks to locate objects and pay less attention to context features. Therefore, we get rough predictions. On the contrary, when we feed these rough predictions into the fine prediction stage, the performance will be greatly improved.
[0107] Among them, "B" represents that ResNet-50 deletes all the proposed modules and only uses the pre-trained model through simple connections. "W / O fine" and "W / O coarse" respectively indicate without the fine prediction sub-network and the coarse prediction sub-network.
[0108] Effectiveness of CAFM and RSKM. We studied the importance of CAFM and RSKM. Figure 5 In it, we observed that the performance was improved by adding CAFM or RSKM to the basic model. (a) Especially when adding CAFM to the basic model, the performance was significantly improved. For example, in the CHAMELEON dataset, adding the cross-level attention fusion module resulted in performance improvements of 1.8%, 4.0%, and 2.1%. To explore whether RSKM could really achieve performance, we studied the effectiveness of RSKM. (a) and (c) We can see that (c) is better than the performance. (a) In the CAMO, CHAMELEON, COD10K, and NC4K datasets rose by 1.1%, 2.4%, 3.6%, and 2.0% respectively. These improvements indicate that introducing the residual selective kernel module can enable our model to achieve accurate camouflaged object detection.
[0109] Effectiveness of the weighted hybrid loss. To verify the effectiveness of the WH loss we proposed, we conducted a set of ablation studies on different losses based on our CFNet. From the results in Table II, it can be seen that higher performance can be obtained in (h) when using the weighted hybrid loss we proposed.
[0110] Computational complexity. We compared the computational complexity of our CFNet and 8 other state-of-the-art models for a more comprehensive evaluation. We resized all images to 352×352 to test the inference speed and MAC based on the NVIDIA Tesla V100 GPU. As Figure 6 shown, we reported the average inference speed to avoid the influence of random factors. Although the parameters of our CFNet are not the smallest, we can still obtain a competitive inference speed.
[0111] Conclusion, we present a novel coarse-to-fine framework, called Coarse-to-Fine Network (CFNet), and propose a new weighted hybrid loss for accurate camouflaged object detection. CFNet consists of a coarse prediction sub-network and a fine prediction sub-network to obtain fine predictions. Specifically, a cross-level attention fusion module (CAFM) is designed to better integrate inconsistent semantic and multi-scale context features, and a residual selective kernel module (RSKM) is developed to further enhance discriminative feature representations. To obtain clearer boundaries and accurate localization, we propose a new weighted hybrid loss. Extensive experiments conducted on four widely used datasets verify that our proposed method outperforms state-of-the-art models under four evaluation metrics.
[0112] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0113] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for quickly and accurately detecting camouflaged objects from thick to thin, characterized in that, It includes a rough prediction sub-network and a fine prediction sub-network that fuse multi-scale context information extracted from the backbone and obtain accurate detection results of camouflaged objects; The rough prediction sub-network uses a cross-level attention fusion module to integrate features of different scales and fuse boundary information, and then makes an initial rough prediction; the cross-level attention fusion module includes four branches respectively used to capture remote semantic dependencies, accurate position information, global features, and local features; The fine prediction sub-network uses a residual selective kernel module to enhance discriminative feature representation, and then uses two cross-level attention fusion modules to further capture multi-scale textures and obtain fine-grained prediction results; the residual selective kernel module is constructed by stacking selective kernel networks to enhance the representation ability of the network; In the cross-level attention fusion module, first, the high-level features are upsampled twice, and added element-wise with the low-level features to obtain the context feature X'; two spatial ranges of the pooling kernel are used to encode each channel along the horizontal and vertical coordinates respectively; then the intermediate features along the spatial dimension are split into separate tensors, and a 1×1 convolutional layer and an activation function are added respectively; finally, the following long-range semantic dependencies and precise position information are obtained from two directions, denoted by A l Indicates: A l (i,j) = X′(i,j) × G h (i) × G w (j) (1) where G h and G w are the learned attention weights; Two branches are further divided from X′ to obtain global features and local features. The features of the two branches are respectively operated through point convolution, ReLU, and another point convolution. The obtained global and local attention weights are described as follows: A r = X' × σ(G(X') + L(X')) (2) Where G(X′) and L(X′) represent global and local channel contexts, and σ is the sigmoid function; Finally, the cross-level attention fusion feature is obtained: F f = Conv 3×3 (C(F′,(A l + A r ),F″)) (3) Among them, C is a cascading operation, Conv3×3 represents a 3×3 convolutional layer, F′ and F″ are low-level features and high-level features respectively; A l represents remote semantic dependency relationships and precise location information; Finally, the rough prediction is obtained through the following operations: where f4 is a multi-level feature, which is the cross-layer attention feature after fusion in the coarse prediction stage; In the residual selective kernel module, the selective kernel convolution is integrated into the residual block. The i-th residual block in the j-th residual group is defined as: F j,i = F j,i-1 + SK j,i (X j,i )·X j,i (5) Among which F j,i and F j,i-1 are respectively the input and output of the residual selective kernel module, X j,i is the feature to be learned from the input, and SK j,i is the SK convolution; Then, the residual feature is obtained through two stacked convolutional layers as follows: Among them and are the weight sets of two stacked convolutional layers in the residual selective kernel module; Finally, the fine prediction is obtained through the following operations: where f4′ is a multi-level feature, which is the cross-layer attention feature after fusion in the fine prediction stage.
2. The method for quickly and accurately detecting a camouflaged object from thick to thin as described in claim 1, characterized in that, in In the rough prediction sub-network, ResNet-50 is used as the backbone to extract multi-level features. The multi-level features include f1, f2, f3, and f4. The last three-level features are fed into three 3×3 convolutional blocks with a dilation rate of 3 to effectively increase the receptive field and capture more context feature representations. Then, the cross-level attention fusion module is added to each convolutional layer to fuse these multi-scale features to obtain a rough prediction.
3. The method for quickly and accurately detecting a camouflaged object from thick to thin according to claim 2, characterized in that, In the fine prediction sub-network, the rough prediction feature and the f2 feature are merged through the use of an overall attention module to obtain a refined feature, and two new convolutional block networks are fed after the rough prediction sub-network; then three 3×3 convolutional blocks with a dilation rate of 3 and three residual selective modules are applied to enhance the context features; The cross-level attention fusion module is further used to aggregate multi-scale context features to obtain a fine prediction.
4. The method for quickly and accurately detecting a camouflaged object from thick to thin according to claim 1, characterized in that, After obtaining the fine prediction results, enhance the loss through weighted consistency and weighted boundary loss Combine weighted binary cross-entropy and weighted IoU loss to supervise the network to obtain more accurate predictions.
5. A method for quickly and accurately detecting camouflaged objects from thick to thin as described in claim 4, characterized in that, The weighted consistency enhancement loss is represented by the function as follows: where TP, FP, and FN represent true positive, false positive, and false negative respectively, |·| denotes the calculated area, and γ is a hyperparameter; weights are assigned to the pixels at position (i, j) in the image, and α ij represents the attention weight at position (i, j). Hard pixels correspond to larger attention weights, and simple pixels will be assigned smaller attention weights.
6. The method for quickly and accurately detecting a camouflaged object from thick to thin according to claim 5, characterized in that, The weighted boundary loss is defined as: Among them, P k,i,j and E k,i,j respectively represent the values of the pixels in the predicted boundary map P k and the edge map E k respectively.
7. The method for quickly and accurately detecting camouflaged objects from thick to thin according to claim 6, characterized in that it defines For the weighted mixture loss, it is represented by the following formula: Among them, is the weighted consistency enhancement loss, is the weighted boundary loss, is the weighted binary cross-entropy loss, is the weighted IoU loss.
8. A device for quickly and accurately detecting camouflaged objects from thick to thin, characterized in that, It includes a storage unit and a processing unit. The storage unit is used to store a computer program, and the processing unit is used to execute the method for quickly and accurately detecting camouflaged objects from rough to fine according to any one of claims 1-7 through the computer program stored in the storage unit.
Citation Information
Patent Citations
Bidirectional attention-based camouflage object detection method
CN113553973A
Camouflaged object segmentation method with distraction mining
US20220230324A1