A multi-level edge enhancement fusion method for camouflaged target segmentation
Through the multi-level edge enhancement fusion method, the Res2Net-50 network is used to extract multi-level features and perform fine fusion of edge features, which solves the edge blur and noise problems in camouflage target segmentation, and achieves more efficient camouflage target detection.
Patent Information
- Application Number
- CN202310986423.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-07
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-08-07
AI Technical Summary
The existing camouflage target segmentation method has degraded detection performance when facing challenging scenarios, mainly due to the loss of semantic information of high-level feature maps, the fullness of noise from the feature map extracted by the backbone network, and the blurred edges are difficult to segment, resulting in incomplete segmentation of the camouflage target.
The multi-level edge enhancement fusion method is adopted to extract multi-level features through the Res2Net-50 network, and edge extraction modules are used to obtain edge features. The residual texture enhancement modules are used to refine features, and top-down fusion is carried out through the edge guide fusion modules, combining global prior information to optimize the disguised target segmentation.
It significantly improves the detection performance of camouflage targets, with finer edges and more complete structures, and can effectively segment camouflage targets, which is better than the prediction results of multiple cutting-edge models in recent years.
Smart Images

Figure CN117036389B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a camouflaged target segmentation method with multi-level edge enhancement and fusion. Background Art
[0002] Camouflaged object segmentation is a newly emerging field in computer vision and a key branch of the computer vision segmentation task, attracting considerable research attention in recent years. Camouflaged objects are those that are highly similar to or obscured by the background. They often blend in with their surroundings, making their color and posture closely resemble those of the surroundings, thereby camouflaging themselves and making them difficult to detect. Examples include chameleons living in the desert, polar bears on the ice, and soldiers wearing camouflage uniforms. Camouflaged object segmentation aims to detect camouflaged objects in a visual scene, segment them from the background, and extract useful information about the camouflaged objects. Camouflaged object segmentation has numerous applications. Besides its inherent scientific value, it can also be used in computer vision (for search and rescue, and the discovery of rare animals), medical image segmentation (such as polyp segmentation, lung infection segmentation, and retinal image segmentation), agriculture (disaster detection, locust detection), art (realistic blending, entertainment art), and military reconnaissance. Dozens of existing camouflaged object segmentation methods exist, but early approaches were ineffective and labor-intensive. In recent years, advances in deep learning have significantly improved segmentation performance. Early approaches to camouflaged object recognition used heuristic prior feature extraction to generate segmentation maps. These prior features typically include color contrast, texture information, and center priors. However, these handcrafted features are difficult to extract, their selection is highly subjective, and they struggle to capture the connection between semantic and content information. Consequently, these methods often fail in real-world scenarios and suffer from poor generalization performance. To address this issue, deep learning-based camouflaged object segmentation methods have been proposed in recent years, demonstrating the impressive performance of convolutional neural networks. These deep learning-based model frameworks typically utilize ResNet, VGGNet, and Transformer as backbone networks. These networks are capable of extracting image features at various scales and high-level semantic information, thereby mining more effective features within the image. Finally, through the fusion of different modules and image restoration, the final segmented image of the camouflaged object is generated.
[0003] While the aforementioned methods have achieved promising results and improvements in camouflaged target segmentation, most methods experience reduced detection performance in challenging scenarios. This is primarily due to the following: 1. The high-level feature map extraction network is too deep, and the use of pooling layers causes the feature maps to lose some high-level semantic information, resulting in suboptimal detection results. Furthermore, when high-level semantic information is passed to shallower layers, the shallower layers do not receive sufficient high-level semantic information, which in turn reduces the network's detection capabilities. This lack of global information acquisition has become a challenge, and mitigating the loss of high-level semantic information has become a new area of research. 2. The feature maps extracted by the backbone network are often noisy. 3. Due to the lack of edges or the high degree of fusion between the camouflaged target's edges and the surrounding environment, existing methods often struggle to fully segment the camouflaged target and fail to effectively address edge blurring. For camouflaged targets, the foreground and background are highly similar. This high similarity between the camouflaged target and the background results in more image noise during feature extraction, making boundary information difficult to extract and less scale information to be obtained. This results in noise being incorporated into the fusion of multi-scale features, resulting in inadequate fusion and poor performance. To sum up, this project aims to solve the problems existing in the camouflaged target detection method, such as insufficient global information acquisition, less global information acquisition by shallow networks, insufficient multi-scale feature fusion, and insufficient contextual information extraction of existing methods, which lead to poor camouflaged target segmentation performance. By using deep learning methods, we conduct research on camouflaged target segmentation methods in terms of multi-level feature fusion and high-level feature extraction optimization, which is of great significance to the development of camouflaged target segmentation in the field of computer vision. Summary of the Invention
[0004] In response to the above-mentioned deficiencies in the prior art, the present invention provides a multi-level edge-enhanced fusion camouflaged target segmentation method that solves the problem that the existing methods are difficult to completely segment the camouflaged target due to the lack of edges or the high fusion of the camouflaged target edge with the environment, resulting in blurred edges.
[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a multi-level edge enhancement fusion camouflage target segmentation method, comprising the following steps:
[0006] S1. Extracting multi-level features of an input image, wherein the multi-level features include second to fifth features, and inputting the second to fifth features into an edge extraction module to obtain edge features;
[0007] S2. Input the second to fifth features into the residual texture enhancement module to obtain refined second to fifth features;
[0008] S3. Input the fifth feature into the pyramid pooling module to obtain global prior information;
[0009] S4. Input the edge features, the refined second to fifth features, and the global prior information into the edge-guided fusion module, and perform a top-down fusion operation through the edge-guided fusion module to obtain the final camouflaged target prediction result and complete the camouflaged target segmentation.
[0010] Furthermore: in S1, multi-level features of the input image are extracted through the Res2Net-50 network, and the number of channels of the second to fifth features are 256, 512, 1024 and 2048 respectively.
[0011] Further: In S1, the edge feature F is obtained e The specific expression is:
[0012]
[0013] Where, F o is the integrated feature, which is obtained by integrating the second to fifth features. 1×1 is a 1×1 convolution operation, is a multiplication operation, σ is a Sigmoid function, GAP(·) is a global average pooling operation, Conv1d(·) is a one-dimensional convolution operation with a convolution kernel of k, and k is the size of the adaptive convolution kernel. The specific expression is as follows:
[0014]
[0015] Where C is the set number of channels, |·| odd To take the nearest odd general formula.
[0016] The beneficial effect of the above further scheme is: the method of calculating edge features of the present invention takes into account local channel attention, highlights key channels, thereby removing redundant channels and noise, has lower complexity, and plays a key role in exploring effective edge semantics in noisy features.
[0017] Furthermore: the integrated feature F o The specific expression is:
[0018] F o =Conv 3×3 (CAT(F′2,F′3,F′4,F′5))
[0019] F′2=Conv 1×1 (F2)
[0020] F′3=UP(Conv 1×1 (F3)
[0021] F′4=UP(Conv 1×1(F4)
[0022] F′5=UP(Conv 1×1 (F5)
[0023] In the formula, CAT(·) is the concatenation operation, Conv 3×3 is a 3×3 convolution operation, UP(·) is an upsampling operation, F2 is the second feature, F3 is the third feature, F4 is the fourth feature, F5 is the fifth feature, F′2 is the second feature with 64 channels, F′3 is the third feature with 64 channels, F′4 is the fourth feature with 64 channels, and F′5 is the fifth feature with 64 channels.
[0024] Further: in S2, the residual texture enhancement module includes a main branch and first to fourth negative branches, wherein the inputs of the main branch and the first negative branch are input features, and the inputs of the second to fourth negative branches are the sum of the up-sampled output of the previous negative branch and the input features;
[0025] Wherein, the input feature is the second feature, the third feature, the fourth feature or the fifth feature;
[0026] The method for obtaining the refined second to fifth features is specifically a method for obtaining the refined input features, which is specifically:
[0027] After splicing the outputs of the first to fourth negative branches, the splicing result is convolved by 3×3 and added to the output of the group branch. The addition result is input into the ReLU function to obtain refined input features.
[0028] The beneficial effect of the above further scheme is: since the backbone network uses a large number of convolution operations, it is unable to extract features containing rich contextual information, which is not conducive to the segmentation of camouflaged targets. Therefore, when using the features extracted by the backbone network, it is necessary to refine the features first. Therefore, the present invention designs a residual texture enhancement module based on TEM to refine the features of the backbone network to provide more effective prior information in subsequent fusion.
[0029] Furthermore: the expansion rates of the first to fourth negative branches are 1, 3, 5 and 7 respectively;
[0030] The main branch and the first negative branch are both provided with a 1×1 convolution;
[0031] The second negative branch includes a 1×1 convolution, a 1×3 convolution, a 3×1 convolution, and a 3×3 convolution connected in sequence;
[0032] The third negative branch includes a 1×1 convolution, a 1×5 convolution, a 5×1 convolution, and a 5×5 convolution connected in sequence;
[0033] The fourth negative branch includes a 1×1 convolution, a 1×7 convolution, a 7×1 convolution, and a 7×7 convolution connected in sequence.
[0034] Furthermore: in S4, the edge-guided fusion module performs top-down fusion operations of the first to third levels, and the top-down fusion operations of each level obtain corresponding fusion features and camouflaged target prediction results;
[0035] The S4 comprises the following sub-steps:
[0036] S41, inputting the edge features and global prior information into an edge-guided fusion module, inputting the refined fourth feature and the refined fifth feature as low-resolution features and high-resolution features, respectively, into the edge-guided fusion module, performing a first-level top-down fusion operation, obtaining a first-level fused feature, and using it as a first camouflaged target prediction;
[0037] S42, inputting the edge features and the global prior information into the edge-guided fusion module, inputting the refined third features and the first-level fused features into the edge-guided fusion module as low-resolution features and high-resolution features, respectively, performing a second-level top-down fusion operation, obtaining a second-level fused feature, and using it as the second camouflaged target prediction;
[0038] S43, inputting the edge features and the global prior information into the edge-guided fusion module, inputting the refined second features and the second-level fused features into the edge-guided fusion module as low-resolution features and high-resolution features, respectively, performing a third-level top-down fusion operation, obtaining a third-level fused feature, and using it as the third camouflaged target prediction;
[0039] S44: Taking the third disguised target prediction as the final disguised target prediction result, completing the disguised target segmentation.
[0040] The aforementioned further solution has the following beneficial effects: The edge-guided fusion module is used to integrate the extracted edge features into the network to guide learning and fuse features at different levels. This module can fuse features at different levels, leveraging multi-scale information to mitigate the scale variations introduced by fusing different prior information. This smooths the aggregated features and produces the most comprehensive feature representation, thereby enhancing the ability to segment camouflaged targets.
[0041] Furthermore, the method for performing the top-down fusion operation of the first to third levels is the same, which is specifically as follows:
[0042] SA1, calculate edge-enhanced features based on edge features, low-resolution features, and high-resolution features;
[0043] SA2, calculate the fusion features based on the edge enhanced features, global prior information, low-resolution features and high-resolution features, and complete the top-down fusion operations of the first to third levels.
[0044] Further: In SA1, the edge-enhanced feature F′ is calculated e The specific expression is:
[0045]
[0046] Where, F l is a low-resolution feature, F h is a high-resolution feature, UP(·) is an upsampling operation, ⊕ is an addition operation, and F e is the edge feature;
[0047] In SA2, the fusion feature F is calculated b The specific expression is:
[0048]
[0049] Where, F p is the global prior information, For the subtraction operation, is a multiplication operation, Conv 3×3 (·) is a 3×3 convolution operation, and M(·) is a multi-scale channel attention module.
[0050] Further: In S4, the edge-guided fusion module performs a top-down fusion operation supervised by the truth map, and the loss function L of the supervision of the truth map is total The specific expression is:
[0051]
[0052] Where, is the weighted binary cross entropy loss, is the weighted intersection-over-union loss, L dice (·) is the dice loss, P2 is the first disguised target prediction, P3 is the second disguised target prediction, P4 is the third disguised target prediction, P e is the edge prediction, G is the true value map, G e is the edge truth map, is a trade-off parameter.
[0053] The beneficial effects of the present invention are as follows: the present invention provides a multi-level edge-enhanced fusion camouflaged target segmentation method, which uses edge features to guide learning, significantly improving the performance of camouflaged target detection. The residual texture enhancement module is used to refine the backbone features, the edge extraction module is used to obtain edge features, and finally the edge-guided fusion module is used to fuse the edge features with global prior information, making the final result with finer edges and more complete structures, and can completely segment the camouflaged target and avoid edge blur. Extensive experiments conducted on three challenging benchmark datasets show that under four widely used evaluation indicators, the camouflaged target segmentation method of the present invention outperforms the prediction results of multiple cutting-edge models in recent years, achieving advanced results. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a flow chart of a camouflaged target segmentation method with multi-level edge enhancement and fusion according to the present invention.
[0055] Figure 2 It is a schematic diagram of the overall structure of the network model of the present invention.
[0056] Figure 3 Schematic diagram of the edge extraction module structure of the present invention.
[0057] Figure 4 Schematic diagram of the residual texture enhancement module structure of the present invention.
[0058] Figure 5 This is a structural diagram of the boundary-guided fusion module of the present invention.
[0059] Figure 6 Schematic diagram of performance comparison experimental results of the present invention. DETAILED DESCRIPTION
[0060] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0061] Example 1:
[0062] like Figure 1 As shown, in one embodiment of the present invention, a multi-level edge enhancement fusion camouflaged target segmentation method includes the following steps:
[0063] S1. Extracting multi-level features of an input image, wherein the multi-level features include second to fifth features, and inputting the second to fifth features into an edge extraction module to obtain edge features;
[0064] S2. Input the second to fifth features into the residual texture enhancement module to obtain refined second to fifth features;
[0065] S3. Input the fifth feature into the pyramid pooling module to obtain global prior information;
[0066] S4. Input the edge features, the refined second to fifth features, and the global prior information into the edge-guided fusion module, and perform a top-down fusion operation through the edge-guided fusion module to obtain the final camouflaged target prediction result and complete the camouflaged target segmentation.
[0067] The present invention optimizes edge segmentation by directly using the edge information of the camouflaged target, thereby enhancing the performance of camouflaged target segmentation. The COD results are improved mainly by using the residual texture enhancement module, the edge extraction module and the edge-guided fusion module. The residual texture enhancement module can increase the receptive field to capture richer feature information in the backbone network layer and provide effective prior information for subsequent fusion. The edge extraction module integrates features at different levels and explores and captures effective edge features under the supervision of the edge truth map through the channel attention mechanism, and transmits them to each layer of the network for enhancement. The edge-guided fusion module can combine edge features with features at different levels to guide learning, so as to enhance the amount of edge information included and fuse features of different scales from high to low layers.
[0068] In S1, the multi-level features of the input image are extracted through the Res2Net-50 network, and the number of channels of the second to fifth features are 256, 512, 1024 and 2048 respectively.
[0069] like Figure 2 As shown, in this embodiment, Res2Net-50 is used as the backbone network to extract the input image F∈R H ×W×3 Multi-level features are extracted from the image processing unit, and the multi-level features are specifically the first feature F1, the second feature F2, the third feature F3, the fourth feature F4 and the fifth feature F5.
[0070] like Figure 3 As shown, in S1, the edge feature F is obtained e The specific expression is:
[0071]
[0072] Where, F o is the integrated feature, which is obtained by integrating the second to fifth features. 1×1 is a 1×1 convolution operation, is a multiplication operation, σ is a Sigmoid function, GAP(·) is a global average pooling operation, Conv1d(·) is a one-dimensional convolution operation with a convolution kernel of k, and k is the size of the adaptive convolution kernel. The specific expression is as follows:
[0073]
[0074] Where C is the set number of channels, |·| odd To take the nearest odd general formula.
[0075] In this embodiment, the number of channels is 256, so the convolution kernel k=5 is obtained. The method for calculating edge features of the present invention takes into account local channel attention and highlights key channels, thereby removing redundant channels and noise, having lower complexity, and playing a key role in exploring effective edge semantics in noisy features.
[0076] The integration feature F o The specific expression is:
[0077] F o =Conv 3×3 (CAT(F′2,F′3,F′4,F′5))
[0078] F′2=Conv 1×1 (F2)
[0079] F′3=UP(Conv 1×1 (F3)
[0080] F′4=UP(Conv 1×1 (F4)
[0081] F′5=UP(Conv 1×1 (F5)
[0082] In the formula, CAT(·) is the concatenation operation, Conv 3×3 is a 3×3 convolution operation, UP(·) is an upsampling operation, F2 is the second feature, F3 is the third feature, F4 is the fourth feature, F5 is the fifth feature, F′2 is the second feature with 64 channels, F′3 is the third feature with 64 channels, F′4 is the fourth feature with 64 channels, and F′5 is the fifth feature with 64 channels.
[0083] In this embodiment, in order to better mine edge semantic information, the present invention combines low-level features with high-level features and designs an edge extraction module. The principle of obtaining edge features is specifically as follows:
[0084] The features F2-F5 are reduced to 64 channels through a 1×1 convolution, and then F3-F5 are upsampled to the same size as F2 for channel splicing, and then a 3×3 convolution is performed to obtain the integrated feature F o ∈R W×H×C , the number of channels of integrated features is 256.
[0085] like Figure 4 As shown, in S2, the residual texture enhancement module includes a main branch and first to fourth negative branches, wherein the inputs of the main branch and the first negative branch are input features, and the inputs of the second to fourth negative branches are the sum of the up-sampled output of the previous negative branch and the input features;
[0086] Wherein, the input feature is the second feature, the third feature, the fourth feature or the fifth feature;
[0087] The method for obtaining the refined second to fifth features is specifically a method for obtaining the refined input features, which is specifically:
[0088] After splicing the outputs of the first to fourth negative branches, the splicing result is convolved by 3×3 and added to the output of the group branch. The addition result is input into the ReLU function to obtain refined input features.
[0089] In this embodiment, since the backbone network uses a large number of convolution operations, it is unable to extract features containing rich contextual information, which is not conducive to the segmentation of camouflaged targets. Therefore, when using the features extracted by the backbone network, it is necessary to refine the features first. Therefore, the present invention designs a residual texture enhancement module based on TEM to refine the features of the backbone network to provide more effective prior information in subsequent fusion.
[0090] The expansion rates of the first to fourth negative branches are 1, 3, 5 and 7 respectively;
[0091] The main branch and the first negative branch are both provided with a 1×1 convolution;
[0092] The second negative branch includes a 1×1 convolution, a 1×3 convolution, a 3×1 convolution, and a 3×3 convolution connected in sequence;
[0093] The third negative branch includes a 1×1 convolution, a 1×5 convolution, a 5×1 convolution, and a 5×5 convolution connected in sequence;
[0094] The fourth negative branch includes a 1×1 convolution, a 1×7 convolution, a 7×1 convolution, and a 7×7 convolution connected in sequence.
[0095] like Figure 4As shown in the figure, the outputs of all sub-branches are concatenated and then a 3×3 convolution is performed and the main branch features are added. Then the entire module is sent to a ReLU function to obtain refined input features. Compared with TEM, the residual texture enhancement module adds three residual branch structures, which can better fuse information of different scales to enhance feature representation.
[0096] In S4, the edge-guided fusion module performs top-down fusion operations on the first to third levels, and the top-down fusion operations on each level obtain corresponding fusion features and camouflaged target prediction results;
[0097] In this embodiment, the edge-guided fusion module is used to integrate the extracted edge features into the network to guide learning and fuse features at different levels. This module can fuse features at different levels, leveraging multi-scale information to mitigate the scale variations introduced by fusing different prior information. This smooths the aggregated features and produces the most comprehensive feature representation, thereby enhancing the ability to segment camouflaged objects.
[0098] The S4 comprises the following sub-steps:
[0099] S41, inputting the edge features and global prior information into an edge-guided fusion module, inputting the refined fourth feature and the refined fifth feature as low-resolution features and high-resolution features, respectively, into the edge-guided fusion module, performing a first-level top-down fusion operation, obtaining a first-level fused feature, and using it as a first camouflaged target prediction;
[0100] S42, inputting the edge features and the global prior information into the edge-guided fusion module, inputting the refined third features and the first-level fused features into the edge-guided fusion module as low-resolution features and high-resolution features, respectively, performing a second-level top-down fusion operation, obtaining a second-level fused feature, and using it as the second camouflaged target prediction;
[0101] S43, inputting the edge features and the global prior information into the edge-guided fusion module, inputting the refined second features and the second-level fused features into the edge-guided fusion module as low-resolution features and high-resolution features, respectively, performing a third-level top-down fusion operation, obtaining a third-level fused feature, and using it as the third camouflaged target prediction;
[0102] S44: Taking the third disguised target prediction as the final disguised target prediction result, completing the disguised target segmentation.
[0103] like Figure 5 As shown, the method for performing the top-down fusion operation of the first to third levels is the same, which is specifically as follows:
[0104] SA1, calculate edge-enhanced features based on edge features, low-resolution features, and high-resolution features;
[0105] SA2, calculate the fusion features based on the edge enhanced features, global prior information, low-resolution features and high-resolution features, and complete the top-down fusion operations of the first to third levels.
[0106] In SA1, the edge-enhanced feature F′ is calculated. e The specific expression is:
[0107]
[0108] Where, F l is a low-resolution feature, F h is a high-resolution feature, UP(·) is an upsampling operation, ⊕ is an addition operation, and F e is the edge feature;
[0109] In this embodiment, the edge-guided fusion module considers the fusion of cross-layer features, edge features, and global features, and multiplies the edge features F e Injected into the features of each layer to obtain the edge-enhanced features F′ e .
[0110] In SA2, the fusion feature F is calculated b The specific expression is:
[0111]
[0112] Where, F p is the global prior information, For the subtraction operation, is a multiplication operation, Conv 3×3 (·) is a 3×3 convolution operation, and M(·) is a multi-scale channel attention module.
[0113] The present invention also uses a multi-scale channel attention module to aggregate the fused features and mine contextual information to enhance the detection effect. The multi-scale channel attention module is based on a dual-branch structure, in which one branch uses global average pooling to obtain global context to emphasize large objects with global distribution, while the other branch maintains the original feature size to obtain local context to avoid ignoring small objects, and finally aggregates multi-scale contextual information. The weights obtained by the multi-scale channel attention module are multiplied with the low-resolution features, and the attention background is multiplied with the high-resolution features. After adding and fusing them, they pass through a 3×3 convolution, and then the global context information F is added. p Then the feature F is output after a 3×3 convolution b .
[0114] In S4, the edge-guided fusion module performs a top-down fusion operation supervised by the truth map, and the loss function L of the supervision of the truth map is total The specific expression is:
[0115]
[0116] Where, is the weighted binary cross entropy loss, is the weighted intersection-over-union loss, L dice (·) is the dice loss, P2 is the first disguised target prediction, P3 is the second disguised target prediction, P4 is the third disguised target prediction, P e is the edge prediction, G is the true value map, G e is the edge truth map, is a trade-off parameter.
[0117] In this embodiment, the weight parameter is set to 3. The present invention has four output predictions, which are the first to third disguised target predictions and edge predictions. The edge predictions are obtained based on the edge features of the edge extraction module. The first to third disguised target predictions come from the edge-guided fusion module. The true value map and the edge true value map are both input from the outside.
[0118] The implementation process of the method of the present invention is specifically as follows:
[0119] Under the supervision of the ground-truth edge map, the edge extraction module extracts edge semantic information related to the camouflaged target from low-level features containing edge semantic information and high-level features containing global structural information. Simultaneously, the second to fifth features are input into the residual texture enhancement module. This module expands the receptive field and performs residual fusion, further refining the features from the backbone network, removing excess noise from the features, and reducing the number of channels to 64. The pyramid pooling module expands the network's receptive field through multi-branch global adaptive pooling, mining global contextual information. The fifth feature is input into the pyramid pooling module to obtain effective global prior information and reduce the number of channels to 64. The edge-guided fusion module fuses edge features, refined second to fifth features, and global prior information in a top-down manner at each level. Each output of the edge-guided fusion module is supervised by the ground-truth map. The final camouflaged target prediction result is used as the network's prediction result to complete camouflaged target segmentation.
[0120] Example 2:
[0121] This example is a performance comparison experiment set up for Example 1:
[0122] The algorithm of the present invention is evaluated on three public benchmark datasets, which are also the three most widely used datasets in the field of COD: (1) CAMO, containing 1250 images (1000 training sets and 250 test sets), (2) COD10K, containing 5066 disguised images (3040 training sets and 2026 test sets); and (3) NC4K, containing 4121 images, which is the largest test set. The present invention is consistent with the settings of most COD models, using the training sets of CAMO and COD10K as the training sets of the model, and using the test sets of CAMO, COD10K and NC4K as the test sets. In order to conduct a better comprehensive comparison, the four most widely used indicators in COD are used to evaluate the method proposed in the present invention: E-measure, S-measure, Weighted F-measure, and MAE.
[0123] To demonstrate the effectiveness of our proposed algorithm, we compare it with 15 cutting-edge COD models from recent years, including PoolNet, EGNet, UCNet, PraNet, SINet, C2FNet, PFNet, UGTR, LSR, ERRNet, PreyNet, BSANet, TPRNet, FAPNet, and SINetV2. For a fair comparison, all evaluation data for these methods was obtained from individual papers or retrained using open-source code.
[0124] As shown in Table 1, the method of the present invention is quantitatively compared with 15 advanced models in four evaluation indicators on three benchmark datasets. Obviously, the method of the present invention outperforms the other 15 models in all four indicators on the three datasets. Among them, the methods that also use edge information to improve COD capabilities (ERRNet, BSANet, FAPNet) are compared with the method of the present invention. The method of the present invention significantly improves ERRNet by 9% on COD10K-test. On CAMO-test, it significantly improves the performance by 3.6% compared to BSANet. Improved by 2.6% over FAPNet This proves that the method of the present invention is more effective in extracting and using edge information. Compared with the second best SINetV2, it is better overall. Although there are two indicators that are lower than SINetV2, they are very close, and the visualization results are also more detailed than SINetV2.
[0125] Table 1 provides a quantitative comparison of four evaluation metrics with 15 state-of-the-art models on three benchmark datasets.
[0126]
[0127] like Figure 6 The figure shows a qualitative comparison of the prediction results of seven methods for several typical samples in the dataset. As can be seen from the figure, the method of the present invention provides accurate camouflaged target segmentation results with a more refined and complete object structure. In particular, for some images with blurred edges, the present invention can also finely segment the edges. For example, in the third and fourth rows of the figure, the edges of the camouflaged target are particularly blurred. In this case, it can be seen that the other models are inaccurate in segmentation, while the model of the present invention can still accurately detect the camouflaged target with rich edge details, demonstrating that edge information enhancement improves the detection ability of camouflaged targets.
[0128] Example 3:
[0129] This embodiment is an ablation experiment set up for Example 1:
[0130] To verify the effectiveness of the designed modules, we conducted five ablation experiments, as shown in Table 2. The baseline model removed all of our designed modules, retaining only the texture enhancement module and the pyramid pooling module. Feature fusion was performed using simple additive fusion. In the second experiment, we added an edge extraction module to the baseline model and fused it with features from each layer in a multiplicative manner. The addition of the edge extraction module improved all performance aspects, demonstrating the effectiveness of the edge extraction module. In the third experiment, we added an edge-guided fusion module to the baseline model. This change in fusion method resulted in a gradual improvement in all performance aspects, demonstrating the effectiveness of the edge-guided fusion module. In the fourth experiment, we added both the edge extraction module and the edge-guided fusion module to the baseline model. Overall performance was superior to that of either module alone. Finally, we replaced the texture enhancement module in the fourth experiment with the residual texture enhancement module designed in this paper. This achieved the best performance, with steady improvements across all performance aspects, demonstrating the effectiveness of the improvements to the texture enhancement module. It should be noted that all of the above experiments included the pyramid pooling module.
[0131] Table 2 Ablation experiment results on three benchmark datasets.
[0132]
[0133] To verify the effectiveness of the pyramid pooling module placed at the top layer of the network, we also designed an ablation experiment. As shown in Table 3, the first row shows the experimental results after removing the pyramid pooling module from the complete model. Clearly, the model with the pyramid pooling module achieves higher performance across all metrics, demonstrating the effectiveness of the global contextual information captured by the pyramid pooling module in improving camouflaged object detection performance.
[0134]
[0135] Table 3 Pyramid Pooling Module (PPM) ablation experiment results.
[0136] The beneficial effects of the present invention are as follows: the present invention provides a multi-level edge-enhanced fusion camouflaged target segmentation method, which utilizes edge features to guide learning, significantly improving the performance of camouflaged target detection. The residual texture enhancement module refines backbone features, the edge extraction module acquires edge features, and finally, the edge-guided fusion module fuses edge features with global prior information, resulting in a final result with finer edges and a more complete structure. Extensive experiments conducted on three challenging benchmark datasets demonstrate that, under four widely used evaluation metrics, the camouflaged target segmentation method of the present invention outperforms the prediction results of multiple cutting-edge models in recent years, achieving advanced results.
[0137] In the description of the present invention, it should be understood that the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying the relative importance or the number of technical features implicitly specified. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of such features.
Claims
1. A multi-level edge enhancement fusion camouflage target segmentation method, characterized in that: The following steps are involved: S1. Extracting multi-level features of an input image, wherein the multi-level features include second to fifth features, and inputting the second to fifth features into an edge extraction module to obtain edge features; The multi-level features of the input image are extracted through the Res2Net-50 network, where the number of channels of the second to fifth features are 256, 512, 1024, and 2048 respectively; Edge features The specific expression is: Where, is an integrated feature, which is obtained by integrating the second to fifth features. is a 1×1 convolution operation, is the multiplication operation, is the Sigmoid function, GAP (·) is the global average pooling operation, The convolution kernel is k One-dimensional convolution operation, k is the adaptive convolution kernel size, and its specific expression is as follows: Where, C is the set number of channels, To take the nearest odd general formula; S2. Input the second to fifth features into the residual texture enhancement module to obtain refined second to fifth features; The residual texture enhancement module includes a main branch and first to fourth negative branches, wherein the inputs of the main branch and the first negative branch are input features, and the inputs of the second to fourth negative branches are the sum of the up-sampled output of the previous negative branch and the input features; Wherein, the input feature is the second feature, the third feature, the fourth feature or the fifth feature; The method for obtaining the refined second to fifth features is specifically a method for obtaining refined input features, which is specifically: After splicing the outputs of the first to fourth negative branches, perform a 3×3 convolution on the spliced result and add it to the output of the group branch. The added result is input into the ReLU function to obtain refined input features; S3. Input the fifth feature into the pyramid pooling module to obtain global prior information; S4. Input the edge features, refined second to fifth features, and global prior information into the edge-guided fusion module, perform a top-down fusion operation through the edge-guided fusion module, obtain the final camouflaged target prediction result, and complete the camouflaged target segmentation.
2. The camouflaged target segmentation method with multi-level edge enhancement fusion according to claim 1, characterized in that: The integrated features The specific expression is: Where, For splicing operation, is a 3×3 convolution operation, is the upsampling operation, For the second feature, For the third feature, The fourth characteristic is The fifth characteristic is is the second feature with 64 channels, is the third feature with 64 channels, is the fourth feature with 64 channels, It is the fifth feature with 64 channels.
3. The camouflaged target segmentation method with multi-level edge enhancement fusion according to claim 1, characterized in that: The expansion rates of the first to fourth negative branches are 1, 3, 5, and 7, respectively; The main branch and the first negative branch are both provided with a 1×1 convolution; The second negative branch includes a 1×1 convolution, a 1×3 convolution, a 3×1 convolution, and a 3×3 convolution connected in sequence; The third negative branch includes a 1×1 convolution, a 1×5 convolution, a 5×1 convolution, and a 5×5 convolution connected in sequence; The fourth negative branch includes a 1×1 convolution, a 1×7 convolution, a 7×1 convolution, and a 7×7 convolution connected in sequence.
4. The camouflaged target segmentation method with multi-level edge enhancement fusion according to claim 1, characterized in that: In S4, the edge-guided fusion module performs top-down fusion operations on the first to third levels, and the top-down fusion operations on each level obtain corresponding fusion features and camouflaged target prediction results; The S4 comprises the following sub-steps: S41, inputting the edge features and global prior information into an edge-guided fusion module, inputting the refined fourth feature and the refined fifth feature as low-resolution features and high-resolution features, respectively, into the edge-guided fusion module, performing a first-level top-down fusion operation, obtaining a first-level fused feature, and using it as a first camouflaged target prediction; S42, inputting the edge features and the global prior information into the edge-guided fusion module, inputting the refined third features and the first-level fused features into the edge-guided fusion module as low-resolution features and high-resolution features, respectively, performing a second-level top-down fusion operation, obtaining a second-level fused feature, and using it as the second camouflaged target prediction; S43, inputting the edge features and the global prior information into the edge-guided fusion module, inputting the refined second features and the second-level fused features into the edge-guided fusion module as low-resolution features and high-resolution features, respectively, performing a third-level top-down fusion operation, obtaining a third-level fused feature, and using it as the third camouflaged target prediction; S44: Taking the third disguised target prediction as the final disguised target prediction result, completing the disguised target segmentation.
5. The camouflaged target segmentation method with multi-level edge enhancement fusion according to claim 4, characterized in that: The method for performing the top-down fusion operation of the first to third levels is the same, which is specifically as follows: SA1, calculate edge-enhanced features based on edge features, low-resolution features, and high-resolution features; SA2, calculate the fusion features based on the edge enhanced features, global prior information, low-resolution features and high-resolution features, and complete the top-down fusion operations of the first to third levels.
6. The camouflaged target segmentation method with multi-level edge enhancement fusion according to claim 5, characterized in that: In SA1, the edge-enhanced features are calculated. The specific expression is: Where, is a low-resolution feature, For high-resolution features, is the upsampling operation, For the addition operation, is the edge feature; In SA2, the fusion feature is calculated The specific expression is: Where, is the global prior information, For the subtraction operation, is the multiplication operation, is a 3×3 convolution operation, and M(·) is a multi-scale channel attention module.
7. The camouflaged target segmentation method with multi-level edge enhancement fusion according to claim 5, characterized in that: In S4, the edge-guided fusion module performs a top-down fusion operation supervised by the truth map, and the loss function of the supervision of the truth map is The specific expression is: Where, is the weighted binary cross entropy loss, is the weighted intersection-over-union loss, is the dice loss, For the first camouflage target prediction, For the second camouflage target prediction, For the third camouflage target prediction, For edge prediction, G is the true value graph, is the edge truth map, is a trade-off parameter.
Citation Information
Patent Citations
Camouflage target detection method based on attention mechanism and convolutional neural network
CN116228702A
Medical image segmentation using an integrated edge guidance module and object segmentation network
US10482603B1