Camouflage target detection method based on boundary feature enhancement

By explicitly enhancing boundary features and guiding attention through a three-stage network architecture, combined with multi-scale feature fusion and loss function optimization, the problems of blurred boundaries and low feature discrimination in camouflaged target detection are solved, achieving higher accuracy and robustness in camouflaged target detection.

CN120876885APending Publication Date: 2025-10-31NANJING YIQICHUANG INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510922225.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing deep learning methods suffer from poor detection accuracy and boundary segmentation in camouflaged target detection due to blurred boundaries and low feature discrimination, and cannot effectively handle the low contrast difference between camouflaged targets and the background.

Method used

A three-stage network architecture is adopted, including a boundary enhancement module (BEM) to explicitly predict the target contour, a boundary guidance module (BGM) to dynamically guide attention, and a multi-level feature aggregation module (FAM) to fuse multi-scale information. The model is optimized by combining weighted binary cross-entropy and weighted cross-union ratio loss functions.

Benefits of technology

It significantly improves the detection accuracy and boundary segmentation quality of camouflaged targets, enabling more accurate identification and segmentation of camouflaged targets, adapting to camouflaged targets of different scales, and exhibiting stronger robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876885A_ABST
    Figure CN120876885A_ABST
Patent Text Reader

Abstract

The invention discloses a camouflage target detection method based on boundary feature enhancement. The method comprises the steps of firstly extracting multi-level features of an image by using a backbone network; then, through a boundary enhancement module, low-level fine texture features and high-level semantic features are fused, and the boundary enhancement module is specially used for generating a high-quality boundary prediction map; then, a boundary guiding module takes the boundary diagram as attention guidance, so that the network pays more attention to a boundary region which is crucial for distinguishing a target from a background in subsequent processing; and finally, effectively fusing the guided hierarchical features through a multi-level feature aggregation module so as to adapt to camouflage targets of different scales. According to the method, by explicitly enhancing and utilizing the boundary features, the detection precision of the camouflage target and the integrity of boundary segmentation are remarkably improved, excellent performance is shown on a plurality of public data sets, and the method has important academic research value and wide application prospects and can be applied to the fields of ecological monitoring, security reconnaissance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and deep learning technology, and in particular to a method and system for object detection and segmentation in images.

[0002] More specifically, this invention provides an innovative technical solution that can significantly improve detection accuracy and boundary segmentation quality for a highly challenging task—Camouflaged Object Detection (COD). Background Technology

[0003] In nature, many organisms have evolved camouflage abilities that blend seamlessly with their surroundings in terms of color, texture, and shape to avoid predators or hunt. This phenomenon results in extremely low visual differences between camouflaged targets and their backgrounds, posing a significant challenge to both human observers and computer vision systems. Camouflaged target detection (COD) aims to accurately identify and segment these "invisible" objects from complex natural scene images. It has significant theoretical research value and broad practical application prospects in fields such as ecological species monitoring, biodiversity conservation, military reconnaissance, and intelligent security.

[0004] Traditional image processing methods, such as those based on color, texture, or saliency models, perform reasonably well when dealing with ordinary targets of high contrast, but struggle with camouflaged targets. These methods rely on hand-designed low-level features, while camouflaged targets inherently weaken the discriminative power of these features, resulting in low detection rates and inaccurate boundary localization using traditional methods.

[0005] With the rise of deep learning, methods represented by convolutional neural networks (CNNs) have achieved breakthroughs in various visual tasks. However, directly applying CNN models designed for general object detection or segmentation to COD tasks encounters several core technical bottlenecks:

[0006] Boundary information loss problem: Modern CNN architectures (such as ResNet, VGG, etc.) commonly use pooling layers or strided convolutions for spatial downsampling to build deep networks, expand the receptive field, and reduce computational costs. This operation inevitably reduces the spatial resolution of feature maps during layer-by-layer propagation, leading to the blurring or complete loss of fine boundary details. For camouflaged targets, whose boundaries are already blurred, any loss of information will severely affect the accuracy and completeness of the final segmentation.

[0007] Limitations of general attention mechanisms: While attention mechanisms (such as SE-Net and CBAM) allow networks to focus on more important features, most mechanisms compute channel weights based on global context (such as global average pooling). This approach emphasizes global statistical features and is insufficient in responding to subtle boundary differences that are crucial for distinguishing camouflaged targets from the background and often exist in local regions. When dealing with low-contrast camouflage edges, this global attention is prone to weight diffusion, failing to achieve precise focusing.

[0008] Challenges of multi-scale feature fusion: camouflaged targets may appear at arbitrary scales in images. Although structures such as Feature Pyramid Networks (FPN) attempt to fuse multi-scale features, semantic gaps still exist between different levels during feature propagation, which may lead to semantic information fragmentation or insufficient feature fusion, resulting in poor prediction results for targets at specific scales or sensitivity to changes in target scale.

[0009] In the prior art, Chinese patent applications such as "A method for detecting camouflaged targets in hyperspectral images based on deep learning" (application number 202010135340.0), "A method for detecting camouflaged targets based on improved YOLO algorithm" (application number 202110097503.5), and "A method for detecting and recognizing camouflaged targets based on deep neural networks" (application number 202110812766.X) all discuss the use of deep learning for camouflaged target detection.

[0010] In summary, existing technical solutions have failed to effectively address the core problems in camouflaged target detection caused by blurred boundaries and low feature discrimination. Therefore, there is an urgent need for a novel deep learning method that can explicitly model, enhance, and utilize boundary information, while possessing strong cross-scale feature fusion capabilities, to overcome the performance bottlenecks of current camouflaged target detection technologies. Summary of the Invention

[0011] The core objective of this invention is to address the challenges presented in the aforementioned background technology by providing an innovative camouflage target detection method based on boundary feature enhancement. This method, through a meticulously designed three-stage network architecture, can significantly improve the detection accuracy of camouflage targets, particularly the completeness and accuracy of boundary segmentation.

[0012] To achieve the above objectives, the technical solution proposed by this invention includes the following core steps:

[0013] First, this invention explicitly transforms the fuzzy boundary detection problem into a learnable subtask.

[0014] A boundary enhancement module (BEM) is designed to fuse low-level features rich in spatial detail and high-level features rich in global semantics. Its output is a high-quality initial boundary prediction map that clearly depicts the outline of the camouflaged target, providing strong prior information for subsequent processing.

[0015] Secondly, this invention proposes a mechanism for attention guidance using prior knowledge. By designing a boundary guidance module (BGM), the boundary prediction map generated by BEM is used as a dynamic guidance signal. Based on this, the BGM generates channel attention weights and adaptively adjusts the channel responses of subsequent feature maps, enabling the network to "focus" computational resources on the boundary regions that are crucial for distinguishing the target from the background. This effectively suppresses the interference of background texture and improves the discriminative power of the features.

[0016] Furthermore, to address the diverse scales of camouflaged targets, this invention designs a multi-level feature aggregation module (FAM). This module employs multiple sets of parallel dilated convolutional branches with different dilation rates, enabling efficient capture of multi-scale contextual information without sacrificing spatial resolution. By effectively fusing this information, the model ensures robust detection capabilities for camouflaged targets of varying sizes.

[0017] Finally, this invention employs a carefully designed joint loss function during the model training phase. This function consists of weighted binary cross-entropy (WBCE) loss and weighted intersection-union ratio (WIoU) loss, which synergistically optimizes both pixel classification and region shape, further enhancing the overall performance of the model.

[0018] Based on the above ideas, this invention proposes a camouflage target detection method based on boundary feature enhancement. It aims to solve the technical problem in the field of computer vision where the camouflage target and the surrounding environment are highly integrated in terms of color, texture and shape, resulting in extremely low contrast between the target and the background and blurred boundary contours, making it difficult to accurately segment by existing technologies. The root cause of this problem is that existing deep learning models inherently lose key boundary detail information during the feature extraction downsampling process, and the general attention mechanism is insufficient in responding to subtle differences between the camouflage target and the background.

[0019] This invention employs an innovative, boundary-centric, three-stage processing architecture to explicitly model and enhance the boundary features of camouflaged targets, dynamically guide and focus them, and robustly aggregate them across multiple scales. The method sequentially performs the following processing steps:

[0020] Step 1: Hierarchical Feature Extraction and Representation. The input image to be detected is acquired, and a pre-trained deep convolutional backbone network is used to perform forward propagation processing on the input image. This process aims to transform the raw pixel information into a hierarchical, multi-dimensional feature representation, generating a set of feature maps with different spatial resolutions and semantic depths. This set of feature maps forms a rich information foundation for subsequent boundary awareness and object segmentation. Low-level feature maps preserve high-resolution spatial details, while high-level feature maps contain abstract global contextual semantics.

[0021] Step 2: Explicit Enhancement and Prediction of Boundary Features. To fundamentally address the problem of lost boundary information, at least one low-level feature map and one high-level feature map from the multiple layers of feature maps are input into a specially designed boundary enhancement module (BEM). This module generates a high-quality initial boundary prediction map that highlights the precise contours of the camouflaged target by fusing the fine spatial details carried by the low-level features and the global semantic guidance provided by the high-level features in parallel, using an explicit modeling approach. This prediction map will then serve as strong prior knowledge for attention guidance in subsequent steps.

[0022] Step 3: Dynamic Feature Guidance Based on Boundary Prior. To enable the network to intelligently focus on boundary regions crucial for segmentation, the initial boundary prediction map and at least one intermediate-level feature map from the multiple layers of feature maps are input into a boundary guidance module (BGM). This module uses the boundary prediction map as a dynamic, spatially variable guidance signal, and calculates and generates a channel attention weight vector accordingly. Subsequently, this weight vector is applied to the intermediate-level feature map to adaptively and selectively enhance the response of feature channels highly correlated with the target boundary, while suppressing irrelevant or interfering background features, thereby obtaining a more discriminative feature map guided and optimized by boundary information.

[0023] Step 4: Multi-scale aggregation of guided features. To ensure robust detection capabilities against camouflaged targets of different sizes and shapes, the guided feature maps and feature maps from other levels are input into a multi-level feature aggregation module (FAM). This module efficiently captures and integrates contextual information from different receptive fields without sacrificing spatial resolution through multiple parallel convolutional branches with different dilation rates. A bottom-up aggregation path effectively fuses the guided and optimized multi-level features, ultimately generating a final feature representation that retains fine boundary details and contains rich multi-scale context.

[0024] Step 5: Generation and output of the final target mask. The final feature representation is fed into a classification head network to perform binary classification (target or background) on each pixel, thereby generating a final detection result mask of the same size as the input image, representing the precise location and complete shape of the camouflaged target, thus completing the entire detection process.

[0025] Specifically:

[0026] In step 2, the boundary enhancement module (BEM) used to generate the initial boundary prediction map has the following internal structure in order to effectively integrate low-level spatial details and high-level semantic information: a multi-scale context information extraction branch, which receives the low-level feature map and applies max pooling and average pooling operations to it in parallel to capture significant local features and regional average responses, respectively, and then integrates the two results to obtain richer context information.

[0027] A semantic information delivery branch receives the high-level feature map and restores its semantic information to the same spatial resolution as the low-level feature map through upsampling techniques such as bilinear interpolation to achieve feature map alignment; a fusion output unit concatenates and fuses the outputs of the multi-scale context information extraction branch and the semantic information delivery branch in the channel dimension, and performs deep information interaction through at least one convolutional layer. Finally, the output is normalized to a probability value between 0 and 1 through the Sigmoid activation function to form the initial boundary prediction map.

[0028] In step 3, the boundary guidance module (BGM) used to generate the guided feature map, in order to dynamically and adaptively enhance the response of key feature channels by utilizing boundary prior information, specifically includes the following processing steps:

[0029] The initial boundary prediction map is downsampled to match its spatial dimensions with the intermediate-level feature map. The downsampled initial boundary prediction map and the intermediate-level feature map are concatenated along the channel dimension to form a combined feature map containing the original features and boundary priors. Global average pooling is applied to this combined feature map to compress its spatial dimensions, obtaining a channel descriptor vector representing global information. This vector is then passed through at least one fully connected layer and a sigmoid activation function to learn and generate a channel attention weight vector. The channel attention weight vector is element-wise multiplied with the original intermediate-level feature map to dynamically weight different channels. The weighted result is then added to the original feature map via a residual connection to ensure information flow stability and prevent gradient vanishing. Finally, the guided feature map is output.

[0030] In step 4, the multi-level feature aggregation module (FAM) used to generate the final feature representation has the following internal structure to effectively capture and fuse target information at different scales without reducing spatial resolution: multiple parallel dilated convolution branches, each using a dilated convolution kernel with a different dilation rate, wherein the dilation rate is set to an incremental value to construct multiple effective receptive fields of different sizes, thereby enabling simultaneous perception of local details and wide-area context of the target; the dilated convolution branches use asymmetric convolution kernels, such as a combination of 1xN and Nx1, to more efficiently capture features in the horizontal and vertical directions; the outputs of all parallel branches are fused and combined with the input features of the module through residual connections to generate aggregated features that are highly adaptable to changes in target scale.

[0031] In order to effectively evaluate the performance and generalization ability of the constructed neural network model during the training process, this invention divides the overall dataset into three parts: training set, validation set, and test set.

[0032] During the model training phase, a feature dataset highly relevant to the target task is first selected from the full dataset. This feature set is then used to train the deep neural network model. The training process iteratively adjusts the model weights and parameters, enabling the network to learn the feature distribution and hidden patterns in the dataset, thereby achieving accurate predictions of unseen data.

[0033] After each training round, a validation set, independent of the training set, is used to validate the model's performance. The validation set is used only to evaluate model performance and optimize hyperparameters, and does not participate in model parameter updates throughout the process. Testing on the validation set allows for real-time monitoring of whether the model is overfitting or underfitting. During this process, the network parameters with the best validation performance are continuously recorded and saved to determine the model state that performs best on the validation set, laying the foundation for subsequent testing phases.

[0034] After the model has completed training and validation, the optimized network model is used to evaluate predictions on the remaining dataset (i.e., the test set). Performance metrics analysis on the test set provides a clear picture of the model's generalization ability to unknown data. The evaluation results on the test set are a key indicator of the model's effectiveness in real-world applications and can effectively help determine the model's predictive reliability in real-world scenarios.

[0035] Specifically, in order to synergistically optimize pixel-level classification accuracy and region-level shape matching, the model training process of this invention adopts a specially designed joint loss function, which specifically includes: a weighted binary cross-entropy (WBCE) loss term, which effectively solves the problem of extreme imbalance between positive and negative samples that is common in camouflaged target detection by assigning smaller weights to the background pixels that are more numerous and larger weights to the target pixels that are less numerous.

[0036] A weighted intersection-over-union (WIoU) loss term is used to assign weights to pixels based on their gradient values ​​in the real target mask when calculating the WIoU loss. This gives pixels located on the target boundary a higher weight in the loss calculation, thereby enhancing the model's ability to learn the precise geometry of the target.

[0037] The core of the camouflaged target detection method based on boundary feature enhancement of the present invention lies in solving the technical problems of boundary ambiguity and low feature discrimination through an innovative three-stage network architecture.

[0038] This method first uses a boundary enhancement module (BEM) to explicitly predict and enhance the target contour; then, a boundary guidance module (BGM) uses this boundary prior as a dynamic attention signal to guide the network to focus on key features; finally, a multi-level feature aggregation module (FAM) fuses multi-scale contextual information to adapt to targets of different sizes. By combining a joint loss function that collaboratively optimizes pixel classification and region shape, this invention can significantly improve the detection accuracy and boundary segmentation quality of camouflaged targets, and its overall performance is superior to existing technologies. Attached Figure Description

[0039] Figure 1 This is a flowchart illustrating the overall network architecture of the camouflaged target detection method (BFENet) proposed in this invention.

[0040] Figure 2 This is a detailed internal structure and data flow diagram of the boundary enhancement module (BEM) in this invention;

[0041] Figure 3 This is a detailed internal structure and data flow diagram of the Boundary Guidance Module (BGM) in this invention;

[0042] Figure 4 This is a detailed internal structure and data flow diagram of the multi-level feature aggregation module (FAM) in this invention;

[0043] Figure 5 This is a flowchart of the neural network training process provided in an embodiment of the present invention. Detailed Implementation

[0044] Traditional detection methods and some deep learning models often suffer from incomplete detection or inaccurate localization when dealing with camouflaged targets due to the loss of boundary information or low feature discrimination. This invention aims to address the problem that naturally camouflaged targets are difficult to detect accurately using existing computer vision technologies because they blend seamlessly with the background and have blurred boundaries.

[0045] This invention discloses a camouflaged target detection method based on boundary feature enhancement. First, a backbone network is used to extract multi-level features from an image. Then, a boundary enhancement module fuses low-level fine texture features and high-level semantic features to generate a high-quality boundary prediction map. Next, a boundary guidance module uses the aforementioned boundary map as attention guide, causing the network to focus more on boundary regions crucial for distinguishing targets from the background in subsequent processing. Finally, a multi-level feature aggregation module effectively fuses the guided features from each level to adapt to camouflaged targets of different scales. This invention significantly improves the detection accuracy and boundary segmentation integrity of camouflaged targets by explicitly enhancing and utilizing boundary features. It demonstrates superior performance on multiple public datasets, possessing significant academic research value and broad application prospects, and can be used in fields such as ecological monitoring and security reconnaissance.

[0046] Based on the aforementioned design concept, the present invention will be described in detail below with reference to specific implementation methods:

[0047] This invention discloses a camouflaged target detection method based on boundary feature enhancement, referring to... Figure 1 In this example, the method is implemented using a specially designed deep neural network (referred to as BFENet in this embodiment).

[0048] Step 1. Multi-level feature extraction.

[0049] This method takes an RGB image to be detected as input. First, a powerful deep convolutional neural network is used as the backbone, such as Res2Net-50 pre-trained on ImageNet. Res2Net-50 is chosen because its unique bottleneck structure contains residual-like multi-scale connections, enabling the representation of multi-scale features at a finer-grained level, which is highly advantageous for recognizing camouflaged targets. The backbone performs forward propagation on the input image, outputting a set of hierarchical feature maps. This set of feature maps forms a feature pyramid: and These are low-level features with high spatial resolution, preserving rich details such as edges and textures; while , , These are mid- to high-level features with lower spatial resolution but a larger receptive field, containing richer abstract semantic information about objects and the global scene.

[0050] Step 2. Implementation of the Boundary Enhancement Module (BEM).

[0051] Reference Figure 2To overcome the core challenge of ambiguous boundaries, this invention designs a boundary enhancement module (BEM), which aims to generate a high-quality initial boundary prediction map. BEM receives low-level feature maps. and high-level feature maps Internally, it serves as the input, employing a meticulously designed two-branch fusion architecture: Spatial detail branch: the input's... First, a 1x1 convolutional layer is used for channel dimensionality reduction to reduce computation. Then, the feature maps are fed in parallel into max pooling and average pooling layers. Max pooling captures the most salient local features (potentially corresponding to sharp boundaries), while average pooling summarizes the average response of the region (providing smoother context). The outputs of these two pooling operations are summed element-wise (ADD) and then fused again through a 1x1 convolutional layer. This branch aims to fully extract... Multi-scale spatial details; semantically guided branches: input First, a 1x1 convolutional layer is used for channel adjustment, and then bilinear interpolation is used for upsampling to restore its spatial resolution to that of [previous layer]. Processed features Figure 1 The purpose of this branch is to pass down high-level global semantic information to guide the determination of which edges belong to the target boundary.

[0052] Fusion and Output: The outputs of the two branches are concatenated (C) along the channel dimension to form a powerful feature rich in both detail and semantics. This feature is then passed through two convolutional layers (a 3×3 convolution for deep feature interaction and a 1×1 convolution for final integration) and a sigmoid activation function to output a single-channel probability map. In the image, the value of each pixel is between [0,1], which represents the probability that the pixel is the boundary of the target.

[0053] Step 3. Implementation of the Boundary Guidance Module (BGM).

[0054] Reference Figure 3 BEM-generated boundary map As a powerful dynamic prior, it is used to guide the network to focus on subsequent features. The BGM module receives... and any intermediate layer feature map output by the backbone network (like , , () is used as input. Its workflow is as follows:

[0055] First, the boundary map By downsampling, such as through average pooling or strided convolution, its spatial dimensions are made similar to those of the input feature map. Align and then align the results. and The features are concatenated (C) along the channel dimension to form a combined feature that explicitly includes the original features and their corresponding boundary information. This combined feature is then fed into a channel attention generator. Specifically, a global average pooling (GAP) operation compresses the spatial dimension of the feature map to 1×1, resulting in a channel descriptor vector. This vector is then passed through two fully connected layers (FC, forming a bottleneck structure to learn non-linear cross-channel relationships) and a sigmoid activation function to finally generate a channel attention weight vector. . This attention vector Compared with the original feature map Performing element-wise multiplication (X) is equivalent to multiplying by the product of the two products. Each channel is dynamically and differentially reweighted. Guided by boundary information, feature channels related to the target boundary receive higher weights, while channels related to background noise are suppressed. Finally, to maintain the integrity of the information flow and prevent gradient vanishing, a residual connection is used to connect the weighted features with the original feature map. The results are then added together. This result is then smoothed and integrated using a 3x3 convolution, outputting the final guided feature map. .

[0056] Step 4. Multi-level Feature Aggregation Module (FAM) and Result Generation.

[0057] Reference Figure 4 To robustly handle targets of different scales, this invention employs a bottom-up path to aggregate multi-level features guided by the BGM. The multi-level feature aggregation module (FAM) receives feature maps from two adjacent levels (e.g., ... and ) as input, where Typically, upsampling is performed. The core of FAM is a parallel multi-branch dilated convolutional structure. It contains four sets of parallel branches, each employing a pair of asymmetric dilated convolutions (1×N and N×1) with different dilation rates D (e.g., D=1, 3, 5, 7). Dilated convolutions can exponentially expand the receptive field without increasing the number of parameters or computational cost. Different dilation rates correspond to different receptive field sizes, allowing the module to simultaneously capture local details, moderate-range structures, and wide-area contextual information of the target. Using asymmetric convolutional kernels is more efficient than using symmetric N×N kernels. The outputs of all parallel branches are element-wise summed and fused, and residual connections are made with the module's input to ensure stable training and efficient information flow. In this way, the guided feature maps are aggregated layer by layer from high-level to low-level features. Finally, at the highest resolution level, a final feature representation incorporating all scale and boundary information is obtained. This feature is then processed by a simple classification head (such as a 3x3 convolution and a sigmoid function) to generate the final pixel-level camouflage target detection mask (Output).

[0058] Step 5. Define the loss function.

[0059] To effectively train BFENet, this embodiment employs a composite loss function.

[0060]

[0061]

[0062] in, This represents the total number of pixels in the input image (i.e., the total number of samples). For pixel index, For pixels The true label, 0 represents the background, and 1 represents the target. For the model to pixels The predicted probability, and The dynamic weighting coefficients for positive and negative samples satisfy... , The total number of target pixels (number of positive samples). This represents the total number of background pixels (number of negative samples). , .

[0063]

[0064] Spatial weight Defined as:

[0065] Indicates the actual label in pixels The gradient magnitude at the point is adjusted by λ=2.5. This design can exponentially decrease the loss contribution of low gradient regions (flat background) while significantly enhancing the weight of high gradient regions (target edge), forming a strong gradient focus on the boundary region.

[0066] The weights are dynamically adjusted based on the ratio of positive to negative samples, which alleviates the class imbalance problem caused by the fact that the camouflaged target pixels usually only account for a small part of the image. The weighted intersection-union ratio (IoU) modulates the IoU loss using a spatial weight map. This weight map is calculated based on the gradient of the ground truth mask; pixels with larger gradient values ​​in boundary regions have larger weights, thus forcing the network to pay more attention to boundary alignment. In this embodiment, the balance factor α is experimentally set to 0.8 to achieve the best co-optimization effect.

[0067] refer to Figure 5 The main implementation steps of this invention are as follows:

[0068] Data Preparation: The implementation of this invention is based on systematic data preparation. This process begins by selecting an industry-recognized publicly available benchmark dataset for camouflage target detection. To ensure the stability and efficiency of model training, all data undergoes rigorous preprocessing, including image size unification and pixel value standardization. To improve the model's generalization ability and suppress overfitting, this invention further employs a series of data augmentation strategies, such as multi-scale training and stochastic geometric transformations, to expand sample diversity. Finally, the processed dataset is strictly divided into three independent subsets: training, validation, and testing, providing standardized data support for subsequent model parameter learning, hyperparameter optimization, and objective evaluation of final performance.

[0069] Model Construction: The camouflaged target detection model (BFENet) of this invention is implemented using the Python programming language and the PyTorch deep learning framework. Its network structure design integrates a boundary enhancement module (BEM), a boundary guidance module (BGM), and a multi-level feature aggregation module (FAM), aiming to efficiently process the boundary information of camouflaged targets and fuse multi-scale features through a specialized modular design.

[0070] Model Training: Model training follows the principle of iterative optimization. In each training iteration, a forward propagation process is first executed. Training data is input, and the model calculates and generates a prediction mask based on the current parameters. Subsequently, the prediction result is compared with the ground truth, and the total loss value is calculated according to the joint loss function of this invention (including weighted binary cross-entropy (WBCE) and weighted intersection-union ratio (WIoU)). Next, in the backpropagation phase, optimization algorithms such as Adam are used to perform backpropagation based on the gradient of the loss function, and the learnable parameters of the model are updated. This iterative process is repeated until the model converges or reaches the preset number of training epochs, thereby enabling the model to gradually learn deep patterns and rules in the data.

[0071] Model Performance Assessment: After completing one or more training cycles, this invention evaluates the model's current performance using an independent validation set. The core of the evaluation is to verify whether the model's generalization ability and prediction accuracy have reached preset performance thresholds. If the model's various evaluation metrics (such as S-measure, MAE, etc.) perform excellently, the current model can be considered a candidate optimal model and enter the final evaluation stage. Conversely, if the performance does not meet the standards, it indicates that the model has room for optimization, and the model improvement process should be initiated. If the model performance is stable but not optimal, it can also be prioritized to enter the hyperparameter optimization process to explore better performance.

[0072] Hyperparameter tuning: To achieve optimal model performance, a series of key hyperparameters need to be finely tuned and optimized. Hyperparameters include, but are not limited to: learning rate (e.g., initial value set to 1e-4), batch size (e.g., 16), optimizer selection (e.g., Adam), weight decay coefficient (e.g., 0.1), and total number of training iterations (e.g., 50). Precise hyperparameter configuration can significantly improve the model's convergence speed and generalization ability, effectively suppressing overfitting. After tuning, the model training and validation process needs to be re-executed. By comparing performance metrics under different configurations, a set of optimal hyperparameter combinations can be finally determined.

[0073] Model Improvement: To enhance the performance of deep learning models, adjustments can be made across multiple dimensions, including network architecture design, parameter tuning, activation function selection, regularization strategies, data augmentation techniques, and optimization algorithm improvements. This iterative optimization process requires systematically trying different combinations of techniques and quantifying the optimization effects using various evaluation metrics. After performance evaluation, a new phase of model building will begin. This phase involves reconstructing the network architecture, refining parameter tuning, selecting adaptive activation functions, applying targeted regularization methods and data augmentation strategies, and combining these with efficient optimization algorithms to continuously strengthen the model's generalization ability and robustness, thereby significantly improving its performance in real-world applications.

[0074] Model Evaluation: When evaluating deep learning models, in addition to the core accuracy metric, the model's time and space complexity are equally important. Time complexity focuses on the computational resource consumption during model training and inference, directly reflected in training time and inference response speed. Space complexity, on the other hand, emphasizes the storage resources used by the model during operation, encompassing the size of the model's parameter files and runtime memory usage. A comprehensive analysis of time and space complexity allows for a thorough evaluation of the model's operational efficiency and scalability, enabling the precise matching of the most suitable model solution to different application scenarios.

[0075] Summarize

[0076] This invention proposes a camouflaged target detection method based on boundary feature enhancement, aiming to significantly improve the recognition accuracy and segmentation quality of camouflaged targets in natural environments. Traditional visual detection algorithms often perform poorly when faced with targets that are highly integrated with the background due to blurred boundaries and low feature discrimination. To overcome these limitations, this invention innovatively constructs a three-stage deep learning architecture centered on boundaries. Specifically, the method first explicitly predicts and enhances the target contour through a boundary enhancement module, providing reliable prior information for subsequent analysis. Based on this, a boundary guidance module uses this contour as a dynamic attention signal to guide the network to intelligently focus on the key boundary region between the target and the background. Finally, a multi-level feature aggregation module fuses multi-scale contextual information to ensure robust detection of targets of different sizes. This method can automatically and accurately segment camouflaged targets, significantly improving detection accuracy and boundary integrity, providing strong technical support for the intelligent development of fields such as ecological protection and military reconnaissance. This detection method, which deeply integrates boundary feature enhancement strategies, not only improves the algorithm's automated processing capabilities but also enhances its robustness in complex scenarios, showing broad application prospects.

Claims

1. A camouflaged target detection method based on boundary feature enhancement, which uses a neural network detection model to process the input image and outputs the detection result; wherein, The training process of a neural network model is as follows: first, prepare the data; then, build the detection model; and finally, train, validate, and evaluate the detection model. The detection model is characterized by employing a three-stage network architecture to process multi-level feature maps extracted from the input image; In a three-stage network architecture: First, the problem of fuzzy boundary detection is transformed into a learnable subtask: The Boundary Enhancement Module (BEM) is used to fuse low-level feature maps rich in spatial details and high-level feature maps rich in global semantics, resulting in a high-quality initial boundary prediction map. The initial boundary prediction map clearly depicts the outline of the camouflaged target, providing prior information for subsequent processing. Secondly, attention can be guided using prior knowledge: The Boundary Guidance Module (BGM) is used to take the boundary prediction map generated by BEM as a dynamic guidance signal. Based on this, BGM generates channel attention weights and adaptively adjusts the channel response of subsequent feature maps, so that the detection model can focus its computational resources on the boundary regions that are crucial to distinguishing the target from the background. Secondly, the multi-level feature aggregation module (FAM): FAM employs multiple sets of parallel dilated convolutional branches with different dilation rates to capture multi-scale contextual information without sacrificing spatial resolution; this information is then fused. In addition, during the training of the detection model, weighted binary cross-entropy (WBCE) loss and weighted intersection-union ratio (WIoU) loss are used to optimize the model in a coordinated manner from the two levels of pixel classification and region shape.

2. The camouflaged target detection method based on boundary feature enhancement according to claim 1, characterized in that, The detection model uses a deep neural network, BFENet, as its network structure. For the input image to be detected, the processing steps in BFENet include: 1) Hierarchical feature extraction and representation: A pre-trained deep convolutional backbone network is used to perform forward propagation processing on the input image, transforming the original pixel information in the input image into a hierarchical, multi-dimensional feature representation, generating a set of feature maps with different spatial resolutions and semantic depths. 2) Explicit enhancement and prediction of boundary features: Take at least one low-level feature map and one high-level feature map from the multiple layers of feature maps obtained in step 1), and input them into the boundary enhancement module BEM. BEM generates an initial boundary prediction map by fusing low-level and high-level feature maps in parallel, and serves as strong prior knowledge for attention-guided step 3). 3) Dynamic feature guidance based on boundary priors: Input the initial boundary prediction map obtained in step 2) and at least one intermediate layer feature map from the multiple layers of feature maps obtained in step 1) into the boundary guidance module BGM. The BGM module uses the boundary prediction map as a guiding signal to calculate and generate the channel attention weight vector; then it applies this weight vector to the intermediate layer feature map to obtain the guided feature map. 4) Multi-scale aggregation of guided features: The guided feature map obtained in step 3) and other feature maps from other levels are input into the multi-level feature aggregation module FAM; FAM captures and integrates contextual information from different receptive fields through multiple parallel convolutional branches with different dilation rates; then, through a bottom-up aggregation path, it fuses the guided feature map with feature maps from other levels to generate the final feature representation. 5) Generation and output of the final target mask: The final feature representation obtained in step 4) is fed into a classification head network to perform binary classification of each pixel as either target or background, thereby generating the final detection result mask and completing the entire detection process.

3. The camouflaged target detection method based on boundary feature enhancement according to claim 2, characterized in that, In step 2), the structure of BEM includes: A multi-scale context information extraction branch: This branch receives low-level feature maps and applies max pooling and average pooling operations in parallel to the low-level feature maps to capture salient local features and regional average responses, respectively; the results of the two operations are then fused to obtain richer context information; A semantic information delivery branch: This branch receives high-level feature maps and uses upsampling techniques to restore the semantic information of the high-level feature maps to the same spatial resolution as the low-level feature maps, so as to achieve feature map alignment; A fusion output unit: This unit concatenates and fuses the outputs of the multi-scale context information extraction branch and the semantic information transfer branch in the channel dimension; and performs deep information interaction through at least one convolutional layer; finally, the output is normalized to a probability value between 0 and 1 by the Sigmoid activation function to obtain the initial boundary prediction map.

4. The camouflaged target detection method based on boundary feature enhancement according to claim 2, characterized in that, In step 3), the data processing procedure for the BGM includes: The initial boundary prediction map of the input is downsampled to match its spatial size with the intermediate level feature map of the input. The downsampled initial boundary prediction map is concatenated with the intermediate layer feature map along the channel dimension to form a combined feature map containing the original features and boundary priors; Global average pooling is applied to the combined feature map to compress its spatial dimension, resulting in a channel descriptor vector that represents global information. This vector is then passed through at least one fully connected layer and a sigmoid activation function to learn and generate a channel attention weight vector; The channel attention weight vector is multiplied element-wise with the original intermediate-level feature map to achieve dynamic weighting of different channels. The weighted result is then added to the original feature map through residual connections to finally output the guided feature map.

5. The camouflaged target detection method based on boundary feature enhancement according to claim 2, characterized in that, In step 4), the structure of FAM includes: Multiple parallel dilated convolution branches, each using a dilated convolution kernel with a different dilation rate, wherein the dilation rate is set to a set of incrementing values ​​to construct multiple effective receptive fields of different sizes; Any dilated convolution branch uses an asymmetric convolution kernel; The outputs of all parallel dilated convolution branches are fused and combined with the input features of FAM through residual connections to generate aggregated features that are highly adaptable to changes in target scale.

6. The camouflaged target detection method based on boundary feature enhancement according to claim 2, characterized in that, In training the detection model, a joint loss function is used, which includes: A weighted binary cross-entropy (WBCE) loss term is used, which assigns small weights to the abundant background pixels and large weights to the scarce target pixels. A weighted intersection-union ratio (WIoU) loss term is used to assign weights to pixels based on their gradient values ​​in the real target mask when calculating the WIoU loss, so that pixels located on the target boundary have a high weight in the loss calculation.

Citation Information

Patent Citations

  • Hyperspectral image camouflage target detection method based on deep learning

    CN111368712A

  • Camouflage target detection method based on improved YOLO algorithm

    CN112801169A

  • A camouflaged target detection and recognition method based on deep neural network

    CN113449727B