Camouflage target detection method based on initial position guidance and adjacent feature enhancement network

Through the combination of PVT network, position prior extraction module and adaptive morphological perception module, the problem of insufficient accuracy of camouflage object detection in complex backgrounds is solved, and the accurate detection of camouflage object is achieved.

CN120374944APending Publication Date: 2025-07-25河南信息科技学院筹建处
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510449343.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing camouflage object detection methods lack detection accuracy in complex backgrounds, and feature modeling limitations, cross-level information attenuation, and initial positioning deviations make it difficult to accurately distinguish between the target and the background.

Method used

Multi-scale features are extracted using PVT network, combined with position prior extraction module and adaptive morphological perception module, feature information is enhanced through adjacent layer feature guidance mechanism, feature fusion is used by decoder, and accurate camouflage object detection results are generated.

Benefits of technology

It improves the accuracy of camouflage object detection, effectively acquires the initial position prior knowledge of the target, enhances the feature fusion ability, and ensures the accuracy and robustness of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374944A_ABST
    Figure CN120374944A_ABST
Patent Text Reader

Abstract

The invention discloses a camouflage target detection method based on initial position guidance and an adjacent feature enhancement network, belongs to the technical field of deep learning image detection, and aims to solve the problem that an existing camouflage target detection method is insufficient in precision. According to the method, firstly, a PVT network is used as a backbone network to extract features, and a local-to-global dependency relationship is captured in a progressive mode; then, a position prior extraction module (PPEM) is utilized to aggregate low-level and high-level features to obtain rough contour information of a target, and initial position prior knowledge is provided for the model; and deep fine-grained information is extracted by using an adaptive morphological perception module (AMPM) through an adjacent layer feature guide mechanism, diversified multi-scale context information is obtained, and the accuracy of a prediction result is improved. And finally, fusing the position priori knowledge of the camouflage target with the target features by using a decoder to finish the accurate detection of the camouflage target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning image detection. Specifically, it is a camouflaged object detection method based on initial position guidance and adjacent feature enhancement network. Background Technique

[0002] Camouflage is a survival strategy formed through evolution, enabling organisms to blend into the environment to avoid natural enemies or approach prey stealthily. The core task of camouflaged object detection (COD) is to segment objects that are completely fused with the background, and its challenge is significantly higher than that of salient object detection (SOD). Currently, COD technology has been widely applied in fields such as industrial defect detection, agricultural pest and disease identification, and polyp segmentation in medical images.

[0003] Traditional COD methods mainly achieve object recognition based on manually designed features, such as optical characteristics, color, and gradient information. However, in human visual scenes, camouflaged objects and the background often have a high degree of intrinsic similarity, and accurate object perception is crucial for image-level semantic reasoning and cognition. Due to the inherent limitations of manual features (such as insufficient ability to identify foreground elements in complex scenes) and external interference (such as manual annotation features being vulnerable to errors), traditional methods often result in the loss of key information during the detection process and are difficult to accurately distinguish objects from the background.

[0004] In recent years, the introduction of deep learning methods has significantly improved the performance of COD. However, the existing methods still face the following problems: (1) Limitations in feature modeling: Manual features are difficult to capture the subtle differences between objects and the background, and CNN-based models are vulnerable to noise interference in object-background assimilation scenarios; (2) Cross-level information attenuation: During the multi-scale feature fusion process, the simple splicing of shallow details and deep semantics easily leads to the loss of small object structures; (3) Propagation of initial positioning deviation: The candidate box generation error of two-stage detectors will cascade and affect the subsequent classification and regression accuracy. Therefore, the present invention proposes a camouflaged object detection method based on initial position guidance and adjacent feature enhancement network. Summary of the Invention

[0005] The present invention proposes a camouflaged object detection method based on initial position guidance and adjacent feature enhancement network, aiming to solve the problems faced by camouflaged object detection mentioned in the above background technique and improve the detection accuracy of camouflaged objects in complex backgrounds.

[0006] To achieve the above objective, the present invention provides the following technical solution: A camouflaged object detection method based on initial position guidance and adjacent feature enhancement network, including the following implementation steps:

[0007] S1: Obtain a dataset of camouflaged target detection images. The dataset is taken from three existing standard camouflaged target detection datasets, namely CAMO, COD10K, and NC4K, and contains various target images with natural and artificial camouflage.

[0008] S2: Use the PVT network as the backbone network for feature extraction to obtain multi-scale features f i (i = 1, 2, 3, 4);

[0009] S3: Use the Position Prior Extraction Module (PPEM) to aggregate the low-level feature f1 and the high-level feature f4 to obtain the rough contour information of the target, providing the model with initial position prior knowledge f c ;

[0010] S4: Use the Adaptive Morphological Perception Module (AMPM) to extract deep fine-grained information through the adjacent layer feature guidance mechanism, generating enhanced multi-view features f i1 (i = 1, 2, 3, 4);

[0011] S5: Use the Decoder to fuse the initial position prior knowledge f c of the camouflaged target with the enhanced multi-view features f i1 (i = 1, 2, 3, 4) to generate the prediction map P i (i = 1, 2, 3, 4);

[0012] S6: Use convolution and addition operations to fuse the prediction results P i (i = 1, 2, 3, 4) at each level to complete the precise detection of the camouflaged target and obtain the final prediction map;

[0013] S7: Based on the camouflaged target image set, train the network to obtain a camouflaged target detection model, and detect the camouflaged target image to be detected;

[0014] Furthermore, in S2, when performing feature extraction, the PVT network combines the pyramid structure and the spatial attention mechanism, has the ability to capture global dependencies and multi-scale feature extraction capabilities, and obtains the multi-scale features f i (i = 1, 2, 3, 4);

[0015] Furthermore, in S3, the Position Prior Extraction Module (PPEM) aggregates the low-level feature f1 and the high-level feature f4 through convolution, and combines multiplication and addition operations to generate the initial position prior knowledge f c , and the specific operation is: fa = CBR3(CBR3(Conv1(f1))) f m2 = CBR3(Conv1(f4)) f c = Conv1(CBR3(f a3 )) where: CBR3 represents Conv 3×3 + BN + ReLU, Conv 3×3 is the convolutional layer, BN is the batch normalization layer, and ReLU is the activation function; Conv1 represents Conv 1×1;

[0016] Furthermore, in S4, the Adaptive Morphology Perception Module (AMPM) extracts deep fine-grained information through the adjacent layer feature guidance mechanism, deeply mines the deep information and morphological region features to capture high-level semantic clues in sensitive image regions, and generates enhanced multi-view features f i1 (i = 1, 2, 3, 4);

[0017] Furthermore, the Adaptive Morphology Perception Module (AMPM) introduces a Field-of-view Contraction Unit (FCU), which generates local detail features f0 by shrinking the receptive field, reducing background noise interference, and purifying key features. The specific operation is as follows: where: GAP represents global average pooling; CBR1 represents Conv 1×1 + BN + ReLU, Conv 1×1 is the convolutional layer, BN is the batch normalization layer, and ReLU is the activation function;

[0018] Furthermore, the Adaptive Morphology Perception Module (AMPM) introduces a Double Attention Unit (DAU) combined with a multi-branch strategy to capture deep semantic information from f0 while accelerating model inference. Specifically, each branch is composed of a Double Attention Unit (DAU) and CBR n (n = 1, 3, 5) or average pooling (Pooling) to extract and fuse into multi-scale and multi-morphology feature information f1, f3, f p 、f5, and connect them into high-order semantic clues f e . The specific operation is as follows: f1 = CBR1(DAU(f0)) f e=Concat(f1,f3,f p ,f5) Among them: Pooling represents average pooling; Concat represents feature concatenation;

[0019] Furthermore, the dual-path attention unit (DAU) includes spatial attention (SA) and channel attention (CA), which enhances the model's ability to extract and fuse feature information. The specific operation is as follows:

[0020] Furthermore, in S5, the decoder needs two steps to complete the accurate detection of the camouflage target. The first step: fuse the prior knowledge f c of the initial position of the camouflage target with the target feature f i1 (i = 1, 2, 3, 4) to achieve feature enhancement and contour refinement, obtaining P f1 and P f2 . The specific operation is as follows: Among them: Sigmoid represents the activation function; The second step: apply Conv 1×1 to the feature maps P f1 and P f2 to extract the global spatial information f 1s , fuse P f1 and P f2 to generate a rough target structure representation f 2c , and then through pooling and convolution operations, and after activation by the Sigmoid function, obtain the channel attention weight f sc . Finally, fuse f sc with f 1s for recalibration fusion to generate the prediction map P i . The specific operation is as follows: f sc =Sigmoid(Conv1(Pooling(Concat(P f1 ,P f2 )))) Among them: Conv1 represents Conv 1×1; Concat represents feature concatenation; Pooling represents average pooling;

[0021] Furthermore, in S6, use convolution and addition operations to fuse the prediction results P i (i = 1, 2, 3, 4) at each level to complete the accurate detection of the camouflage target and obtain the final prediction map P. The specific operation is as follows:

[0022] Further, in S7, the process of the network model training strategy is as follows:

[0023] 7.1 Obtain the training dataset;

[0024] 7.2 Resize the training dataset images to 384×384 and input them into the network;

[0025] 7.3 Select the Adam algorithm to update the optimizer parameters during the training process;

[0026] 7.4 Use a composite loss function of binary cross entropy loss (BCE) and intersection over union loss (IoU) to supervise the training of the network during the training process, and use the overall loss function to represent it;

[0027] 7.5 Use four metrics S α 、M、 to evaluate the camouflage target detection model;

[0028] Further, during the model training process, the initial position prior knowledge f c in S3 calculates the supervision loss L1 with the ground truth (GT) and performs backpropagation to update the network parameters. The specific operation is as follows: L1 = L BCE (f c , G) + L IoU (f c , G) where: f c represents the initial position prior knowledge, L BCE represents the binary cross entropy loss, L IoU represents the intersection over union loss, and G represents the ground truth;

[0029] Further, during the model training process, the predicted map P i in S5 calculates the supervision loss L2 with the ground truth (GT) and performs backpropagation to update the network parameters. The specific operation is as follows: where: P i represents the predicted map, that is, the output of the decoder, L BCE represents the binary cross entropy loss, L IoU represents the intersection over union loss, and G represents the ground truth;

[0030] Further, during the model training process, the overall loss function in S7.4 is L. The specific operation is as follows: L = L1 + L2

[0031] Further, the optimized camouflage target detection model in S7.5 is evaluated using four widely used metrics, namely: S-measure (S α ), which is used to measure the structural similarity between regions and objects, and the higher the value, the better the performance; Mean Absolute Error (M), which is used to measure the pixel-level difference between the predicted value and the true value, and the lower the value, the better the performance; Evaluates the overall and local accuracy of the camouflage target detection results by comparing the differences between the predicted map and the true map, and the higher the value, the better the performance; Used to calculate the relationship between precision and recall, and the higher the value, the better the performance.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows: A camouflage target detection method based on initial position guidance and adjacent feature enhancement network provided by the present invention systematically improves the detection accuracy of the model by using initial position prior knowledge and adjacent feature enhancement mechanisms; uses the Position Prior Extraction Module (PPEM) to aggregate low-level and high-level features to effectively obtain the initial position prior knowledge of the target; uses the Adaptive Morphology Perception Module (AMPM) to achieve feature enhancement and contour refinement; uses the decoder to achieve comprehensive interaction of features at different levels, ensuring the accuracy of camouflage target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is the overall architecture diagram of camouflage target detection and the framework diagram of the Position Prior Extraction Module (PPEM) provided by the embodiment of the present invention;

[0034] Figure 2 It is the framework diagram of the Adaptive Morphology Perception Module (AMPM) of camouflage target detection provided by the embodiment of the present invention;

[0035] Figure 3 It is the framework diagram of the decoder of camouflage target detection provided by the embodiment of the present invention;

[0036] Figure 4 It is the schematic diagram of quantitative evaluation of camouflage target detection provided by the embodiment of the present invention;

[0037] Figure 5 It is the schematic diagram of visual contrast analysis of camouflage target detection provided by the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] To solve the problem of insufficient accuracy in existing camouflaged target detection methods, a camouflaged target detection method based on initial position guidance and adjacent feature enhancement network is proposed, which specifically includes the following steps:

[0040] Step 1, obtain a camouflaged target detection data image set, preprocess the image data, and divide the preprocessed image data into a training set and a test set;

[0041] Specifically, in the embodiments of the present invention, three standard camouflaged target detection data sets CAMO, COD10K, and NC4K are used;

[0042] Step 2, use the PVT backbone network to extract 4 layers of features f i (i = 1, 2, 3, 4);

[0043] Step 3, use the Position Prior Extraction Module (PPEM) to aggregate the low-level feature f1 and the high-level feature f4 to obtain the rough contour information of the target, and provide the initial position prior knowledge f c ;

[0044] Specifically, as Figure 1 shown (left), the Position Prior Extraction Module (PPEM) aggregates the low-level feature f1 and the high-level feature f4 through convolution, and generates the initial position prior knowledge f of the target by combining multiplication and addition operations c , and the specific operation is: f a = CBR3(CBR3(Conv1(f1))) f m2 = CBR3(Conv1(f4)) f c = Conv1(CBR3(f a3 )) Where: CBR3 represents Conv 3×3 + BN + ReLU, Conv 3×3 is the convolutional layer, BN is the batch normalization layer, and ReLU is the activation function; Conv1 represents Conv 1×1;

[0045] Step 4: Use the Adaptive Morphological Perception Module (AMPM) to extract deep fine-grained information through the adjacent layer feature guidance mechanism, and generate enhanced multi-field-of-view features f i1 (i = 1, 2, 3, 4);

[0046] Specifically, as Figure 2 shown, concatenate the initial position prior knowledge f c and the low-level feature f1 through channels using Concat, and input them into the Adaptive Morphological Perception Module (AMPM). First, enter the Field-of-view Contraction Unit (FCU), and generate local detail features f0 by shrinking the receptive field, reducing background noise interference, and purifying key features. The specific operations are as follows: Among them: GAP represents global average pooling; CBR1 represents Conv 1×1 + BN + ReLU;

[0047] Then, divide the local detail feature f0 into 4 branches. Each branch is composed of a Double Attention Unit (DAU) and CBR n (n = 1, 3, 5) or average pooling (Pooling) to extract and fuse into multi-scale and multi-morphology feature information f1, f3, f p , f5, and connect them into high-order semantic cues f e . The specific operations are as follows: f1 = CBR1(DAU(f0)) f e = Concat(f1, f3, f p , f5) Among them: The Double Attention Unit (DAU) includes spatial attention (SA) and channel attention (CA), which improve the model's ability to extract and fuse feature information. The specific operations are as follows:

[0048] Finally, generate enhanced features f 11 , and the specific operations are as follows: And so on, fuse feature f1 and feature f2 to generate enhanced feature f 21 , fuse feature f2 and feature f3 to generate enhanced feature f 31 , fuse feature f3 and feature f4 to generate f 41 .

[0049] Step 5: Use a decoder to fuse the prior knowledge f of the initial position of the camouflage target c with the target feature f i1 (i = 1, 2, 3, 4) to generate the prediction map P i (i = 1, 2, 3, 4);

[0050] Specifically, as Figure 3 shown, first fuse the prior knowledge f of the initial position of the camouflage target c with the target feature f 11 to achieve feature enhancement and contour refinement, obtaining P f1 and P f2 . The specific operation is as follows:

[0051] Then apply Conv 1×1 to the feature maps P f1 and P f2 to extract the global spatial information f 1s , fuse P f1 and P f2 to generate a rough target structure representation f 2c , and then obtain the channel attention weight f sc after pooling, convolution, and activation by the Sigmoid function. Finally, recalibrate and fuse f sc with f 1s to generate the prediction map P1. The specific operation is as follows: f sc = Sigmoid(Conv1(Pooling(Concat(P f1 , P f2 )))) And so on, fuse the prior knowledge f of the initial position c with the target features f 21 , f 31 , f 41 respectively to generate the prediction maps P2, P3, and P4;

[0052] Step 6: Use convolution and addition operations to fuse the prediction results P i (i = 1, 2, 3, 4) at each level to complete the accurate detection of the camouflage target and obtain the final prediction map P;

[0053] Specifically, as Figure 1 shown, use CBR1 and addition operations to fuse the prediction results P i(i = 1, 2, 3, 4) are fused to obtain the final predicted map P. The specific operation is as follows:

[0054] Step 7: Based on the camouflage target training image set, train the network to obtain a camouflage target detection model, and detect the camouflage target image set to be detected.

[0055] Specifically, the process of training and optimizing the model is as follows:

[0056] 7.1 Obtain the training data sets of the data sets CAMO, COD10K, and NC4K respectively;

[0057] 7.2 Resize the training data set images to 384×384 and input them into the network;

[0058] 7.3 Select the Adam algorithm to update the optimizer parameters during the training process;

[0059] 7.4 Use a composite loss function of binary cross entropy loss (Binary Cross Entropy, BCE) and intersection over union loss (Intersection over Union, IoU) to supervise the training of the network during the training process, and use the overall loss function to represent it;

[0060] Specifically, during the model training process, first calculate the supervision loss L1 between the initial position prior knowledge f c obtained in step 3 and the ground truth map (Ground Truth, GT), and perform backpropagation to update the network parameters. The specific operation is as follows: L1 = L BCE (f c , G) + L IoU (f c , G) where: f c represents the initial position prior knowledge, L BCE represents the binary cross entropy loss, L IoU represents the intersection over union loss, and G represents the ground truth map;

[0061] Then calculate the supervision loss L2 between the predicted map P i (i = 1, 2, 3, 4) obtained in step 5 and the ground truth map (GT) respectively, and perform backpropagation to update the network parameters. The specific operation is as follows: where: P i represents the predicted map, that is, the output of the decoder, L BCE represents the binary cross entropy loss, L IoURepresents the intersection over union loss, and G represents the ground truth map;

[0062] Finally, calculate the overall loss function as L. The specific operation is as follows: L = L1 + L2

[0063] 7.5 Four metrics S α , M, are used to evaluate the camouflage target detection model;

[0064] Specifically, the optimized camouflage target detection model through training is evaluated using four widely used metrics, namely: S-measure (S α ), which is used to measure the structural similarity between regions and objects, and the higher the value, the better the performance; Mean Absolute Error (M), which is used to measure the pixel-level difference between the predicted value and the ground truth value, and the lower the value, the better the performance; The overall and local accuracy of the camouflage target detection results is evaluated by comparing the differences between the predicted map and the ground truth map, and the higher the value, the better the performance; is used to calculate the relationship between precision and recall, and the higher the value, the better the performance.

[0065] The effects of the present invention can be further illustrated by the following experiments:

[0066] Experimental conditions

[0067] All frameworks of the present invention are implemented using the PyTorch framework. The training process uses NVIDIA RTX3090 24G, and the hyperparameters are set as follows: initial learning is 5e-5, batch size is set to 8, the image size of all training and testing processes is 384×384, total training epochs is 150, and the Adam optimizer is used to optimize the loss function.

[0068] After training the model for multiple rounds, the parameters of the best-performing round are saved. Subsequently, the saved best parameters are loaded into the model, the test set images are input into the model, and four metrics S α , M, are used to evaluate the camouflage target detection model. Comparative evaluations are conducted with 12 state-of-the-art methods, including FEDER, DTINet, OSFormer, FDCOD, HitNet, FPNet, FSPNet, OAFormer, diffCOD, IPNet, VSCode, and CSIF. To ensure the fairness of the evaluation, the predicted maps and evaluation metrics of all methods are from the resources provided by the original authors.

[0069] Quantitative evaluation (see Figure 4), the comparative analysis is as follows: Our model has achieved the current best performance in all evaluation metrics on the COD10K dataset; compared with the current optimal CSIF method, the differences in metrics such as S α and M are extremely small, indicating that our model has strong adaptability to scenarios with changes in the scale of camouflaged targets. Especially on the NC4K dataset, in metrics and it outperforms the CSIF method with significant advantages of 5.82% and 0.11% respectively; it has achieved a near-optimal performance on the CAMO dataset. Generally speaking, the position prior extraction module (PPEM) can effectively roughly locate camouflaged targets, and through the adaptive morphology perception module (AMPM), deep fusion of multi-scale features is carried out, and finally accurate detection of camouflaged targets is achieved.

[0070] Visual comparative analysis (see Figure 5 ) shows the qualitative comparison results with 5 leading methods. It can be observed that our model can generate accurate predictions for input images in various challenging scenarios. Specifically, in scenarios involving background color fusion (the first row), small target detection (the third row), and complex background interference (the fourth, sixth, and eighth rows), etc., camouflaged targets are successfully identified, intuitively verifying the superiority of the method; the visualization results in the seventh and eighth rows further show that our model can more accurately locate camouflaged targets and capture more subtle position details, verifying the effectiveness of the position prior extraction module (PPEM); as shown in the third and seventh rows, our model more effectively filters background interference and retains richer context semantic information.

[0071] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0072] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A camouflaged target detection method based on initial position guidance and adjacent feature enhancement network, characterized in that, It includes the following implementation steps: S1: Obtain a camouflaged target detection data image set. The data set is taken from three existing standard camouflaged target detection data sets, namely CAMO, COD10K, and NC4K, and contains various natural and artificial camouflaged target images; S2: Use the PVT network as the backbone network for feature extraction to obtain the multi-scale features f of the image i (i = 1, 2, 3, 4); S3: Aggregate the low-level feature f1 and the high-level feature f4 using the Position Prior Extraction Module (PPEM) to obtain the rough contour information of the target, providing the model with the initial position prior knowledge f c ; S4: Use the Adaptive Morphological Perception Module (AMPM) to extract deep fine-grained information through the adjacent layer feature guidance mechanism, and generate enhanced multi-view features f i1 (i = 1, 2, 3, 4); S5: Use a decoder to fuse the prior knowledge f of the initial position of the camouflage target c with the enhanced multi-view feature f i1 (i = 1, 2, 3, 4) to generate prediction maps P at all levels i (i = 1, 2, 3, 4); S6: Use convolution and addition operations to fuse the prediction results P at each level i (i = 1, 2, 3, 4) to complete the accurate detection of the camouflaged target and obtain the final prediction map; S7: Based on the camouflaged target image set, train the network to obtain a camouflaged target detection model, and detect the to-be-detected camouflaged target image set.

2. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 1, wherein: In S2, when performing feature extraction, the PVT network combines a pyramid structure and a spatial attention mechanism, has the ability to capture global dependencies and multi-scale feature extraction capabilities, and obtains multi-scale features f of the image i (i = 1, 2, 3, 4).

3. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 1, wherein: In the above S3, the Position Prior Extraction Module (PPEM) aggregates the low-level feature f1 and the high-level feature f4 through convolution, and generates the initial target position prior knowledge f by combining multiplication and addition operations c , and the specific operation is as follows: f a = CBR3(CBR3(Conv1(f1))) f m2 = CBR3(Conv1(f4)) f c = Conv1(CBR3(f a3 )) Among them: CBR3 represents Conv3×3 + BN + ReLU, where Conv3×3 is a convolutional layer, BN is a batch normalization layer, and ReLU is an activation function; Conv1 represents Conv1×1.

4. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 1, wherein: In S4, an Adaptive Morphological Perception Module (AMPM) is used to extract deep fine-grained information through an adjacent layer feature guidance mechanism, deeply mining deep information and morphological region features to capture high-level semantic clues in sensitive image regions and generate enhanced multi-view features f i1 (i = 1, 2, 3, 4).

5. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 4, wherein: The adaptive morphology perception module (AMPM) introduces a field-of-view contraction unit (FCU). By shrinking the receptive field, reducing background noise interference, and purifying key features, it generates local detail features f0. The specific operation is as follows: Among them: GAP represents global average pooling; CBR1 represents Conv1×1 + BN + ReLU, where Conv1×1 is a convolutional layer, BN is a batch normalization layer, and ReLU is an activation function.

6. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 4, characterized in that: The adaptive morphology perception module (AMPM) introduces a double-path attention unit (DAU) combined with a multi-branch strategy to capture deep semantic information from f0 while accelerating model inference. Specifically, each branch consists of a double-path attention unit (DAU) and a CBR n (n = 1, 3, 5) or a pooling layer, extracting and fusing into multi-scale and multi-morphology feature information f1, f3, f p , f5, and connecting them into high-order semantic cues f e . The specific operation is as follows: f1 = CBR1(DAU(f0)) f e = Concat(f1, f3, f p , f5) Among them: Pooling represents adaptive average pooling; Concat represents feature concatenation.

7. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 6, wherein: The dual-path attention unit (DAU) includes spatial attention (SA) and channel attention (CA), which are used to improve the model's ability to extract and fuse feature information. The specific operation is as follows: 。 8. A camouflaged target detection method based on an initial position guidance and adjacent feature enhancement network according to claim 1, characterized in that: In the above S5, the decoder needs two steps to complete the accurate detection of the camouflage target. The first step is to fuse the prior knowledge f of the initial position of the camouflage target c with the target feature f i1 (i = 1, 2, 3, 4) to achieve feature enhancement and contour refinement, and obtain P f1 and P f2 . The specific operation is as follows: Among them: Sigmoid represents an activation function; The second step is to apply operations to the feature map P f1 and P f2 by using a 1×1 convolution to extract global spatial information f 1s . Then, P f1 and P f2 are fused to generate a rough target structure representation f 2c . After pooling and convolution operations and activation by the Sigmoid function, the channel attention weight f sc is obtained. Finally, f sc is recalibrated and fused with f 1s to generate the prediction map P i . The specific operation is as follows: f sc = Sigmoid(Conv1(Pooling(Concat(P f1 ,P f2 )))) Among them: Conv1 represents Conv1×1, Concat represents feature concatenation; Pooling represents average pooling.

9. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 1, wherein: In S6, the prediction results P of each level are fused by convolution and addition operations i (i = 1, 2, 3, 4) to complete the accurate detection of the camouflaged target and obtain the final prediction map P. The specific operation is as follows: 。 10. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 1, wherein: In S7, the process of training the network model is as follows: 7.1 Obtain the training data set; 7.2 Resize the training data set images to 384×384 and input them into the network; 7.3 Select the Adam algorithm to update the optimizer parameters during training; 7.4 During training, use a composite loss function of binary cross entropy loss (BCE) and intersection over union loss (IoU) to supervise the training of the network, and use the overall loss function to represent it; 7.5 Four metrics S α , M, are used to evaluate the camouflage target detection model.

11. A camouflaged target detection method based on an initial position guidance and adjacent feature enhancement network according to claim 1, characterized in that: During the model training process, the prior knowledge f of the initial position in S3 c needs to calculate the supervised loss L1 with the ground truth (GT) and perform backpropagation to update the network parameters. The specific operation is as follows: L1 = L BCE (f c , G) + L IoU (f c , G) Among them: f c represents the prior knowledge of the initial position, L BCE represents the binary cross-entropy loss, L IoU represents the intersection over union loss, and G represents the ground truth map.

12. The camouflaged target detection method based on the initial position guidance and adjacent feature enhancement network according to claim 1, wherein: During the model training process, the predicted graph P in S5 i needs to calculate the loss L2 with the ground truth graph (GT) and perform backpropagation to update the network parameters. The specific operation is as follows: Where: P i represents the prediction map, i.e., the output of the decoder, L BCE represents the binary cross-entropy loss, L IoU represents the intersection over union loss, and G represents the ground truth map.

13. A camouflaged target detection method based on an initial position guidance and adjacent feature enhancement network according to claim 10, characterized in that: During the model training process, the overall loss function in 7.4 is L. The specific operation is as follows: L = L1 + L2.

14. A camouflaged target detection method based on an initial position guidance and adjacent feature enhancement network according to claim 10, characterized in that: The four metrics for evaluating the camouflage target detection model in 7.5 are: S-measure (S α ), which is used to measure the structural similarity between regions and objects, and the higher the value, the better the performance; Mean Absolute Error (M), which is used to measure the pixel-level difference between the predicted value and the true value, and the lower the value, the better the performance; evaluates the overall and local accuracy of the camouflage target detection results by comparing the differences between the predicted map and the true map, and the higher the value, the better the performance; is used to calculate the relationship between precision and recall, and the higher the value, the better the performance.