Camouflaged target detection method based on edge feature fusion and high-order spatial interaction

By introducing edge feature fusion and high-order spatial interaction in camouflage object detection, the problem of low detection accuracy when camouflage object is similar to the background in the prior art is solved, and more efficient camouflage object detection performance is achieved.

CN116310693BActive Publication Date: 2025-05-20FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310356445.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-05-20
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

When the existing camouflage object detection method deals with the similar color and texture of the camouflage object to the background, the detection accuracy is low and is greatly affected by interference factors such as lighting changes and background movement.

Method used

A camouflage object detection method based on edge feature fusion and high-order spatial interaction is designed. Edge masks and edge features are generated through edge perception modules, combined with edge enhancement modules and edge feature fusion modules, and built high-order spatial interaction modules and context aggregation modules to enhance feature representation and detection performance.

Benefits of technology

It significantly improves the performance of camouflage object detection, can more effectively distinguish camouflage object from background, and reduces sensitivity to light changes and background movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310693B_ABST
    Figure CN116310693B_ABST
Patent Text Reader

Abstract

The present invention relates to a camouflaged target detection method based on edge feature fusion and high-order spatial interaction. The method comprises: performing data preprocessing, including data pairing and data enhancement processing, to obtain a training data set; designing a camouflaged target detection network based on edge feature fusion and high-order spatial interaction, the network consisting of an edge perception module, an edge enhancement module, an edge feature fusion module, a high-order spatial interaction module, and a context aggregation module; designing a loss function to guide the parameter optimization of the network designed in step B; using the training data set obtained in step A to train the camouflaged target detection network based on edge feature fusion and high-order spatial interaction in step B, converging to Nash equilibrium, and obtaining a trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction; inputting an image to be tested into the trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction, and outputting a mask image of the camouflaged target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image and video processing and computer vision, and particularly to a camouflaged target detection method based on edge feature fusion and high-order spatial interaction. Background Art

[0002] With the development of technology, digital image processing has been widely applied to all aspects of human social life. Moreover, it can also be applied to many aspects such as scientific research. Camouflaged target detection is a new digital image processing task, the purpose of which is to accurately and efficiently detect camouflaged targets embedded in the surrounding environment, segment an image into camouflaged targets and the background, so as to discover the camouflaged targets therein. Camouflage phenomena widely exist in nature. Organisms in nature use their own structures and physiological characteristics to blend into the surrounding environment to avoid predators. Camouflaged target detection can help discover camouflaged organisms in nature and help scientists better study natural organisms. The applicable fields of camouflaged target detection are very broad. In addition to its academic value, camouflaged object detection also helps to promote the search and detection of camouflaged hidden targets in the military, the judgment of diseases in the medical field, and the invasion of locusts in agricultural remote sensing, etc.

[0003] Early camouflaged target detection methods distinguish camouflaged targets and the background based on handcrafted low-level features such as color, texture, geometric gradient, frequency domain, and motion. However, most camouflaged targets are extremely similar in color to the background. Color-based methods can only solve the situation where there is a color difference between the object and the background. Texture-based methods have better detection effects when the color is very close to the background, but perform poorly when the texture of the camouflaged target is similar to the background. Motion-based detection methods rely on motion information, and locate camouflaged targets according to the change differences between the background colors and textures formed by motion. However, this type of method is greatly affected by interference factors and may have problems such as false detection and missed detection due to light changes or background movement. Although the camouflaged target detection methods based on hand-designed features can achieve certain effects, they often fail in complex scenes.

[0004] In recent years, with the in-depth application of deep learning in various fields of computer vision, many camouflaged object detection models based on convolutional neural networks have emerged. These models can model camouflaged object information with powerful feature extraction capabilities and autonomous learning capabilities, improve the accuracy of camouflaged object detection, and enhance the generalization of camouflaged object detection models. The effect is significantly improved compared with traditional camouflage detection methods. The mainstream method is to input an image into the backbone network, extract image features from the backbone network, and then predict the mask of the camouflaged object based on these features to discover the camouflaged object. These methods make full use of the semantic information of the convolutional neural network and expand the receptive field to detect camouflaged objects. However, due to the high similarity between the camouflaged object and the background in terms of color and texture, it is difficult for the camouflaged object detection model based on the convolutional neural network to learn the features of the camouflaged object to distinguish the foreground and the background. Therefore, some methods introduce additional clues, such as edge information, on the original basis to help the camouflaged object detection model based on the convolutional neural network better distinguish the camouflaged object and the background. Therefore, making good use of these additional information can effectively improve the accuracy of camouflaged object detection. The present invention designs a camouflaged object detection method based on edge feature fusion and high-order spatial interaction. This method first extracts image features through the backbone network, then designs an edge perception module to generate an edge mask and edge features, then designs an edge enhancement module and an edge feature fusion module, constructs a high-order spatial interaction module and a context aggregation module, and finally uses the designed network to generate a camouflaged object mask. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a camouflaged object detection method based on edge feature fusion and high-order spatial interaction, which is beneficial to significantly improving the performance of camouflaged object detection by fusing edge features and performing high-order spatial interaction.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A camouflaged object detection method based on edge feature fusion and high-order spatial interaction, comprising the following steps:

[0007] Step A: Perform data preprocessing, including data pairing and data augmentation processing, to obtain a training data set;

[0008] Step B: Design a camouflaged object detection network based on edge feature fusion and high-order spatial interaction. This camouflaged object detection network consists of an edge perception module, an edge enhancement module, an edge feature fusion module, a high-order spatial interaction module, and a context aggregation module;

[0009] Step C: Design a loss function to guide the parameter optimization of the network designed in Step B;

[0010] Step D: Use the training dataset obtained in Step A to train the camouflaged target detection network in Step B until it converges to the Nash equilibrium, obtaining a trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction;

[0011] Step E: Input the image to be tested into the trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction, and output the mask image of the camouflaged target.

[0012] In a preferred embodiment, the specific implementation steps of Step A are as follows:

[0013] Step A1: Combine each original image, its corresponding label image, and edge label image to form an image triple;

[0014] Step A2: Randomly flip each group of image triples left and right, randomly crop, and randomly rotate them; perform color enhancement on the original image by setting random values as parameters to adjust the brightness, contrast, saturation, and sharpness of the original image; add random black or white dots as random noise to the label image corresponding to the original image;

[0015] Step A3: Scale each image in the dataset to an image of the same size with dimensions H×W.

[0016] In a preferred embodiment, the specific implementation steps of Step B are as follows:

[0017] Step B1: Construct an image feature extraction network and use the constructed network to extract image features;

[0018] Step B2: Design an edge perception module and use the designed module to generate an edge mask and edge features;

[0019] Step B3: Design an edge enhancement module and an edge feature fusion module. Use the edge enhancement module to enhance the feature representation with the semantic information of the edge structure of the camouflaged target, and use the edge feature fusion module to generate features that fuse edge information;

[0020] Step B4: Construct a high-order spatial interaction module and a context aggregation module. Use the high-order spatial interaction module to suppress the attention to the background and promote the attention to the foreground, and use the context aggregation module to mine context semantics to enhance object detection;

[0021] Step B5: Design a camouflaged target detection network based on edge feature fusion and high-order spatial interaction, including an edge perception module, an edge feature fusion module, an edge enhancement module, a high-order spatial interaction module, and a context aggregation module. Use the designed network to generate the final camouflaged target mask.

[0022] In a preferred embodiment, the specific implementation steps of step B1 are as follows:

[0023] Step B1: Using Res2Net-50 as the backbone network, extract features from the original image I with an input size of H×W×3. Specifically, denote the feature maps output after the first stage, second stage, third stage, and fourth stage of the original image I as F 1 , F 2 , F 3 , and F 4 , where the size of the feature map F 1 is The size of the feature map F 2 is The size of the feature map F 3 is The size of the feature map F 4 is C = 256.

[0024] In a preferred embodiment, the specific implementation steps of step B2 are as follows:

[0025] Step B21: Design an edge perception module. The input of this edge perception module is the first-stage feature map F 1 and the fourth-stage feature map F 4 extracted in step B1. The output of this edge perception module is the edge feature map F e and the edge mask M e ;

[0026] Step B22: Design a feature fusion block in the edge perception module; the input of this edge perception module is the feature maps F 1 and F 4 extracted in step B1. The input feature map F 1 passes through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to reduce the number of channels and obtain the feature map The input feature map F 4 passes through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to reduce the number of channels and obtain the feature map Use bilinear interpolation to adjust the width and height of the feature map F' 4 to the same width and height as F″ 1 to obtain the feature map Concatenate F' 1 and F″ 4 ″ along the channel dimension and then pass through a channel attention module to obtain the edge feature map The specific formula is expressed as follows:

[0027] F 1 = ReLU(BN(Conv1(F1 )))

[0028] F 4 = ReLU(BN(Conv1(F 4 )))

[0029] ”'

[0030] F 4 = Up(F 4 )

[0031] Fe = SE(Concat(F 1 , F 4 ))

[0032] where Conv1(·) is a convolutional layer with a kernel size of 1×1, BN(·) is a batch normalization operation, ReLU(·) is a ReLU activation function, Up(·) is a bilinear interpolation upsampling, Concat(·, ·) is a concatenation operation along the channel dimension, and SE(·) is a channel attention module;

[0033] Step B23: Design the convolutional block in the edge perception module; the input is the edge feature map F obtained in step B22 e , which successively passes through a 3×3 convolution, a BN layer, a ReLU activation function, a 3×3 convolution, a BN layer, a ReLU activation function, and a 1×1 convolution to finally generate an edge mask The specific formula is as follows:

[0034] Me = Conv1(ReLU(BN(Conv3(ReLU(BN(Conv3(Fe))))))))

[0035] where Conv3(·) is a convolutional layer with a kernel size of 3×3, BN(·) is a batch normalization operation, ReLU(·) is an activation function, and Conv1(·) is a convolution with a kernel size of 1×1.

[0036] In a preferred embodiment, the specific implementation steps of step B3 are as follows:

[0037] Step B31: Design the edge enhancement module. First, design the edge guidance operation in the edge enhancement module; the input is the edge mask M obtained in step B2 e and the feature map obtained in step B1 Downsample the input edge mask M e using bilinear interpolation to the same width and height as the feature map F i to obtain a mask Multiply the mask M' e by the feature map F i and then multiply by Fi Add them up, and then successively pass through a 3×3 convolution, a BN layer, and a ReLU activation function to obtain an edge-guided feature map The specific formula is as follows:

[0038] M' e = Down(M e )

[0039]

[0040] where Down(·) is a bilinear interpolation downsampling operation, is a matrix multiplication operation, is a matrix addition operation, Conv3(·) is a convolution layer with a convolution kernel size of 3×3, BN(·) is a batch normalization operation, and ReLU(·) is an activation function;

[0041] Step B32: Construct the CBAM attention sub-module in the edge enhancement module. The edge enhancement module consists of a serial channel attention SE and a spatial attention SA. The input feature map is the feature map F obtained in Step B32 guide , and obtain edge-enhanced features The specific formula is as follows:

[0042] F ee = SA(SE(F guide ))

[0043] where SE(·) is a channel attention module and SA(·) is a spatial attention module;

[0044] Step B33: Design an edge feature fusion module. The input is the first-stage feature map extracted in Step B1 The edge feature map obtained in Step B2 and the edge mask Multiply the edge mask M e with the feature map F 4 , and then add it to F 4 to obtain a feature map Pass the edge feature map F e successively through a 3×3 convolution, a BN layer, and a ReLU activation function to obtain a feature map with reduced channels Pass F M and F' e concatenate them along the channel dimension and then successively pass through a 3×3 convolution, a Swish activation function, a 3×3 convolution, a Swish activation function, an SE module, and a 3×3 convolution. Finally, add the feature map F' e to obtain a feature map that fuses edge information for the first time Pass the feature map F' e through the SE module and then combine it with Concatenate along the channel dimension, and then pass through a 3×3 convolution to obtain the feature map for the second fusion of edge information Finally, the feature map is added to the feature map F 1 to obtain the final feature map for the fused edge information The specific formula is as follows:

[0045]

[0046] F' e = ReLU(BN(Conv3(F e )))

[0047]

[0048]

[0049]

[0050] where is the matrix multiplication operation, is the matrix addition operation, Conv3(·) is a convolutional layer with a convolutional kernel size of 3×3, BN(·) is the batch normalization operation, ReLU(·) is the activation function, Swish(·) is the Swish activation function, SE(·) is the channel attention module, and Concat(·,·) is the concatenation operation along the channel dimension.

[0051] In a preferred embodiment, the specific implementation steps of step B4 are as follows:

[0052] Step B41: First, construct the gated convolution module in the high-order spatial interaction module, and denote the input feature map of this module as Normalize the input feature map F α (denoted as LN 1 ) to obtain the normalized feature map Then, pass through a 1×1 convolution to double its channels, obtaining the feature map Split along the channel into two feature maps Input q into the depthwise separable convolution to obtain the feature map Then split it into n (n is the order) feature maps where Multiply the feature map p 0 by the feature map q 0 and then pass through a 1×1 convolution to double its channels, obtaining the first spatial interaction feature map Multiply the feature map p 1 by the feature map q 1 and then expand its channels to twice the original through a 1×1 convolution to obtain the second spatial interaction feature map Then iterate successively to the feature map p n-1 and the feature map q n-1 Multiply them and then pass through a convolution layer with the same number of input channels and output channels and a convolution kernel size of 1×1 to obtain the n - th spatial interaction feature map Finally, add the input feature map F α and p n to get the intermediate output feature map The specific formula is as follows:

[0053]

[0054]

[0055]

[0056] Q = DWConv(q)

[0057]

[0058]

[0059]

[0060] where Split(·) is splitting along the channel dimension, DWConv(·) is depth - wise separable convolution, Conv1(·) is a convolution layer with a convolution kernel size of 1×1, is matrix multiplication operation, is matrix addition operation;

[0061] Step B42: Construct the feed - forward module in the high - order spatial interaction module. The input is the feature map F obtained in step B41 mid , perform layer normalization on F mid , denoted as LN 2 , and then input it into two fully - connected layers, denoted as MLP. The output of the two fully - connected layers is added to the feature map F mid to obtain the high - order spatial interaction feature The specific formula is as follows:

[0062]

[0063] where is matrix addition operation;

[0064] Step B43: Construct the channel reduction module in the high-order spatial interaction module. The input is F obtained in Step B42 hsi , and pass F hsi through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to obtain the high-order spatial interaction feature map with reduced channels The specific formula is as follows:

[0065] F’ hsi = ReLU(BN(Conv1(F hsi )))

[0066] where Conv1(·) is a convolutional layer with a kernel size of 1×1, BN(·) is a batch normalization operation, and ReLU(·) is an activation function;

[0067] Step B44: First, construct the convolutional block in the context aggregation module. Denote the input of this context aggregation module as two feature maps of different scales and First, perform bilinear interpolation upsampling on the feature map F high to adjust its width and height to the same width and height as F low , and then concatenate it with F low along the channel dimension, and then pass through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to obtain the feature map Then, divide F cat into four feature maps along the channel dimension on average and Add and and then pass through a 3×3 convolution, a BN layer, and a ReLU activation function in sequence to obtain the feature map Add and The sum of the three is then passed through a 3×3 convolution with a dilation rate of 2, a BN layer, and a ReLU activation function in sequence to obtain the feature map Add and The sum of the three is then passed through a 3×3 convolution with a dilation rate of 3, a BN layer, and a ReLU activation function in sequence to obtain the feature map Add and and then pass through a 3×3 convolution with a dilation rate of 4, a BN layer, and a ReLU activation function in sequence to obtain the feature map Then, concatenate and along the channel dimension and then pass through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to obtain the feature map Finally, F cat and F' catAfter addition, it successively passes through a 3×3 convolution, a BN layer, and a ReLU activation function to obtain a context feature map The specific formula is as follows:

[0068] F cat = ReLU(BN(Conv1(Concat(F low , Up(F high ))))

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076] Among them, Up(·) is a bilinear interpolation upsampling operation, Concat(·,·) and Concat(·,·,·,·) are concatenation operations along the channel dimension, is a matrix addition operation, Conv3(·) is a convolutional layer with a convolution kernel size of 3×3, Conv3 d=i (·) is a 3×3 convolution with a dilation rate of i, Conv1(·) is a convolutional layer with a convolution kernel size of 1×1, BN(·) is a batch normalization operation, ReLU(·) is an activation function, and Split(·) is an equal splitting operation along the channel dimension.

[0077] In a preferred embodiment, the specific implementation steps of step B5 are as follows:

[0078] Step B5: Design a camouflage target detection network based on edge feature fusion and high-order spatial interaction, including an edge perception module, an edge feature fusion module, an edge enhancement module, a high-order spatial interaction module, and a context aggregation module; input the original image, and obtain four feature maps of different scales through the backbone network in step B1 and Input F 1 and F 4 into the edge perception module in step B2 to obtain an edge feature map and an edge mask Then construct three edge enhancement modules in step B3, denoted as EEM1 , EEM 2 and EEM 3 , where the input of EEM 1 is the fourth-stage feature map F extracted in step B1 4 and the edge mask M obtained in step B2 e , and the output is the edge-enhanced feature EEM 2 's input is the third-stage feature map F extracted in step B1 3 and the edge mask M obtained in step B2 e , and the output is the edge-enhanced feature EEM 3 's input is the second-stage feature map F extracted in step B1 2 and the edge mask M obtained in step B2 e , and the output is the edge-enhanced feature Then, construct an edge feature fusion module in step B3, with the input being the first-stage feature map F extracted in step B1 1 , the edge feature map F obtained in step B2 e and the edge mask M e , and the output is the feature map integrating edge information Then, construct four high-order spatial interaction modules in step B4, denoted as HSIM 1 , HSIM 2 , HSIM 3 and HSIM 4 , and their inputs are the feature maps obtained in step B3 and The outputs are respectively and Immediately afterwards, construct three context aggregation modules in step B4, denoted as CAM 1 , CAM 2 and CAM 3 , where the input of CAM 1 is the feature map and The output is the context feature map CAM 2 's input is the output of CAM 1 and the feature map The output is the context feature map CAM 3 's input is the output of CAM 2 and the feature map ​​The output is the context feature map For the edge mask M e , perform bilinear interpolation upsampling to enlarge it by 4 times to obtain the final edge mask M edge ; For the context feature map After compressing it to 1 channel through a 1×1 convolution, perform bilinear interpolation upsampling to enlarge it by 16 times to obtain the first-stage camouflage target mask For the context feature map After compressing it to 1 channel through a 1×1 convolution, perform bilinear interpolation upsampling to enlarge it by 8 times to obtain the second-stage camouflage target mask For the context feature map After compressing it to 1 channel through a 1×1 convolution, perform bilinear interpolation upsampling to enlarge it by 4 times to obtain the final camouflage target mask The specific formula is as follows:

[0079] M edge = Up scale=4 (M e )

[0080]

[0081]

[0082]

[0083] Among them, Up scale=4 (·) is bilinear interpolation upsampling with a multiple of 4, Up scale=8 (·) is bilinear interpolation upsampling with a multiple of 8, Up scale=16 (·) is bilinear interpolation upsampling with a multiple of 16, and Conv1(·) is a convolutional layer with a kernel size of 1×1 and an output channel number of 1.

[0084] In a preferred embodiment, the specific implementation steps of step C are as follows:

[0085] Step C: Design a loss function as a constraint to optimize the camouflage target detection network based on edge feature fusion and high-order spatial interaction. The specific formula is as follows:

[0086]

[0087] Among them, G camo represents the label image corresponding to the original image I, and G edge represents the edge label image corresponding to the original image I, represents the total loss function, represents the weighted binary cross-entropy loss, represents the weighted intersection over union loss, represents the Dice coefficient loss, and λ represents the weight of this loss.

[0088] In a preferred embodiment, the specific implementation steps of step D are as follows:

[0089] Step D1: Randomly divide the training data set obtained in step A into several batches, and each batch contains N pairs of images;

[0090] Step D2: Input the original image I, and after passing through the camouflaged target detection network based on edge feature fusion and high-order spatial interaction in step B, obtain the edge mask M edge , the camouflaged target mask and Calculate the loss using the formula in step C

[0091] Step D3: Calculate the gradients of the parameters in the network using the backpropagation method according to the loss, and update the network parameters using the Adam optimization method;

[0092] Step D4: Repeat steps D1 to D3 in batches until the value of the objective loss function of the network converges to the Nash equilibrium, save the network parameters, and obtain the camouflaged target detection model based on edge feature fusion and high-order spatial interaction; for the tested camouflaged target image, use the one with the largest resolution among the three camouflaged target masks predicted by the model as the final camouflaged target mask.

[0093] Compared with the prior art, the present invention has the following beneficial effects: On the basis of making good use of edge information, the present invention better fuses edge information with backbone features, and at the same time performs high-order spatial interaction on the features, which can better learn the relationship between the camouflaged target and the background in the image. The present invention proposes a camouflaged target detection method based on edge feature fusion and high-order spatial interaction, which respectively generates edge features and edge masks through an edge perception module, fuses edge information in an edge enhancement module and an edge feature fusion module, performs high-order spatial interaction on the fused features in a high-order spatial interaction module, and finally aggregates features at different levels in a context aggregation module, and finally can output high-quality camouflaged target masks. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 is the implementation flowchart of the method in the preferred embodiment of the present invention.

[0095] Figure 2 is the structural diagram of the camouflaged target detection network based on edge feature fusion and high-order spatial interaction in the preferred embodiment of the present invention.

[0096] Figure 3It is the structural diagram of the edge perception module in the preferred embodiment of the present invention.

[0097] Figure 4 It is the structural diagram of the edge enhancement module in the preferred embodiment of the present invention.

[0098] Figure 5 It is the structural diagram of the edge feature fusion module in the preferred embodiment of the present invention.

[0099] Figure 6 It is the structural diagram of the high-order spatial interaction module in the preferred embodiment of the present invention.

[0100] Figure 7 It is the structural diagram of the context aggregation module in the preferred embodiment of the present invention. Detailed implementation manners

[0101] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0102] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0103] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0104] The present invention provides a camouflaged target detection method based on edge feature fusion and high-order spatial interaction, as Figure 1-7 shown, including the following steps:

[0105] Step A: Perform data preprocessing, including data pairing and data augmentation processing, to obtain a training data set;

[0106] Step B: Design a camouflaged target detection network based on edge feature fusion and high-order spatial interaction. This network is composed of an edge perception module, an edge enhancement module, an edge feature fusion module, a high-order spatial interaction module, and a context aggregation module;

[0107] Step C: Design a loss function to guide the parameter optimization of the network designed in Step B;

[0108] Step D: Use the training dataset obtained in Step A to train the camouflaged target detection network in Step B until it converges to the Nash equilibrium, and obtain a trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction;

[0109] Step E: Input the image to be tested into the trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction, and output the mask image of the camouflaged target.

[0110] Furthermore, Step A includes the following steps:

[0111] Step A1: Combine each original image, its corresponding label image, and edge label image into an image triple.

[0112] Step A2: Randomly flip each group of image triples left and right, randomly crop, and randomly rotate; perform color enhancement on the original image by setting random values as parameters to adjust the brightness, contrast, saturation, and sharpness of the original image; add random black or white dots as random noise to the label image corresponding to the original image.

[0113] Step A3: Scale each image in the dataset to an image of the same size with dimensions H×W.

[0114] Furthermore, Step B includes the following steps:

[0115] Step B1: Construct an image feature extraction network and use the constructed network to extract image features.

[0116] Step B2: Design an edge perception module and use the designed module to generate an edge mask and edge features.

[0117] Step B3: Design an edge enhancement module and an edge feature fusion module. Use the edge enhancement module to enhance the feature representation with the semantic of the edge structure of the camouflaged target, and use the edge feature fusion module to generate features that fuse edge information.

[0118] Step B4: Construct a high-order spatial interaction module and a context aggregation module. Use the high-order spatial interaction module to suppress the attention to the background and promote the attention to the foreground, and use the context aggregation module to mine context semantics to enhance object detection.

[0119] Step B5: Design a camouflaged target detection network based on edge feature fusion and high-order spatial interaction, including an edge perception module, an edge feature fusion module, an edge enhancement module, a high-order spatial interaction module, and a context aggregation module. Use the designed network to generate the final camouflaged target mask.

[0120] Furthermore, Step B1 includes the following steps:

[0121] Step B1: Using Res2Net-50 as the backbone network, extract features from the original image I with an input size of H×W×3. Specifically, denote the feature maps output by the original image I after the first stage, the second stage, the third stage, and the fourth stage as F 1 , F 2 , F 3 , and F 4 , where the size of the feature map F 1 is The size of the feature map F 2 is The size of the feature map F 3 is The size of the feature map F 4 is

[0122] Further, as Figure 3 shown, Step B2 includes the following steps:

[0123] Step B21: Design an edge perception module. The input of this module is the first-stage feature map F 1 and the fourth-stage feature map F 4 extracted in Step B1. The output of this module is the edge feature map F e and the edge mask M e .

[0124] Step B22: Design a feature fusion block in the edge perception module. The input of this module is the feature maps F 1 and F 4 extracted in Step B1. The input feature map F 1 sequentially passes through a 1×1 convolution, a BN layer, and a ReLU activation function to reduce the number of channels to obtain the feature map The input feature map F 4 sequentially passes through a 1×1 convolution, a BN layer, and a ReLU activation function to reduce the number of channels to obtain the feature map Use bilinear interpolation to adjust the width and height of the feature map F' 4 to the same width and height as F' 1 to obtain the feature map Concatenate F’ 1 and F”’ 4 along the channel dimension and then pass through a channel attention module to obtain the edge feature map The specific formula is expressed as follows:

[0125] F’ 1 = ReLU(BN(Conv1(F 1 )))

[0126] F’ 4= ReLU(BN(Conv1(F 4 )))

[0127] F” 4 = Up(F’ 4 )

[0128] F e = SE(Concat(F’ 1 , F” 4 ))

[0129] Among them, Conv1(·) is a convolutional layer with a kernel size of 1×1, BN(·) is a batch normalization operation, ReLU(·) is a ReLU activation function, Up(·) is a bilinear interpolation upsampling, Concat(·, ·) is a concatenation operation along the channel dimension, and SE(·) is a channel attention module.

[0130] Step B23: Design the convolutional block in the edge perception module. The input is the edge feature map F obtained in step B22 e , which successively passes through a 3×3 convolution, a BN layer, a ReLU activation function, a 3×3 convolution, a BN layer, a ReLU activation function, and a 1×1 convolution to finally generate an edge mask The specific formula is as follows:

[0131] M e = Conv1(ReLU(BN(Conv3(ReLU(BN(Conv3(F e ))))))))

[0132] Among them, Conv3(·) is a convolutional layer with a kernel size of 3×3, BN(·) is a batch normalization operation, ReLU(·) is an activation function, and Conv1(·) is a convolution with a kernel size of 1×1.

[0133] Furthermore, as shown in Figure 4 and Figure 5 , step B3 includes the following steps:

[0134] Step B31: Design the edge enhancement module. First, design the edge guidance operation in the edge enhancement module. The input is the edge mask M obtained in step B2 e and the feature map obtained in step B1 Perform bilinear interpolation downsampling on the input edge mask M e to the same width and height as the feature map F i to obtain a mask Multiply the mask M' e with the feature map F i , and then with F iAdd them up, and then successively pass through a 3×3 convolution, a BN layer, and a ReLU activation function to obtain an edge-guided feature map The specific formula is as follows:

[0135] M' e = Down(M e )

[0136]

[0137] where Down(·) is a bilinear interpolation downsampling operation, is a matrix multiplication operation, is a matrix addition operation, Conv3(·) is a convolution layer with a convolution kernel size of 3×3, BN(·) is a batch normalization operation, and ReLU(·) is an activation function.

[0138] Step B32: Construct the CBAM attention sub-module in the edge enhancement module. This module consists of a serial channel attention SE and a spatial attention SA. The input feature map is the feature map F obtained in step B32 guide , and obtain the edge-enhanced feature The specific formula is as follows:

[0139] F ee = SA(SE(F guide ))

[0140] where SE(·) is a channel attention module and SA(·) is a spatial attention module.

[0141] Step B33: Design an edge feature fusion module. The input is the first-stage feature map extracted in step B1 The edge feature map obtained in step B2 and the edge mask Multiply the edge mask M e with the feature map F 4 , and then add it to F 4 to obtain the feature map Pass the edge feature map F e successively through a 3×3 convolution, a BN layer, and a ReLU activation function to obtain a feature map with reduced channels Concatenate F M and F' e along the channel dimension and then successively pass through a 3×3 convolution, a Swish activation function, a 3×3 convolution, a Swish activation function, an SE module, and a 3×3 convolution. Finally, add the feature map F' e to obtain the feature map with edge information fused for the first time Pass the feature map F' eAfter passing through the SE module and then concatenating along the channel dimension, and then passing through a 3×3 convolution to obtain the feature map for fusing edge information for the second time Finally, the feature map is added to the feature map F 1 to obtain the final feature map for fusing edge information The specific formula is as follows:

[0142]

[0143] F' e = ReLU(BN(Conv3(F e )))

[0144]

[0145]

[0146]

[0147] where is matrix multiplication operation, is matrix addition operation, Conv3(·) is a convolutional layer with a convolution kernel size of 3×3, BN(·) is batch normalization operation, ReLU(·) is an activation function, Swish(·) is Swish activation function, SE(·) is a channel attention module, and Concat(·,·) is concatenation operation along the channel dimension.

[0148] Furthermore, as Figure 6 and Figure 7 shown, step B4 includes the following steps:

[0149] Step B41: First, construct the gated convolution module in the high-order spatial interaction module, and denote the input feature map of this module as Perform layer normalization (denoted as LN α 1 ) on the input feature map Fto obtain the normalized feature map Then, expand its channels to twice the original by passing through a 1×1 convolution to obtain the feature map Split along the channel into two feature maps Input q into the depthwise separable convolution to obtain the feature map Then split it into n (n is the order) feature maps where The feature map p 0 is combined with the feature map q 0After multiplication, a 1×1 convolution is used to double its number of channels, obtaining the first spatial interaction feature map Feature map p 1 is multiplied by feature map q 1 After multiplication, a 1×1 convolution is used to double its number of channels, obtaining the second spatial interaction feature map Then, it is iteratively applied to feature map p n-1 and feature map q n-1 After multiplication, a convolution layer with the same number of input and output channels and a kernel size of 1×1 is used to obtain the nth spatial interaction feature map Finally, the input feature map F α is added to p n to obtain the intermediate output feature map The specific formula is as follows:

[0150]

[0151]

[0152]

[0153] Q = DWConv(q)

[0154]

[0155]

[0156]

[0157] where Split(·) splits along the channel dimension, DWConv(·) is depthwise separable convolution, Conv1(·) is a convolution layer with a kernel size of 1×1, is matrix multiplication operation, is matrix addition operation.

[0158] Step B42: Construct the feed-forward module in the high-order spatial interaction module. The input is the feature map F obtained in Step B41 mid , perform layer normalization on F mid (denoted as LN 2 ), and then input it into two fully connected layers (denoted as MLP). The output of the two fully connected layers is added to the feature map F mid to obtain the high-order spatial interaction feature The specific formula is as follows:

[0159]

[0160] where is matrix addition operation.

[0161] Step B43: Construct the channel reduction module in the high - order spatial interaction module. The input is F obtained in Step B42 hsi , and pass F hsi through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to obtain the high - order spatial interaction feature map after channel reduction The specific formula is as follows:

[0162] F’ hsi = ReLU(BN(Conv1(F hsi )))

[0163] where Conv1(·) is a convolutional layer with a kernel size of 1×1, BN(·) is a batch normalization operation, and ReLU(·) is an activation function.

[0164] Step B44: First, construct the convolutional block in the context aggregation module. Denote the input of this module as two feature maps of different scales and First, perform bilinear interpolation upsampling on the feature map F high to adjust its width and height to the same width and height as F low , and then concatenate it with F low along the channel dimension. Then, pass through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to obtain the feature map Then, divide F cat into four feature maps along the channel dimension on average and Add and and then pass through a 3×3 convolution, a BN layer, and a ReLU activation function in sequence to obtain the feature map Add and and the sum of the three passes through a 3×3 convolution with a dilation rate of 2, a BN layer, and a ReLU activation function in sequence to obtain the feature map Add and and the sum of the three passes through a 3×3 convolution with a dilation rate of 3, a BN layer, and a ReLU activation function in sequence to obtain the feature map Add and and the sum passes through a 3×3 convolution with a dilation rate of 4, a BN layer, and a ReLU activation function in sequence to obtain the feature map Then, concatenate and along the channel dimension and pass through a 1×1 convolution, a BN layer, and a ReLU activation function in sequence to obtain the feature map Finally, Fcat and F' cat After adding them together, they successively pass through a 3×3 convolution, a BN layer, and a ReLU activation function to obtain a context feature map The specific formula is as follows:

[0165] F cat = ReLU(BN(Conv1(Concat(F low , Up(F high )))))

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172]

[0173] Among them, Up(·) is a bilinear interpolation upsampling operation, Concat(·,·) and Concat(·,·,·,·) are concatenation operations along the channel dimension, is a matrix addition operation, Conv3(·) is a convolutional layer with a convolution kernel size of 3×3, Conv3 d=i (·) is a 3×3 convolution with a dilation rate of i, Conv1(·) is a convolutional layer with a convolution kernel size of 1×1, BN(·) is a batch normalization operation, ReLU(·) is an activation function, and Split(·) is an equal split operation along the channel dimension.

[0174] Furthermore, as Figure 2 shown, step B5 includes the following steps:

[0175] Step B5: Design a camouflage target detection network based on edge feature fusion and high-order spatial interaction, including an edge perception module, an edge feature fusion module, an edge enhancement module, a high-order spatial interaction module, and a context aggregation module. Input the original image, and obtain four feature maps of different scales through the backbone network in step B1 and Input F 1 and F 4 into the edge perception module in step B2 to obtain an edge feature map and an edge mask Then, construct the edge enhancement modules in step B3, denoted as EEM 1 , EEM 2 and EEM 3 , where the input of EEM 1 is the fourth-stage feature map F 4 extracted in step B1 and the edge mask M e obtained in step B2, and the output is the edge-enhanced feature EEM 2 has the third-stage feature map F 3 extracted in step B1 and the edge mask M e obtained in step B2 as its input, and the output is the edge-enhanced feature EEM 3 takes the second-stage feature map F 2 extracted in step B1 and the edge mask M e obtained in step B2 as its input, and the output is the edge-enhanced feature Next, construct an edge feature fusion module in step B3, with the first-stage feature map F 1 extracted in step B1, the edge feature map F e obtained in step B2, and the edge mask M e as its inputs, and the output is the feature map integrating edge information Then, construct four high-order spatial interaction modules in step B4, denoted as HSIM 1 , HSIM 2 , HSIM 3 and HSIM 4 , whose inputs are the feature maps and obtained in step B3 respectively, and the outputs are and Immediately afterwards, construct three context aggregation modules in step B4, denoted as CAM 1 , CAM 2 and CAM 3 , where the input of CAM 1 is the feature map and , and the output is the context feature map CAM 2 takes the output of CAM 1 as its input and the feature map , and the output is the context feature map CAM 3 takes the output of CAM 2 as its input and the feature map The output is the context feature map For the edge mask M e , it is upsampled by bilinear interpolation and magnified by 4 times to obtain the final edge mask M edge . For the context feature map After being compressed to 1 channel by a 1×1 convolution, it will be upsampled by bilinear interpolation and magnified by 16 times to obtain the first-stage camouflage target mask For the context feature map After being compressed to 1 channel by a 1×1 convolution, it is upsampled by bilinear interpolation and magnified by 8 times to obtain the second-stage camouflage target mask For the context feature map After being compressed to 1 channel by a 1×1 convolution, it is upsampled by bilinear interpolation and magnified by 4 times to obtain the final camouflage target mask The specific formula is as follows:

[0176] M edge = Up scale=4 (M e )

[0177]

[0178]

[0179]

[0180] Among them, Up scale=4 (·) is bilinear interpolation upsampling with a multiple of 4, Up scale=8 (·) is bilinear interpolation upsampling with a multiple of 8, Up scale=16 (·) is bilinear interpolation upsampling with a multiple of 16, and Conv1(·) is a convolutional layer with a kernel size of 1×1 and an output channel number of 1

[0181] Furthermore, step C includes the following steps:

[0182] Step C: Design a loss function as a constraint to optimize the camouflage target detection network based on edge feature fusion and high-order spatial interaction. The specific formula is as follows:

[0183]

[0184] Among them, G camo represents the label image corresponding to the original image I, and G edge represents the edge label image corresponding to the original image I represents the total loss function represents the weighted binary cross-entropy loss Expressed as weighted intersection over union loss, represents Dice coefficient loss, and λ represents the weight of this loss.

[0185] Furthermore, step D is implemented as follows:

[0186] Step D1: Randomly divide the training dataset obtained in step A into several batches, and each batch contains N pairs of images.

[0187] Step D2: Input the original image I, and after passing through the camouflaged target detection network based on edge feature fusion and high-order spatial interaction in step B, obtain the edge mask M edge , the camouflaged target mask and calculate the loss using the formula in step C

[0188] Step D3: Calculate the gradients of the parameters in the network using the backpropagation method according to the loss, and update the network parameters using the Adam optimization method.

[0189] Step D4: Repeat steps D1 to D3 in batches until the value of the objective loss function of the network converges to the Nash equilibrium, save the network parameters, and obtain the camouflaged target detection model based on edge feature fusion and high-order spatial interaction. For the tested camouflaged target image, use the one with the largest resolution among the three camouflaged target masks predicted by the model as the final camouflaged target mask.

[0190] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects produced do not exceed the scope of the technical solutions of the present invention, fall within the protection scope of the present invention.

Claims

1. A camouflaged target detection method based on edge feature fusion and high-order spatial interaction, characterized in that: The steps include: Step A: perform data preprocessing, including data pairing and data enhancement, to obtain a training data set; Step B, designing a camouflaged target detection network based on edge feature fusion and high-order spatial interaction, the camouflaged target detection network consists of an edge perception module, an edge enhancement module, an edge feature fusion module, a high-order spatial interaction module, and a context aggregation module; Step C: Design a loss function to guide the parameter optimization of the network designed in step B; Step D: using the training data set obtained in step A to train the camouflaged target detection network based on edge feature fusion and high-order spatial interaction in step B, converge to Nash equilibrium, and obtain a trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction; Step E: input the image to be tested into the trained camouflaged target detection model based on edge feature fusion and high-order spatial interaction, and output a mask image of the camouflaged target; The specific implementation steps of step B are as follows: Step B1, constructing an image feature extraction network, and using the constructed network to extract image features; Step B2: designing an edge perception module, and using the designed module to generate edge masks and edge features; Step B3, designing an edge enhancement module and an edge feature fusion module, using the edge enhancement module to enhance the feature representation with the semantics of the camouflaged target edge structure, and using the edge feature fusion module to generate features that fuse edge information; Step B4: construct a high-order spatial interaction module and a context aggregation module, use the high-order spatial interaction module to suppress attention to the background and promote attention to the foreground, and use the context aggregation module to mine context semantics to enhance object detection; Step B5, designing a camouflaged target detection network based on edge feature fusion and high-order spatial interaction, including an edge perception module, an edge feature fusion module, an edge enhancement module, a high-order spatial interaction module, and a context aggregation module, and using the designed network to generate the final camouflaged target mask; The specific implementation steps of step B1 are as follows: Step B1, using Res2Net-50 as the backbone network, extract features from the original image I with an input size of H×W×3. Specifically, the feature maps output by the original image I after the first stage, the second stage, the third stage, and the fourth stage are recorded as F1, F2, F3, and F4, respectively, where the size of the feature map F1 is The size of feature map F2 is The size of feature map F3 is The size of feature map F4 is C = 256; The specific implementation steps of step B2 are as follows: Step B21, design an edge perception module, the input of which is the first stage feature map F1 and the fourth stage feature map F4 extracted in step B1, and the output of which is the edge feature map F e and edge mask M e ; Step B22, designing a feature fusion block in an edge perception module; The input of the edge perception module is the feature maps F1 and F4 extracted in step B1. The input feature map F1 is successively processed through 1×1 convolution, BN layer, and ReLU activation function to reduce the number of channels to obtain the feature map The input feature map F4 is obtained by successively passing through 1×1 convolution, BN layer, and ReLU activation function to reduce the number of channels. Use bilinear interpolation to adjust the width and height of the feature map F'4 to the same width and height as F'1, and get the feature map F'1 and F'4 ' After splicing along the channel dimension, the edge feature map is obtained by the channel attention module. The specific formula is as follows: F′1=ReLU(BN(Conv1(F1))) F′4=ReLU(BN(Conv1(F4))) F″4=Up(F′4) F e =SE(Concat(F′1,F″4)) Where Conv1(·) is a convolutional layer with a kernel size of 1×1, BN(·) is a batch normalization operation, ReLU(·) is a ReLU activation function, Up(·) is a bilinear interpolation upsampling, Concat(·,·) is a concatenation operation along the channel dimension, and SE(·) is a channel attention module; Step B23: Design the convolution block in the edge perception module; the input is the edge feature map F obtained in step B22 e , after 3×3 convolution, BN layer, ReLU activation function, 3×3 convolution, BN layer, ReLU activation function, 1×1 convolution, the edge mask is finally generated The specific formula is as follows: M e =Conv1(ReLU(BN(Conv3(ReLU(BN(Conv3(F e ))))))) Where Conv3(·) is a convolutional layer with a kernel size of 3×3, BN(·) is a batch normalization operation, ReLU(·) is an activation function, and Conv1(·) is a convolution with a kernel size of 1×1; The specific implementation steps of step B3 are as follows: Step B31, designing an edge enhancement module. First, designing the edge guidance operation in the edge enhancement module; the input is the edge mask M obtained in step B2. e And the feature map obtained in step B1 The input edge mask M e Perform bilinear interpolation downsampling to the same level as the feature map F i Same width and height, get the mask The mask M' e With the feature map F i Multiply and add F i Add, then pass through 3×3 convolution, BN layer, and ReLU activation function in sequence to obtain the edge-guided feature map The specific formula is as follows: M′ e =Down(M e ) Where Down(·) is a bilinear interpolation downsampling operation, is the matrix multiplication operation, is a matrix addition operation, Conv3(·) is a convolutional layer with a kernel size of 3×3, BN(·) is a batch normalization operation, and ReLU(·) is an activation function; Step B32: construct the CBAM attention submodule in the edge enhancement module. The edge enhancement module consists of a serial channel attention SE and a spatial attention SA. The input feature map is the feature map F obtained in step B32. guide , get edge enhancement features The specific formula is as follows: F ee =THIS IS F guide )) Among them, SE(·) is the channel attention module, SA(·) is the spatial attention module; Step B33: Design an edge feature fusion module, with the input being the first-stage feature map extracted in step B1 The edge feature map obtained in step B2 and edge mask The edge mask M e Multiply it with the feature map F4, and then add it to F4 to get the feature map The edge feature map F e After 3×3 convolution, BN layer, and ReLU activation function, the feature map after channel reduction is obtained. F M With F' e The concatenation along the channel dimension is then passed through 3×3 convolution, Swish activation function, 3×3 convolution, Swish activation function, SE module, 3×3 convolution, and finally the feature map F'. e , get the feature map of the first fusion edge information The feature map F' e After the SE module Splicing along the channel dimension, and then passing a 3×3 convolution, to obtain the feature map of the second fusion edge information Finally, the feature map Add to feature map F1 to get the final feature map of fused edge information The specific formula is as follows: F′ e =ReLU(BN(Conv3(F e ))) in is the matrix multiplication operation, is a matrix addition operation, Conv3(·) is a convolutional layer with a kernel size of 3×3, BN(·) is a batch normalization operation, ReLU(·) is an activation function, Swish(·) is a Swish activation function, SE(·) is a channel attention module, and Concat(·,·) is a concatenation operation along the channel dimension; The specific implementation steps of step B4 are as follows: Step B41: First, construct the gated convolution module in the high-order spatial interaction module, and record the feature map input to this module as The input feature map F α Perform layer normalization, recorded as LN1, to obtain the normalized feature map Then After a 1×1 convolution, the channel is expanded to twice its original size to obtain the feature map Will Split into two feature maps along the channel Input q into the depth-separable convolution to get the feature map Then split it into n, where n is the order, feature graph in After multiplying the feature map p0 and the feature map q0, a 1×1 convolution is performed to expand its channel to twice its original value to obtain the first spatial interaction feature map After multiplying the feature map p1 and the feature map q1, a 1×1 convolution is performed to expand its channel to twice its original value to obtain the second spatial interaction feature map Then iterate to the feature map p n-1 With the feature map q n-1 After multiplication, it passes through a convolution layer with the same number of input channels and output channels and a convolution kernel size of 1×1 to obtain the n-th spatial interaction feature map. Finally, the input feature map F α With p n Add to get the intermediate output feature map The specific formula is as follows: Q = DWConv(q) Among them, Split(·) is splitting along the channel dimension, DWConv(·) is a depth-wise separable convolution, and Conv1(·) is a convolutional layer with a convolution kernel size of 1×1. is the matrix multiplication operation, is a matrix addition operation; Step B42: construct a feedforward module in the high-order spatial interaction module, and input the feature map F obtained in step B41 mid , for F mid The layer is normalized, denoted as LN2, and then input into two fully connected layers, denoted as MLP. The output of the two fully connected layers is then combined with the feature map F mid After adding, we get the high-order spatial interaction features The specific formula is as follows: in is a matrix addition operation; Step B43: construct a channel reduction module in the high-order spatial interaction module, and input the F obtained in step B42 h si , F h si After 1×1 convolution, BN layer, and ReLU activation function, we get the high-order spatial interaction feature map after channel reduction. The specific formula is as follows: F′ h si =ReLU(BN(Conv1(F h si ))) Where Conv1(·) is a convolutional layer with a kernel size of 1×1, BN(·) is a batch normalization operation, and ReLU(·) is an activation function; Step B44: First, construct the convolution block in the context aggregation module, and record the context aggregation module input as two feature maps of different scales. and First, the feature map F high Perform bilinear interpolation upsampling and adjust its width and height to match F low The same width and height, and F low Splicing along the channel dimension, and then passing through 1×1 convolution, BN layer, and ReLU activation function in sequence to obtain the feature map Then F cat Divided equally into four feature maps along the channel dimension and Will and After addition, the feature map is obtained by 3×3 convolution, BN layer, and ReLU activation function. Will and After adding the three together, they are sequentially subjected to a 3×3 convolution with a dilation rate of 2, a BN layer, and a ReLU activation function to obtain a feature map. Will and After adding the three together, they are sequentially subjected to a 3×3 convolution with a dilation rate of 3, a BN layer, and a ReLU activation function to obtain a feature map. Will and After addition, the feature map is obtained by sequentially passing through a 3×3 convolution with a dilation rate of 4, a BN layer, and a ReLU activation function. Then and After splicing along the channel dimension, the feature map is obtained by passing through 1×1 convolution, BN layer, and ReLU activation function in sequence. Finally, F cat and F' cat After addition, the context feature map is obtained by 3×3 convolution, BN layer, and ReLU activation function. The specific formula is as follows: F cat =ReLU(BN(Conv1(Concat(F low ,Up(F high ))))) Among them, Up(·) is a bilinear interpolation upsampling operation, Concat(,·) and Concat(·,·,·,·) are concatenation operations along the channel dimension, is a matrix addition operation, Conv3(·) is a convolution layer with a kernel size of 3×3, Conv3 d=i (·) is a 3×3 convolution with dilation rate i, Conv1(·) is a convolution layer with kernel size 1×1, BN(·) is a batch normalization operation, ReLU(·) is an activation function, and Split(·) is an equal split operation along the channel dimension; The specific implementation steps of step B5 are as follows: Step B5: Design a camouflaged target detection network based on edge feature fusion and high-order spatial interaction, including edge perception module, edge feature fusion module, edge enhancement module, high-order spatial interaction module, and context aggregation module; input the original image, and obtain feature maps of four different scales through the backbone network of step B1. and Input F1 and F4 into the edge perception module in step B2 to obtain the edge feature map and edge mask Then, three edge enhancement modules in step B3 are constructed, which are denoted as EEM1, EEM2 and EEM3 respectively, where the input of EEM1 is the fourth stage feature map F4 extracted in step B1 and the edge mask M obtained in step B2. e , the output is edge enhancement feature The input of EEM2 is the third-stage feature map F3 extracted in step B1 and the edge mask M obtained in step B2. e , the output is edge enhancement feature The input of EEM3 is the second-stage feature map F2 extracted in step B1 and the edge mask M obtained in step B2. e , the output is edge enhancement feature Next, an edge feature fusion module in step B3 is constructed, and the input is the first stage feature map F1 extracted in step B1 and the edge feature map F obtained in step B2. e and edge mask M e , the output is a feature map that integrates edge information Then, four high-order spatial interaction modules in step B4 are constructed, denoted as HSIM1, HSIM2, HSIM3 and HSIM4, whose inputs are the feature maps obtained in step B3. and The outputs are and Next, we construct three context aggregation modules in step B4, which are denoted as CAM1, CAM2 and CAM3, respectively. The input of CAM1 is the feature map and Output is context feature map The input of CAM2 is the output of CAM1 and feature map Output is context feature map The input of CAM3 is the output of CAM2 and feature map Output is context feature map For the edge mask M e , and then upsample it by 4 times through bilinear interpolation to obtain the final edge mask M edge ; For context feature map After 1×1 convolution, it is compressed into 1 channel and then upsampled by 16 times by bilinear interpolation to obtain the first stage camouflage target mask. For context feature map After 1×1 convolution, it is compressed into 1 channel and then upsampled by 8 times by bilinear interpolation to obtain the second stage camouflage target mask. For context feature map After 1×1 convolution, it is compressed into 1 channel and then upsampled by 4 times by bilinear interpolation to obtain the final camouflaged target mask. The specific formula is as follows: M edge =Up scale=4 (M e ) Among them, Up scale=4 (·) is a bilinear interpolation upsampling with a multiple of 4. scale=8 (·) is a bilinear interpolation upsampling with a multiple of 8. scale=16 (·) is a bilinear interpolation upsampling with a multiple of 16, Conv1(·) is a convolution layer with a convolution kernel size of 1×1 and an output channel number of 1; The specific implementation steps of step C are as follows: Step C: Design a loss function as a constraint to optimize the camouflaged target detection network based on edge feature fusion and high-order spatial interaction. The specific formula is as follows: Among them, G camo Represents the label image corresponding to the original image I, G edge represents the edge label image corresponding to the original image I, Expressed as the total loss function, represents the weighted binary cross entropy loss, Expressed as weighted intersection-over-union loss, represents the Dice coefficient loss, and λ represents the weight of the loss.

2. The camouflaged target detection method based on edge feature fusion and high-order spatial interaction according to claim 1 is characterized in that: The specific implementation steps of step A are as follows: Step A1, each original image, its corresponding label image and edge label image are combined into an image triplet; Step A2: randomly flip left and right, randomly crop, and randomly rotate each group of image triplets; perform color enhancement on the original image, and adjust the brightness, contrast, saturation, and clarity of the original image by setting random values ​​as parameters; and add random black or white dots as random noise to the label image corresponding to the original image; Step A3: Scale each image in the dataset to an image of the same size of H×W.

3. The camouflaged target detection method based on edge feature fusion and high-order spatial interaction according to claim 1 is characterized in that: The specific implementation steps of step D are as follows: Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N pairs of images; Step D2: Input the original image I, and obtain the edge mask M after passing through the camouflage target detection network based on edge feature fusion and high-order spatial interaction in step B. edge , camouflage target mask and Calculate the loss using the formula in step C Step D3: Calculate the gradient of the parameters in the network using the back propagation method according to the loss, and update the network parameters using the Adam optimization method; Step D4, repeatedly perform steps D1 to D3 in batches until the target loss function of the network converges to the Nash equilibrium, save the network parameters, and obtain a camouflaged target detection model based on edge feature fusion and high-order spatial interaction; for the tested camouflaged target image, the one with the largest resolution among the three camouflaged target masks predicted by the model is used. As the ultimate camouflaged target mask.