A camouflage target detection method based on multi-dimensional polarization information

By constructing a deep convolutional neural network for camouflaged target detection based on multi-dimensional polarization information and using changes in polarization states to distinguish reflection sources, the problem of low camouflaged target detection accuracy in complex scenarios is solved, achieving more efficient camouflaged target detection.

CN119625775BActive Publication Date: 2025-10-10HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411294517.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-10-10
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

Existing camouflaged target detection methods have difficulty in accurately segmenting the target and background in complex scenes, especially in scenes with small differences in external features. The detection accuracy is not high. Traditional methods are time-consuming and labor-intensive, and deep learning-based methods ignore the intrinsic feature information of the scene.

Method used

A deep convolutional neural network for camouflaged target detection based on multi-dimensional polarization information is constructed. Polarization information and structural details are extracted through the polarization dynamic attention module, Res2Net backbone network and receptive field module. Combined with the feature edge fusion module and high and low gating selection module, multi-level features are integrated and the polarization state changes are used to distinguish the reflection source.

Benefits of technology

It improves the precision and accuracy of camouflaged target detection, can effectively detect camouflaged targets in complex and changing environments, solves the detection problem in scenes with small differences in external features, and enhances the network's ability to utilize the intrinsic features of the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625775B_ABST
    Figure CN119625775B_ABST
Patent Text Reader

Abstract

The application discloses a camouflage target detection method based on multi-dimensional polarization information, comprising the following steps: 1, acquiring a polarization image dataset with pixel-level annotation; 2, processing four original polarization angle images, learning the exclusive weight of the feature of each image through a polarization dynamic attention module, simultaneously extracting polarization information and structural details, and fusing with the original image; 3, constructing a deep convolutional neural network based on multi-dimensional polarization information, inputting the fused image, training the deep convolutional network, and obtaining an accurate camouflage target detection result. Under the condition of a reasonable data model, the application generates a scene representation with rich texture and rich edge details by utilizing the correlation and difference of group polarization images, thereby effectively solving the problem that DoLP may be weak or non-existent under certain light conditions, which further affects the effect of camouflage target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision, image processing and analysis, and specifically is a camouflaged target detection method based on multi-dimensional polarization information. Background Art

[0002] Camouflage is a common biological phenomenon in nature. Animals use methods like changing their color and shape to blend in with their surroundings and avoid detection by predators. In addition to natural animal camouflage, artificial camouflage also exists in real life, such as soldiers wearing camouflage during military activities and body painting. Examples include chameleons that change color to suit their habitat, seahorses that match their background color and shape, soldiers who achieve camouflage through clothing and coloring, and body art that blends in with the background. The salient feature of these images is a high degree of similarity between the visual features of the target area and the background. However, current mainstream computer vision detection algorithms, such as salient object detection and general object detection algorithms, mostly detect objects that are significantly different from their surroundings. Therefore, when faced with images with camouflaged features, detection results often suffer from missed and false detections, hindering the further development of related object detection and recognition systems. In this context, camouflaged target detection algorithms have emerged. They are dedicated to detecting objects that are "perfectly" integrated into the surrounding environment. They have high visual feature recognition capabilities and can effectively improve the overall performance of target detection systems. They are a very challenging and indispensable visual task.

[0003] With the rapid development of computer vision algorithms, more and more researchers have begun to focus on the research area of ​​camouflaged object detection, especially emerging deep learning-based camouflaged object detection algorithms, which have achieved considerable research results in recent years. The practical application value of camouflaged object detection is also becoming increasingly prominent. In the medical field, most lesions are highly similar to surrounding normal tissue. Polyp segmentation is a typical example, as the boundary between polyps and the surrounding mucosa is often unclear. Therefore, polyps in lesions can be considered as camouflaged objects, and camouflaged object detection can be used to detect and segment them. In the military, enemies often disguise themselves and their equipment to adapt to environmental characteristics to conceal their forces. Camouflaged object detection can detect camouflaged enemies and military equipment from the background, ensuring the successful completion of military reconnaissance missions. In agronomy, camouflaged object detection can be used to detect camouflaged pests from similar backgrounds, assisting intelligent pest monitoring systems to improve the level and effectiveness of pest monitoring and early warning. Furthermore, by leveraging the powerful visual feature resolution capabilities of camouflaged object detection models, they can also be applied to wildlife detection and protection, search and rescue of people in need in the wild, and other unexplored scenarios.

[0004] Currently, methods for detecting camouflaged objects can be broadly categorized into traditional methods based on handcrafted features and learning-based methods. Traditional methods primarily focus on underlying image features, such as color, texture, gradient, and other prior information, and perform detection through manual feature extraction. This manual extraction method typically requires extensive prior knowledge, is often time-consuming and labor-intensive, and is difficult to apply to large-scale data sets. Before the advent of deep learning models, it was the mainstream approach for camouflaged object detection. While traditional methods are generally effective, detection algorithms based on a single feature are not applicable to all scenarios, and algorithms based on multiple features are challenging to implement. When the texture is similar to the background, color features play a role in detecting camouflage. Conversely, when camouflage is due to color similarity, texture features are effective. When the camouflaged object exhibits good motion, its motion can be leveraged for detection.

[0005] With the continuous advancement of deep neural network technology, researchers are moving beyond the traditional method of extracting handcrafted features to identify camouflaged targets. Due to the powerful learning capabilities of deep neural networks, they can adaptively extract potential connections and semantic expressions from massive amounts of samples, reducing the complex process of manual feature extraction. This has taken camouflaged target detection capabilities to a new level, significantly improving both accuracy and real-time performance. Researchers have constructed a variety of excellent deep neural network models by focusing on various approaches, including novel feature fusion methods, improved network structures, attention mechanisms, dataset expansion, and joint learning tasks. The key to camouflaged target detection technology lies in sufficient training samples, an excellent feature extraction network, efficient feature fusion techniques, and interference noise removal methods. These are essential for achieving excellent camouflaged target detection models. However, current learning-based methods typically target traditional intensity images and can only utilize external feature information within the scene. In challenging scenes where the external feature differences between the target and the background are minimal, these methods struggle to accurately segment the target. Traditional methods, on the other hand, often exploit the inherent feature differences between the target and the environment to highlight camouflaged areas, but both detection efficiency and accuracy need to be improved. Summary of the Invention

[0006] In order to address the shortcomings of the existing technology, the present invention provides a camouflaged target detection method based on multi-dimensional polarization information, so as to effectively detect camouflaged targets in complex scenes, thereby improving the precision and accuracy of camouflaged target detection in complex and changing environments, and effectively overcoming the problem that under certain lighting conditions, the polarization angle is weak or non-existent, which affects the detection effect of camouflaged targets.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0008] The present invention is characterized in that a camouflaged target detection method based on multi-dimensional polarization information is performed according to the following steps:

[0009] Step 1: Data collection and processing;

[0010] Step 1.1: Use a polarization camera to capture the relative polarization angle in the nth scene They are , , , A set of original polarization images , thus obtaining N sets of original polarization images under N scenes; among them, Indicates the relative polarization angle in the nth scene The original polarization image under ;

[0011] Step 1.2: For each set of original polarization images under N scenes, the polarization angle is The original polarization image is annotated to obtain a pixel-level annotated image, which is used as the real camouflage image;

[0012] Step 2: Construct a deep convolutional neural network for camouflaged target detection based on multi-dimensional polarization information, including: encoding module and decoding module;

[0013] Step 2.1, the encoder After processing, the structural feature map of the n-th scene and the polarization fusion feature of the n-th scene are obtained;

[0014] Step 2.2: After the decoder processes the structural feature map and polarization fusion features of the nth scene, the nth scene is obtained. Predicted camouflage map F for each scene n ;

[0015] Step 3: Train a camouflaged target detection model based on multi-dimensional polarization information;

[0016] Step 3.1, based on The real camouflage map and the predicted camouflage map F of the scene n Construct binary cross entropy loss;

[0017] Step 3.2, based on The real camouflage map and the predicted camouflage map F of the scene n The intersection and union construction of loss;

[0018] Step 3.3, binary cross entropy loss and After the losses are weighted, the total loss is obtained;

[0019] Step 3.4: Use the gradient descent method to train the deep convolutional neural network for camouflaged target detection, and calculate the total loss to update the network parameters until the total loss converges. This will result in a camouflaged target detection model based on multidimensional polarization information, which can be used to detect camouflaged targets for any multidimensional polarization image to be predicted.

[0020] The present invention also provides a method for detecting camouflaged targets based on multi-dimensional polarization information, which is characterized in that: the encoder in step 2.1 includes: a polarization dynamic attention module, a Res2Net backbone network, and H receptive field modules;

[0021] Step 2.1.1, polarized attention module After processing, the nth polarization enhanced dimensional information image D is obtained n ;

[0022] Step 2.1.2, the Res2Net backbone network is composed of H-level initial convolutional layers, H-level downsampling convolutional layers and H-level residual blocks;

[0023] Among them, the h-th level initial convolution layer consists of: a convolution layer, a BN layer and a ReLU activation function layer;

[0024] when When D n After being processed by the h-th level initial convolution layer, the h-th level downsampling convolution layer and the h-th level residual block, the h-th level structural feature map is output. ;

[0025] when When Level structure feature map Input the hth level initial convolution layer for processing, and then pass through the hth level The first downsampling convolution layer and the After processing the level-h residual block, the h-th level structural feature map is obtained , so that the H-th level residual block outputs the H-th level structural feature map ;

[0026] Step 2.1.3: Construct each receptive field module including K branches, a standard convolutional layer and a ReLU layer;

[0027] when When the h-th level structural feature graph Input into the hth receptive field module and processed by the 2D convolution of the kth branch to obtain the texture features of the kth branch output by the hth receptive field module ;

[0028] when When the h-th level structural feature graph After inputting several 2D convolution layers of the kth branch in the hth receptive field module for processing, the texture features of the kth branch output by the hth receptive field module are obtained. ;

[0029] The texture features of the K branches output by the h-th receptive field module After cascading, the hth detail feature is output through the processing of the standard convolution layer ;

[0030] The second branch feature output by the h-th receptive field module and After addition, it is input into the ReLU layer for processing, and finally the h-th receptive field module outputs the polarization fusion feature of the n-th scene. ; Thus, the H receptive field modules output the H polarization fusion features of a scene .

[0031] Furthermore, the polarized attention module in step 2.1.1 includes: several 3D convolutional layers, an adaptive average pooling layer and a sigmoid function;

[0032] 3D convolutional layer pairs After 3D convolution processing, the difference features between the polarization angle images in the nth scene are obtained ;

[0033] After processing by the adaptive average pooling layer and the sigmoid function, the weighted features of the group polarization image in the nth scene are obtained. ;

[0034] right and After performing element-by-element multiplication, we get Relative polarization angle in each scene Angle characteristics below , and then After channel cascade operation, the nth polarization enhanced dimensional information image D is obtained n .

[0035] Furthermore, the decoder in step 2.2 includes: a feature edge fusion module, a high and low gating selection module;

[0036] Step 2.2.1, the feature edge fusion module is composed of H fusion blocks, each fusion block includes several convolution layers and an average pooling layer; and As input, it is sent to the feature edge fusion module for processing to obtain the edge fusion feature ;

[0037] when When With the The first scene Level polarization fusion features After channel cascading, The input is sent to the hth fusion block for processing to obtain the hth edge feature map ;

[0038] when When edge feature maps and As an input, it is sent to the hth fusion block for processing to obtain the hth edge feature map , so that the Hth fusion block outputs the Hth edge feature map ;

[0039] Step 2.2.2 The high and low gated selection module consists of a gated selection block and a residual attention block;

[0040] No. Level 1 polarization fusion features of scenes and The input gate selection block is processed to obtain the Cross-layer features ;

[0041] Will After being processed in the residual attention block, the output Predicted camouflage map F for each scene n .

[0042] Furthermore, the gate selection block in step 2.2.2 first selects After the standard convolution operation, batch normalization and ReLU function processing are performed in sequence, the first Refinement Features ;

[0043] right After the transposed convolution operation, the upsampling operation is performed to obtain the first Perceptual features ;

[0044] right and After cascading on the channel, we get Gated Map ;

[0045] right After the transposed convolution operation, After cascading in the channel dimension, we get Cross-layer features .

[0046] The electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the camouflaged target detection method, and the processor is configured to execute the program stored in the memory.

[0047] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the camouflaged target detection method when the computer program is executed by a processor.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention constructs a deep neural network for camouflaged target detection based on multi-dimensional polarization information and uses labeled data to supervise the deep neural network for learning, thereby obtaining a robust polarization camouflaged target detection feature model. This solves the problem of using clues such as color, depth, and background priors in statistical models for model design while ignoring the intrinsic feature information of the scene, and the problem of low detection accuracy in scenes with small external differences.

[0050] 2. Due to the relative differences in the refraction and reflection characteristics of light caused by different materials, surface roughness, and structures, the polarization state of light will change. The camouflaged target detection model based on multi-dimensional polarization information constructed by the present invention distinguishes the reflection source in the scene according to the change in polarization state, and by introducing the intrinsic polarization information characteristics in the scene, it can provide richer information for scene understanding and provide more basis for distinguishing targets, thereby effectively segmenting the camouflaged target from the background.

[0051] 3. The deep neural network based on multi-dimensional polarization information constructed by the present invention takes the multi-dimensional polarization image as the network input, uses the polarization dynamic attention module to extract polarization information and structural details, and makes full use of the polarization information in the scene; thereby solving the problem that most deep learning-based camouflage target methods are difficult to accurately segment the target in scenes where the external features of the target and the background are slightly different.

[0052] 4. This invention effectively explores both low-level and high-level features through multi-level fusion and different aggregation strategies. The feature edge fusion module fuses boundary information with high-level features, enabling the network to focus on global and contextual information at a high level to accurately locate camouflaged objects. Furthermore, a high-low gate selection module is designed to effectively fuse high-level and low-level features, simultaneously filtering out redundant information, reducing network burden, and incorporating useful information, resulting in more accurate prediction of camouflaged targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Flowchart for camouflaged target detection based on multi-dimensional polarization information camouflaged target model;

[0054] Figure 2 Schematic diagram of the deep neural network structure of polarization images with multi-dimensional polarization information;

[0055] Figure 3 Schematic diagram of a single receptive field module;

[0056] Figure 4 This is a diagram of camouflaged target prediction results using the method of the present invention and other camouflaged target detection methods on a polarization dataset. DETAILED DESCRIPTION

[0057] In this embodiment, a camouflaged target detection method based on multi-dimensional polarization information is designed to locate "indistinguishable targets" with highly similar backgrounds. Polarization information is introduced to expand the difference between the object and its surroundings in the camouflaged target detection task, and a deep neural network based on CNN is constructed to process polarization images. By constructing a more refined camouflaged target detection model based on multi-dimensional polarization, the intrinsic feature differences between the target and the background can be fully utilized to better complete the camouflaged target detection task. Specifically, Figure 1 As shown, the method is performed in the following steps:

[0058] Step 1: Data collection and processing;

[0059] Step 1.1: Use a polarization camera to capture the relative polarization angle in the nth scene They are , , , A set of original polarization images , thus obtaining N sets of original polarization images under N scenes; among them, Indicates the relative polarization angle in the nth scene The original polarization image under ;

[0060] In this embodiment, a Lucid Triton focal plane polarization camera is used to shoot a polarization camouflage target detection dataset, which contains The original resolution of each image is 1024×1024.

[0061] Step 1.2: For each set of original polarization images under N scenes, the polarization angle is The original polarization image is annotated to obtain a pixel-level annotated image, which is used as the real camouflage image;

[0062] In this embodiment, the labeling is performed by using Labelme software. The labeled image is to assign a category label V to each pixel in the polarization image. , respectively, represent the color of the pixel: black, white; black represents the background of the pixel, and white represents the target. The polarization camouflage target detection dataset is divided into training and testing, with the training set containing 511 scenes and the test set containing 128 scenes.

[0063] Step 2: Construct a deep convolutional neural network for camouflaged target detection based on multi-dimensional polarization information, including: encoding module and decoding module;

[0064] Step 2.1. Build the encoder, including: polarized dynamic attention module, Res2Net backbone network, and H receptive field modules;

[0065] Step 2.1.1. The polarized attention module includes: several 3D convolutional layers, a 3D global average pooling layer and a sigmoid function;

[0066] After being processed by the polarization attention module, the nth polarization-enhanced dimensional information image D is obtained. n ;

[0067] In this embodiment, the 3D convolution layer After 3D convolution processing, the difference features between the polarization angle images in the nth scene are obtained ;

[0068] After processing by the adaptive average pooling layer and the sigmoid function, the weighted features of the group polarization image in the nth scene are obtained. ;

[0069] right and After performing element-by-element multiplication, the relative polarization angle in the nth scene is obtained. Angle characteristics below , and then After channel cascade operation, the nth polarization enhanced dimensional information image D is obtained n .

[0070] In the specific implementation, in order to balance the training efficiency and accuracy of the model, The size of the four original polarization images is downsampled to 416×416,

[0071] In this example, the number of samples in a batch is set to 12, each sample has 3 channels, and both the height and width are 416. This means that each batch of training inputs consists of 48 multi-dimensional polarization images of size 416×416, equivalent to four sets of original polarization images of size 352×352. Ultimately, the multi-dimensional polarization image size input to the polarization dynamic attention module for each training batch is 12×3×416×416, and the feature size output by the polarization dynamic attention module is 12×3×352×352.

[0072] Step 2.1.2, the Res2Net backbone network is composed of H-level initial convolutional layers, H-level downsampling convolutional layers and H-level residual blocks;

[0073] Among them, the h-th level initial convolution layer consists of: a convolution layer, a BN layer and a ReLU activation function layer;

[0074] when When D n After being processed by the h-th level initial convolution layer, the h-th level downsampling convolution layer and the h-th level residual block, the h-th level structural feature map is output. ;

[0075] In this embodiment, the first level structural feature map The width, height, and number of channels are 352, 352, and 256.

[0076] when When Level structure feature map Input the hth level initial convolution layer for processing, and then pass through the hth level The first downsampling convolution layer and the After processing the level-h residual block, the h-th level structural feature map is obtained ; Thus, the H-th level residual block outputs the H-th level structural feature map ;

[0077] In this embodiment, the second-level structural feature map The width, height, and number of channels are 176, 176, and 512; the third-level structural feature map The width, height, and number of channels are 88, 88, and 1024; the fourth-level structural feature map The width, height, and number of channels are 44, 44, and 2048.

[0078] Step 2.1.3: Construct each receptive field module including K branches, a standard convolutional layer and a ReLU layer;

[0079] when When the h-th level structural feature graph Input into the hth receptive field module and processed by the 2D convolution of the kth branch to obtain the texture features of the kth branch output by the hth receptive field module ;

[0080] when When the h-th level structural feature graph After inputting several 2D convolution layers of the kth branch in the hth receptive field module for processing, the texture features of the kth branch output by the hth receptive field module are obtained. ;

[0081] The texture features of the K branches output by the h-th receptive field module After cascading and processing through the standard convolution layer, the hth detail feature is output ;

[0082] The second branch feature output by the h-th receptive field module and After addition, it is input into the ReLU layer for processing, and finally the h-th receptive field module outputs the polarization fusion feature of the n-th scene. ; Thus, the H receptive field modules output the H polarization fusion features of a scene ;

[0083] In this example, K=5. The input of the h-th receptive field module is the h-th level structure feature map .

[0084] like Figure 3 As shown, when k=1, 2, the h-th level structural feature map Input a 1×1 standard 2D convolution (consisting of a 2D convolution layer and a BN layer) of the kth branch in the hth receptive field module for processing. The receptive field module The branch and The texture features of the branches are and ;

[0085] when When the h-th level structural feature graph Input the hth receptive field module The branches are processed by several standard convolutional layers to obtain the The receptive field module Texture features output by each branch . No. The branch consists of a convolution kernel of size A standard convolutional layer with a dilation rate of of Standard convolutional layers. And, using Convolutional layers and Combination of convolutional layers instead Convolutional layer to reduce computational overhead. The output texture features are recorded as ;

[0086] Will After cascading, a size of The standard convolution layer outputs the hth detail feature ; Texture features of the output of the second branch of the receptive field module and After adding, it is input into a ReLU layer for processing, and the final The receptive field module outputs the A polarization fusion feature for each scene

[0087] The input of the first receptive field module is the dimension information image D n and the structural feature map output by the first layer of the backbone network After 2 times downsampling, the two are cascaded to obtain the polarization fusion feature with the final output size of 352×352×32 ; The input of the second receptive field module is the structural feature map of the second layer of the backbone network , the final output size is 176×176×32 polarization fusion features ; The input of the third receptive field module is the structural feature map of the third layer of the backbone network , the final output size is 88×88×32 polarization fusion features ; The input of the fourth receptive field module is the structural feature map of the fourth layer of the backbone network , the final output size is 44×44×32 polarization fusion features .

[0088] Step 2.2: Construct a decoder, including a feature edge fusion module and a high and low gating selection module.

[0089] Step 2.2.1, the feature edge fusion module is composed of H fusion blocks, each fusion block includes several convolution layers and an average pooling layer; and As input, it is fed into the feature edge fusion module to obtain the edge fusion feature ;

[0090] when When With the The first scene Level polarization fusion features After channel cascading, The input is sent to the hth fusion block for processing to obtain the hth edge feature map ;

[0091] when When edge feature maps and As an input, it is sent to the hth fusion block for processing to obtain the hth edge feature map ; Thus, the Hth fusion block outputs the Hth edge feature map ;

[0092] Step 2.2.2 The high and low gated selection module consists of a gated selection block and a residual attention block;

[0093] No. Level 1 polarization fusion features of scenes and The input gate selection block is processed to obtain the Cross-layer features In this embodiment, the first level polarization fusion feature The number of channels is 32, and the edge fusion feature The number of channels is 32;

[0094] In the specific implementation, the gate selection block first After the standard convolution operation, batch normalization and ReLU function processing are performed in sequence, the first Refinement Features ; In this embodiment, the refinement feature The dimensions are 44×44×192;

[0095] right After the transposed convolution operation, the upsampling operation is performed to obtain the first Perceptual features In this embodiment, the perception feature The dimensions are 88×88×1;

[0096] n

[0097]

[0098]

[0099] n

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107] ​​​​​​​​​​​​​​​​​​​​​​​​​​In this embodiment, the multi-dimensional polarization images of 2044 scenes after the polarization dataset data enhancement and their corresponding real camouflage images are used for training, and the F output by the output module is converted to n Weighted binary cross entropy and weighted with the real camouflage image The loss calculation yields five training losses, which are summed to form a total loss. This total loss, combined with a gradient descent algorithm, guides network training, resulting in a multi-dimensional network feature model for detecting camouflaged targets using polarization images. Using multi-dimensional polarization images from 128 test scenes in the polarization camouflaged target detection dataset as input, the camouflaged target monitoring model, based on multi-dimensional polarization information, calculates predicted camouflaged target images. These images are then compared with the true camouflaged target images of the corresponding scenes to calculate detection accuracy.

[0108] Table 1 shows the comparison results of the camouflaged target detection method based on multi-dimensional polarization information of the present invention with other current camouflaged target detection methods using the test set of the polarization camouflaged target detection dataset, using "S-measure", "E-measure", "MAE" and "F-measure" as evaluation indicators. "S-measure" can be used to reflect structural characteristics and can be used to evaluate the structural similarity between the prediction result and the real camouflaged target. The closer its value is to 1, the better the camouflaged target detection effect. "E-measure" evaluates the overall and local accuracy and quality of the prediction results in the binary segmentation task by combining pixel-level matching with image-level statistics. The closer its value is to 1, the better the camouflaged target detection effect. "F-measure" is a weighted harmonic mean based on the precision and recall rate. The closer its value is to 1, the better the camouflaged target detection effect. "MAE" is an evaluation indicator widely used in salient target detection. It is used to compare the pixel-by-pixel difference between the binary prediction result and the true value image. The closer its value is to 0, the better the camouflaged target detection effect. "BASNet" stands for Boundary-Aware Segmentation Network; "PraNet" stands for Parallel Reverse Attention Network; "MFFN" stands for Multi-view Feature Fusion Network; SINet-V2 stands for the second version of Search Identification Network; "ZoomNet" comes from the paper Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object Detection; "C 2"FNet-V2" comes from the paper CamouflagedObject Detection via Context-aware Cross-level Fusion; "BGNet" stands for Boundary-Guided Network. According to the quantitative analysis in Table 1, it can be seen that the method proposed in this paper has achieved good results in all evaluation indicators.

[0109] Table 1

[0110]

[0111] Figure 4The results of the multi-dimensional camouflaged target detection method designed for this paper and other current camouflaged target detection methods. Among them, "POL4Net" stands for Polarization four-dimensional network, which is a polarization image camouflage target detection method based on a multi-dimensional network of the present invention; "PGSNet" stands for Polarization Glass Segmentation Network, which proposes a new learning-based glass segmentation network, which utilizes the three-color (RGB) intensity and three-color linear polarization cues of a single photo, uses a novel global guidance and multi-scale self-attention module to dynamically fuse and weight the three-color and polarization cues, and uses global cross-domain context information to achieve robust segmentation; "CMX" stands for Cross-Modal Fusion for RGB-X, which is a unified fusion framework that uses the features of one modality to correct the features of another modality to calibrate the bimodal features. A feature fusion module is deployed by correcting the feature pairs to perform a full exchange of long-range context before mixing; "PopNet" stands for Popping Network, which uses the prior knowledge of objects in 3D to pop-out and uses deep reasoning models for object segmentation. This prior knowledge can assist in reasoning about objects in 3D space and uses 3D information for positioning by adjusting the inferred depth map; "DGNet" stands for Deep GradientNetwork, a new deep framework for camouflaged target detection using target gradient supervision - Deep Gradient Network, decouples the task into two connected branches, namely context and texture encoders; "FAPNet" stands for Feature Aggregation and Propagation Network, a new feature aggregation and propagation network, which explicitly models boundary features through a boundary guidance module to provide boundary enhancement features, represents the multi-scale information of each layer through a multi-scale feature aggregation module, and obtains aggregated feature representations, and integrates features of adjacent layers using a cross-layer fusion and propagation module; "MFFN" stands for Multi-view Feature Fusion Network, a multi-view feature fusion network that imitates the human behavior of searching for unclear objects in images, that is, observing from multiple angles, distances, and perspectives, and capturing key boundary and semantic information by comparing and fusing the extracted multi-view features; "PraNet" stands for Parallel Reverse Attention Network, a parallel reverse attention network that first aggregates features in high layers using parallel partial decoders, and then mines boundary cues through a reverse attention module to establish the relationship between regions and boundary cues.

Claims

1. A camouflaged target detection method based on multi-dimensional polarization information, characterized in that: The steps are as follows: Step 1: Data collection and processing; Step 1.1: Use a polarization camera to capture the relative polarization angle in the nth scene They are , , , A set of original polarization images , thus obtaining N sets of original polarization images under N scenes; among them, Indicates the relative polarization angle in the nth scene The original polarization image under ; Step 1.2: For each set of original polarization images under N scenes, the polarization angle is The original polarization image is annotated to obtain a pixel-level annotated image, which is used as the real camouflage image; Step 2. Construct a deep convolutional neural network for camouflage target detection based on multi-dimensional polarization information, including: an encoder and a decoder; wherein the encoder includes: a polarization dynamic attention module, a Res2Net backbone network, and H receptive field modules; the polarization attention module includes: several 3D convolution layers, an adaptive average pooling layer, and a sigmoid function; the Res2Net backbone network is composed of H-level initial convolution layers, H-level downsampling convolution layers, and H-level residual blocks; each receptive field module includes K branches, a standard convolution layer, and a ReLU layer; the decoder includes: a feature edge fusion module and a high and low gating selection module, wherein the feature edge fusion module is composed of H fusion blocks, each fusion block includes several convolution layers and an average pooling layer; the high and low gating selection module includes a gate selection block and a residual attention block; Step 2.1, the encoder After processing, the structural feature map of the n-th scene and the polarization fusion feature of the n-th scene are obtained; Step 2.2: After the decoder processes the structural feature map and polarization fusion features of the nth scene, the nth scene is obtained. Predicted camouflage map F for each scene n ; Step 3: Train a camouflaged target detection model based on multi-dimensional polarization information; Step 3.1, based on The real camouflage map and the predicted camouflage map F of the scene n Construct binary cross entropy loss; Step 3.2, based on The real camouflage map and the predicted camouflage map F of the scene n The intersection and union construction of loss; Step 3.3, binary cross entropy loss and After the losses are weighted, the total loss is obtained; Step 3.4: Use the gradient descent method to train the deep convolutional neural network for camouflaged target detection, and calculate the total loss to update the network parameters until the total loss converges. This will result in a camouflaged target detection model based on multidimensional polarization information, which can be used to detect camouflaged targets for any multidimensional polarization image to be predicted.

2. A camouflaged target detection method based on multi-dimensional polarization information according to claim 1, characterized in that: The step 2.1 includes: Step 2.1.1, polarized attention module After processing, the nth polarization enhanced dimensional information image D is obtained n ; Step 2.1.2: The h-th initial convolutional layer consists of: a convolutional layer, a batch normalization layer, and a ReLU activation function layer. when When D n After being processed by the h-th level initial convolution layer, the h-th level downsampling convolution layer and the h-th level residual block, the h-th level structural feature map is output. ; when When Level structure feature map Input the hth level initial convolution layer for processing, and then pass through the hth level The first downsampling convolution layer and the After processing the level-h residual block, the h-th level structural feature map is obtained , so that the H-th level residual block outputs the H-th level structural feature map ; Step 2.1.3, when When the h-th level structural feature graph Input into the hth receptive field module and processed by the 2D convolution of the kth branch to obtain the texture features of the kth branch output by the hth receptive field module ; when When the h-th level structural feature graph After inputting several 2D convolution layers of the kth branch in the hth receptive field module for processing, the texture features of the kth branch output by the hth receptive field module are obtained. ; The texture features of the K branches output by the h-th receptive field module After cascading, the hth detail feature is output through the processing of the standard convolution layer ; The second branch feature output by the h-th receptive field module and After addition, it is input into the ReLU layer for processing, and finally the h-th receptive field module outputs the polarization fusion feature of the n-th scene. ; Thus, the H receptive field modules output the H polarization fusion features of each scene .

3. A camouflaged target detection method based on multi-dimensional polarization information according to claim 2, characterized in that: The step 2.1.1 includes: 3D convolutional layer pairs After 3D convolution processing, the difference features between the polarization angle images in the nth scene are obtained ; After processing by the adaptive average pooling layer and the sigmoid function, the weighted features of the group polarization image in the nth scene are obtained. ; right and After performing element-by-element multiplication, we get Relative polarization angle in each scene Angle characteristics below , and then After channel cascade operation, the nth polarization enhanced dimensional information image D is obtained n .

4. A camouflaged target detection method based on multi-dimensional polarization information according to claim 3, characterized in that: The step 2.2 includes: Step 2.2.1, and As input, it is sent to the feature edge fusion module for processing to obtain the edge fusion feature ; when When With the The first scene Level polarization fusion features After channel cascading, The input is sent to the hth fusion block for processing to obtain the hth edge feature map ; when When edge feature maps and As an input, it is sent to the hth fusion block for processing to obtain the hth edge feature map , so that the Hth fusion block outputs the Hth edge feature map ; Step 2.2.2, Level 1 polarization fusion features of scenes and The input gate selection block is processed to obtain the Cross-layer features ; Will After being processed in the residual attention block, the output The predicted camouflage map F of the scene n .

5. A camouflaged target detection method based on multi-dimensional polarization information according to claim 4, characterized in that: The gate selection block in step 2.2.2 is first After the standard convolution operation, batch normalization and ReLU function processing are performed in sequence, the first Refinement Features ; right After the transposed convolution operation, the upsampling operation is performed to obtain the first Perceptual features ; right and After cascading on the channel, we get Gated Map ; right After the transposed convolution operation, After cascading in the channel dimension, we get Cross-layer features .

6. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the disguised target detection method according to any one of claims 1 to 5, and the processor is configured to execute the program stored in the memory.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the disguised target detection method according to any one of claims 1 to 5 are executed.

Citation Information

Patent Citations

  • Camouflage target detection method based on polarization image clues and application thereof

    CN115620049A

  • Camouflage target segmentation method and system based on light intensity and polarization clues

    CN115861608A