Camouflage target detection method based on dynamic Top-k feature selection and cross tensor decomposition

By adopting dynamic Top-k feature selection and cross-tenster decomposition technology in camouflage object detection, combined with global and local feature extraction, the problems of insufficient global context information capture and feature redundancy in traditional methods are solved, and more efficient camouflage object detection is achieved.

CN119942090AActive Publication Date: 2025-05-06NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510367885.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-05-06
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Traditional camouflage object detection methods are difficult to effectively capture the global context information and long-distance dependencies in the image, resulting in incomplete detection of camouflage object, and convolution operations may lose key details information, resulting in low feature fusion efficiency.

Method used

The method based on dynamic Top-k feature selection and cross-tenster decomposition is adopted, and global and local features are extracted through global perception modules and local optimization modules, and feature redundancy information is reduced through cross-tenster decomposition. Finally, cross-layer feature interaction and adaptive weighting are performed in the hierarchical fusion module and the hybrid weighting decoder.

Benefits of technology

Effectively capture the global contextual relationship and local details of the image, reduce feature redundancy learning interference, enhance feature representation ability, and improve the accuracy and robustness of camouflage object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942090A_ABST
    Figure CN119942090A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and target detection, in particular to a camouflage target detection method based on dynamic Top-k feature selection and cross tensor decomposition, which comprises the following steps: inputting a camouflage target image into an encoder network to extract multi-level features, and generating feature maps of different scales; the input global perception module and the local optimization module are used for extracting features of different scales, and global features and local features are output; performing cross tensor decomposition on the global features and the local features, and generating complementary features through low-rank factor matrix expansion and cross combination; inputting the complementary global features, the complementary local features and the previous-layer fusion features into a hierarchical fusion module; the features are input into a hybrid weighted decoder, reverse optimization is carried out in combination with a channel attention mechanism of a DSE module, and finally a camouflage target detection result is output; and carrying out multi-level supervision on the output segmentation image and the final prediction image, and flexibly adjusting the degree of dependence on different features by jointly optimizing model parameters through a loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and target detection, and in particular to a camouflaged target detection method based on dynamic Top-k feature selection and cross tensor decomposition. Background Art

[0002] In nature, there are many animals that are good at "camouflaging" themselves, such as leaf butterflies and stick insects. They blend in with nature to better avoid natural enemies. Animals "camouflage" generally by changing their own color and shape to blend into the background, making it difficult for natural enemies to find them; or by creating a pseudo-edge effect through special patterns, making it difficult to expose the real outline. Artificial "camouflage" is similar to animals, such as hiding oneself through body painting, camouflage clothing, etc.

[0003] Early traditional camouflaged object detection (COD) methods usually rely on hand-crafted features or low-level visual priors (e.g., color, texture, and intensity) to detect camouflaged objects. However, these features or priors may not fully capture the complex structures and contents within the image. In particular, objects occluded by the CAM in natural images often have extremely high similarity with the background, resulting in unsatisfactory segmentation results of traditional strategies.

[0004] Later, due to the availability of large-scale datasets, many deep learning-based COD methods were proposed to solve this difficult segmentation problem. Most of these methods aim to aggregate multi-scale features captured by various convolution operations to gradually distinguish objects from backgrounds. They achieve better performance by integrating spatial information containing different receptive fields. However, the receptive field of convolution operations is often limited, and the extracted clues are often local, making it difficult to establish a global relationship model between all pixels. In particular, when the receptive field of the model is limited to a local perspective and lacks a global perspective, the predicted camouflaged object may be an incomplete whole. In addition, camouflaged objects in images are usually of variable shapes and non-fixed scales, requiring a large amount of contextual information to understand the image content. It is particularly important to model global contextual information in camouflaged object detection and capture long-distance dependencies between pixels. However, traditional convolution operations may lose some key detail information during feature extraction, and the information transmission efficiency between features at different levels is limited, making it difficult to fully utilize the association between features. In the subsequent fusion process, there is often a problem of inability to accurately and efficiently perform feature fusion due to redundant learning of features between different levels. Summary of the invention

[0005] In view of the above-mentioned problems, the present invention provides a camouflaged target detection method based on dynamic Top-k feature selection and cross tensor decomposition, which models the overall contextual relationship of the camouflaged image through the Mamba block in the global perception module and extracts global features. The linear complexity long sequence modeling of the global Mamba can efficiently capture the semantic consistency across regions. The Mamba block in the local optimization module focuses on local details to avoid the loss of fine-grained information caused by global pooling or downsampling. While retaining the high-resolution feature map, the local Mamba can model and enhance the expression ability of fine-grained features. By introducing the dynamic Top-k feature selection mechanism in both the global and local Mamba, the key features are dynamically screened to filter out background features irrelevant to the target, enhance the response of the target area, and suppress irrelevant background noise. By cross-tensor decomposition of global features and local features, the learning interference of redundant information of features between different levels in the subsequent fusion process is effectively reduced. In the hierarchical fusion module, the global, local and previous layer fusion features are processed and cross-fused, which can more comprehensively capture the relationship between features of different levels and types, enhance the feature representation ability, and take different paths. Feature interaction also avoids the loss of important information. In the hybrid weighted decoder, different information is integrated through cross-layer aggregation, dynamic weighting and inverse optimization, which can mine and utilize potentially important information with different characteristics. The dynamic weight mechanism is introduced in the decoder so that features at different levels can be adaptively weighted according to their importance during the fusion process, and the degree of dependence on different features can be flexibly adjusted.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] A camouflaged target detection method based on dynamic Top-k feature selection and cross tensor decomposition, characterized in that the method comprises:

[0008] S1, disguise the target image I∈R H×W×C Input encoder network to extract multi-level features E i,i=1,2,3,4,5 , the generated resolution is Feature maps of different scales;

[0009] S2, input the E2-E5 layer features into the global perception module and the local optimization module respectively to extract features of different scales, screen out global key features and local sensitive features, and output global features G2-G5 and local features L2-L5;

[0010] S3, for the global feature G at the same level i and local feature L i Perform cross tensor decomposition to generate complementary features through low-rank factor matrix expansion and cross merging;

[0011] S4, the complementary global features, local features and the previous layer fusion features Fi-1 Input hierarchical fusion module, which realizes cross-level feature interaction through gated convolution and dynamic weighting mechanism, and outputs fused features F2-F5;

[0012] S5, input the F2-F5 features into the hybrid weighted decoder, combine the channel attention mechanism of the DSE module for reverse optimization, generate and upsample the prediction map layer by layer, and finally output the camouflaged target detection result. The input of the last two decoders not only includes the fusion features of the current level, but also the feature maps output by the first two levels of decoders;

[0013] S6. Perform multi-level supervision on the segmentation map and the final prediction map output by the hybrid weighted decoder, and jointly optimize the model parameters through the loss function.

[0014] Preferably, in step S1, the encoder backbone network uses resnet-50 to extract five initial features, wherein since the resolution of the first-layer feature E1 is greater than the resolution threshold and there is background noise, E1 is discarded and E2-E5 is used to complete the subsequent COD task.

[0015] Preferably, two modules are designed in step S2 to extract local and global feature information, and the spatial local information in the initial feature is increased by the local refinement module, which includes multiple local visual mamba blocks (LVMB) connected together. In LVMB, a dynamic Top-k feature selection mechanism is introduced to dynamically screen the key features in the local features to filter out the background local features that are irrelevant to the target and enhance the sensitivity of the model to the key areas. Local features L2-L5 are output through the LVMB module. The global perception module also includes multiple global visual mamba blocks (GVMB) connected together, and GVMB is used to obtain all pixel relationships from a global perspective. Similar to LVMB, a dynamic Top-k feature selection mechanism is also introduced in GVMB to screen the most critical global features for the detection task and reduce the amount of redundant feature processing. Global features G2-G5 are output through the GVMB module.

[0016] Preferably, local Top-k feature selection (LTTS) and global Top-k feature selection (GTTS) make full use of sparsity by selecting the top k tags with the highest relevance to the query, thereby capturing the most critical disguised target information for extraction. k is dynamically set to a series of values ​​that satisfy ki = (i+1) / (n+2), where i is the current level index, n is the total number of levels, and the difference rate of k values ​​at different levels is not less than 15%. For example, k = 1 / 2, only the 50% of elements with the highest scores can be retained and activated, while the remaining 50% of the elements are masked to 0;

[0017] For the local visual mamba block, the LTTS module is introduced as the feature extractor. Compared with the general visual mamba block feature extractor, the local visual mamba block contains an expansion operator, which can effectively ensure the spatial proximity of adjacent markers in a two-dimensional or three-dimensional array. The window size R=3 is set using the kernel size commonly used in the convolution layer. Since the number of channels of LTTS becomes C', a linear layer is inserted after the state space module (SSM) to project the markers into the original C-dimensional space. For the global visual mamba block, the GTTS module is introduced as the feature extractor. After the deep convolution layer DWC, the activation function SiLU and flattening, it is combined with the Top-k feature selection mechanism to generate features with GRF, and then forwarded to the next SSM. The number of channels in the output remains unchanged, while the sequence length increases, so the linear layer after the SSM is cancelled.

[0018] Preferably, in step S3, the same-level features L i and G i By decomposing the convolutional layer, the local information is selectively integrated into the global information to form complementary global features, and the global information is selectively integrated into the local information to form complementary local features, which solves the problem of redundant learning of features between different layers in the process of camouflaged target detection. i and local feature L i Construct low-rank factor matrices A, a∈R respectively TS×r and Then expand the low-rank factor matrix dimensions to A', a'∈R TS×(r+Δr) and Among them, Δr = 16, S and T represent the number of input channels of the layer and the number of kernels learned in the convolution layer, respectively, D1 and D2 represent the width and height of the convolution kernel space window, respectively, and r represents the rank of the original weight matrix before decomposition; finally, cross-merge to generate enhanced weight matrices M' = A' × b' and m' = a' × B', where the parameter increment ΔP_fac = Δr(TS+D1D2), and the convolution filters K, k can be derived from M', m'.

[0019] Preferably, in order to explicitly capture L i and G i Complementary clues between layers, we maximize the L2 or L1 distance between the feature maps output from each branch of the layer and use it as an additional term in the task training objective. This can be achieved by the following loss Lc: Lc = -||K*x G -ΔK*x G -k*x L -Δk*x L || p ;

[0020] Where p = {1, 2}, * represents convolution, and x represents a vector. The dimensions of K, k and ΔK, Δk are the same.

[0021] Preferably, the hierarchical fusion module in step S4 contains three branches, which can adaptively fuse local features, global representations and previous layer fusion semantic information according to input features. 1. Global weight generation path: GlobalAvgPool→1x1 convolution→BN→ReLU→1x1 convolution→Sigmoid; 2. Local weight generation path: double 1×1 convolution layer series structure and BN batch normalization and GELU loss function; 3. Previous feature weight generation path: double 1×1 convolution layer series structure and batch normalization and ReLU loss function. The outputs of the three paths are spliced, and then redundant information is filtered through gated convolution, and residual connection is performed to improve feature discrimination ability.

[0022] Furthermore, a hybrid weighted decoder is proposed in step S5, which integrates different information through cross-layer aggregation, dynamic weighting and reverse optimization, and can mine and utilize potential important information with different characteristics. The previous fusion features and the segmentation features output by the adjacent decoder are concatenated, convolved, reshaped and reversed, and then the feature map obtained by the addition and multiplication operation is added to the feature map of the previous segmentation feature map after the dynamic weight allocation of the DSE module to finally output the segmentation map D i .

[0023] The DSE module of the hybrid weighted decoder includes a dual attention branch and a weight fusion part. Compared with the traditional SE module, the strong channel attention of DSE enables the model to capture the dependency between feature channels from the perspectives of average pooling and maximum pooling, and suppress irrelevant channels. The dynamic weight mechanism is introduced in the decoder, so that features at different levels can be adaptively weighted according to their importance during the fusion process, and the degree of dependence on different features can be flexibly adjusted.

[0024] Preferably, in step S6, the segmentation map D6 is replaced by the output of the previous stage fusion module, and the output of each stage decoder is also an input of the next stage decoder. D6 and the four decoders are connected adjacently to generate the camouflaged target segmentation map D i,i=2,3,4,5,6 The loss is calculated with the ground truth graph GT of the training set. The selected structured loss function is the weighted binary cross entropy loss. Weighted Intersection-over-Union Loss And tensor loss Lc. To handle multi-scale outputs, all outputs D k=2,3,4,5,6 Upsampling is performed to match the resolution of the segmentation truth map GT. The formula of the total loss function L of the disguised target detection model is as follows:

[0025]

[0026] Where k represents different network layers; D i Represents the segmentation map of the disguised target at different layers; GT represents the true value label of the disguised target.

[0027] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0028] The present invention uses a global Mamba linear complexity long sequence modeling to camouflage the overall contextual relationship of the image in the global perception module, and efficiently captures the semantic consistency across regions; in the local optimization module, local Mamba is used to focus on local details, and local anomalies are captured through fine-grained comparison. While retaining high-resolution feature maps, the fine-grained feature expression ability can also be enhanced by modeling; a dynamic Top-k feature selection mechanism is introduced in the global and local Mamba to dynamically screen key features, which can filter out background features irrelevant to the target, enhance the response of the target area, and suppress irrelevant background noise; in the cross tensor decomposition of global features and local features, the inconsistency in the subsequent fusion process is effectively reduced. Learning interference of redundant information of features at the same level; processing and cross-fusion of global, local and previous layer fusion features in the hierarchical fusion module to more comprehensively capture the relationship between features of different levels and types and enhance feature representation capabilities; in the hybrid weighted decoder, different information is integrated through cross-layer aggregation, dynamic weighting and reverse optimization, which can mine and utilize potentially important information with different characteristics. A dynamic weight mechanism is introduced in the decoder so that features at different levels can be adaptively weighted according to their importance during the fusion process, and the degree of dependence on different features can be flexibly adjusted, thereby enhancing the adaptability and robustness of the model and improving the accuracy of camouflaged target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0030] Figure 1 It is an overall implementation flow chart of the camouflaged target detection method based on dynamic Top-k feature selection and cross tensor decomposition of the present invention;

[0031] Figure 2 It is a flowchart for implementing the local visual mamba block LVMB according to an embodiment of the present invention;

[0032] Figure 3 This is a flowchart for implementing the global visual mamba block GVMB according to an embodiment of the present invention;

[0033] Figure 4 This is a flowchart for implementing the cross tensor decomposition convolutional layer according to an embodiment of the present invention;

[0034] Figure 5This is a flowchart of the implementation of the hierarchical fusion module described in an embodiment of the present invention;

[0035] Figure 6 A flow chart of the implementation of the hybrid weighted decoder according to an embodiment of the present invention;

[0036] Figure 7 This is a flow chart for implementing the dynamic compression and excitation DSE according to an embodiment of the present invention;

[0037] Figure 8 It is a schematic diagram for qualitative comparison between six mainstream camouflage target segmentation models and the method of the present invention in comparative experiments of embodiments of the present invention. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] See also Figure 1-Figure 8 , the present invention provides a technical solution:

[0040] Embodiment 1:

[0041] Figure 1 It is an overall implementation flow chart of a camouflaged target detection method based on dynamic Top-k feature selection and cross tensor decomposition of the present invention;

[0042] like Figure 1 As shown, the camouflaged target detection network model described in the present invention includes an encoder, a global perception module, a local optimization module, a cross tensor decomposition, a hierarchical fusion module and a hybrid weighted decoder.

[0043] The backbone network uses the restnet-50 model in the resnet series to extract multi-scale features E containing disguised target images. i,i=1,2,3,4,5 ,Since the camouflaged target detection algorithms in recent years mostly use resnet-50, in order to verify the effectiveness of the module through comparative experiments, resnet-50 is also selected as the backbone network;

[0044] The global perception module consists of multiple global visual Mamba blocks. From high to low, the outputs of adjacent Mamba blocks are spliced ​​together to output the global feature G. i,i=2,3,4,5 The linear complexity long sequence modeling of Global Mamba can efficiently capture the semantic consistency across regions.

[0045] The local optimization module consists of multiple local visual mamba blocks, which are connected in the same way as the global feature extraction structure and output the local feature L i,i=2,3,4,5 The local Mamba focuses on local details and can model and enhance the fine-grained feature expression capability while retaining the high-resolution feature map. At the same time, a dynamic Top-k feature selection mechanism is introduced in the Mamba block feature extraction stage to dynamically screen key features, filter out background features irrelevant to the target, enhance the target area response, and suppress irrelevant background noise. Then the output global feature G i and local feature L i The cross tensor decomposition is performed to effectively reduce the learning interference of redundant information of features at different levels in the subsequent fusion process. The hierarchical fusion module contains the global, local and previous layer fusion features F i+1,i=2,3,4 The three branches and four fusion modules are adjacent and interconnected, which can more comprehensively capture the relationship between features of different levels and types and enhance the feature representation capability.

[0046] The hybrid weighted decoder obtains the feature F i,i=2,3,4,5 Decode and generate prediction graph D i,i=2,3,4,5 , where D6 is replaced by the output of the first level fusion module. In the decoder, a dynamic weight mechanism is introduced through the DSE module, so that the features of different levels can be adaptively weighted according to their importance during the fusion process, and the segmentation maps D of different features can be flexibly adjusted. i+1 ,D i+2,i=2,3,4 The hybrid weighted decoder integrates different information through cross-layer aggregation, dynamic weighting and inverse optimization, which can mine and utilize potential important information with different characteristics to obtain a more refined camouflaged target segmentation map.

[0047] The camouflaged target detection method of the present invention includes the following contents:

[0048] Step 1: disguise the target image I∈R H×W×C Input encoder network to extract multi-level features E i,i=1,2,3,4,5 , the generated resolution is Feature maps of different scales;

[0049] In this implementation, the encoder backbone network uses resnet-50 to extract five initial features. Since the resolution of the first-layer feature E1 is large and there is background noise, E1 is discarded and E i,i=2,3,4,5 Complete subsequent COD tasks.

[0050] Step 2: E i,i=2,3,4,5 The features are input into the global perception module and the local optimization module respectively to extract features of different scales, screen out the global key features and local sensitive features, and output the global feature G i,i=2,3,4,5 and local feature L i,i=2,3,4,5 ;

[0051] In this implementation, two modules are designed to extract local and global feature information. The local refinement module is used to increase the spatial local information in the initial features. It contains multiple local visual mamba blocks (LVMBs) connected together. In LVMB, a dynamic Top-k feature selection mechanism is introduced to dynamically filter out the key features in the local features to filter out the background local features that are irrelevant to the target and improve the model's sensitivity to the key areas. The local feature L is output by the LVMB module. i,i=2,3,4,5 The global perception module also includes multiple global visual mamba blocks (GVMB) connected together, and GVMB is used to obtain all pixel relationships from a global perspective. Similar to LVMB, GVMB also introduces a dynamic Top-k feature selection mechanism to screen the most critical global features for the detection task and reduce the amount of redundant feature processing. The global feature G is output through the GVMB module. i,i=2,3,4,5 .

[0052] Figure 2 It is a flowchart for implementing the local visual mamba block LVMB according to an embodiment of the present invention;

[0053] like Figure 2 As shown, Local Top-k Feature Selection (LTTS) fully exploits sparsity by selecting the top k tags with the highest relevance to the query, thereby capturing the most critical disguised target information for extraction.

[0054] Specifically, we first use DWC to compress the input channel C by a factor S, so as to avoid the computational overhead of the subsequent operators. After passing the compressed features through SiLU, we use a fixed convolution kernel of size R×R to expand them. Through this operation, we can copy features and preserve the spatial relationship of nearby features. Then, a dense matrix is ​​generated by self-transposed dot product operation. k is dynamically set to a range of values. Taking k1=1 / 2 as an example, only the top 50% of the elements are retained for activation, while the remaining 50% are masked to 0; similarly, when k4=4 / 5, the sparsity rate is 20%. Unlike fixed k values, which lack flexibility in exploring the potential degree of sparsity, the proposed dynamic selection allows the selection process from sparse to dense by setting k to multiple values. Finally, we reshape and flatten the selected spatial dimensions to form a one-dimensional sequence of local features, ensuring that neighboring tokens within the local window are along the channel axis of the query token. The number of output channels is C'=CR 2 / S. For the local visual mamba block, the LTTS module is introduced as a feature extractor. Compared with the general visual mamba block feature extractor, the local visual mamba block contains an expansion operator, which can effectively ensure the spatial proximity of adjacent markers in a two-dimensional or three-dimensional array. The kernel size commonly used in the convolutional layer is used, and the window size R=3 is set. Since the number of channels of LTTS becomes C', a linear layer is inserted after the state space module (SSM) to project the markers into the original C-dimensional space.

[0055] Figure 3 This is a flowchart for implementing the global visual mamba block GVMB according to an embodiment of the present invention;

[0056] like Figure 3 As shown, Global Top-k Feature Selection (GTTS) generates global and spatially independent channel labels.

[0057] Specifically, given an input feature map of size H×W×C′, we use a dilation DWC with a step size of K×K to spatially compress it, and then flatten its spatial size to generate a shape of C′×HW / K 2 The global mark of , and the channel size and spatial size are transposed. Similar to LTTS, the dense matrix M∈R is then generated by the self-transposition dot product operation C'×C' , k is dynamically set to a range of values. For computational efficiency, each set of input channels is compressed into a global tag in all spatial dimensions. The effect of these steps is to learn an approximation of the global context in each input channel, so the problem of fine-grained detail loss associated with dilated DWC is not severe. Finally, we use a linear layer to project these selected tags into the C′-dimensional space and apply the activation function SiLU. For the global visual mamba block, the GTTS module is introduced as a feature extractor, which is combined with the Top-k feature selection mechanism after deep convolutional layers DWC, activation functions SiLU, and flattening to generate features with GRF, which are then forwarded to the next SSM. The number of channels in the output remains unchanged, while the sequence length increases, so the linear layer after SSM is cancelled.

[0058] Step 3: The global features G of the same level i and local feature L i Perform cross tensor decomposition to generate complementary features through low-rank factor matrix expansion and cross merging;

[0059] In this implementation, the same level features L i and G i By decomposing the convolutional layer, local information is selectively integrated into the global information to form complementary global features, and global information is selectively integrated into the local information to form complementary local features, thereby solving the problem of redundant feature learning between different levels in the process of camouflaged target detection.

[0060] Figure 4 This is a flowchart for implementing the cross tensor decomposition convolutional layer according to an embodiment of the present invention;

[0061] like Figure 4 As shown, first we calculate G i and L i Learn two trainable low-rank factor matrices, and generate the convolution weights of each layer through their product. The weights of the convolution layer are expressed as where D1 and D2 represent the width and height of the kernel space window, S and T represent the number of input channels and the number of kernels learned by the layer, respectively. The number of trainable parameters in a standard convolutional layer is given by P = TSD2D1. For a decomposed convolutional layer, we start with two trainable factors A∈R TS×R and Initially, r is their inner dimension, representing the rank of the original weight matrix (before decomposition). In order to better fuse features at different levels, we enhance the network's capabilities by slightly expanding the number of trainable parameters, which is achieved by increasing the columns / rows of the factor matrix. The factor matrices that emerge from the increased columns / rows effectively act as parallel trainable branches, enabling the network to extract key features from different levels to form complementary information. We increase r by Δr (where Δr>0) in matrices A, a, B, and b, resulting in A', a'∈R TS×(r+Δr) and r+Δr as their new internal dimension. Δr=16, and finally cross-merge to generate enhanced weight matrices M'=A'×b'=M+Δm and m'=a'×B'=m+ΔM, where Δm=ΔAΔb, ΔM=ΔaΔB, It can be derived from Δm, the parameter increment ΔP_fac=Δr(TS+D1D2), and the convolution filter K, k can be derived from M', m'.

[0062] In order to clearly capture L i and G i Complementary clues between layers, we maximize the L2 or L1 distance between the feature maps output from each branch of the layer and use it as an additional term in the task training objective. This can be achieved by the following loss Lc: Lc = -||K*x G -ΔK*x G -k*x L -Δk*x L || p ;

[0063] Where p = {1, 2}, * represents convolution, and x represents a matrix vector. The dimensions of K, k and ΔK, Δk are the same.

[0064] Step 4: Complementary global features, local features and previous layer fusion features F i-1 Input hierarchical fusion module, realize cross-level feature interaction through gated convolution and dynamic weighting mechanism, and output fusion feature F i,i=2,3,4,5 ;

[0065] In this embodiment, the hierarchical fusion module includes three branches, which can adaptively fuse local features, global representations, and previous-layer fusion semantic information according to input features. It includes a global weight generation path, a local weight generation path, and a previous-level feature weight generation path. The outputs of the three paths are spliced, and then redundant information is filtered through gated convolution, and residual connections are performed. Weighted fusion of features at different levels enables the model to adaptively adjust the fusion strategy according to the characteristics of the input features, thereby improving the fusion effect. The feature interaction of different paths also avoids the loss of information.

[0066] Figure 5 This is a flowchart of the implementation of the hierarchical fusion module described in an embodiment of the present invention;

[0067] like Figure 5 As shown, the global feature G of the input i First, it goes through the GlobalAvgPooling operation, then through Point-wise Conv, BN, ReLU activation function, Point-wise Conv and BN again, and finally through the Sigmoid function to get a weight value for subsequent weighted operations on features. The local feature Li goes through Point-wise Conv, BN, ReLU, Point-wise Conv and BN again, and then through the Sigmoid function to get another weight value. The previous layer feature F i-1 First, Point-wise Conv is performed, followed by GlobalAvgPooling, and then it is spliced ​​with the intermediate result of the processed global feature branch, and then LN (layer normalization), Point-wise Conv and GELU activation function are performed to obtain an intermediate feature representation. The weight values ​​of the global features and local features obtained by Sigmoid are multiplied with the intermediate features after the previous layer of feature processing, and then the two weighted features are spliced. The spliced ​​features are passed through GConv (gated convolution), and then multiplied and added with the previous spliced ​​features, and finally the fused features Fi are output through Point-wise Conv. Gated convolution is used to filter redundant information to improve the feature discrimination ability, and residual connection is performed to avoid problems such as gradient disappearance, explosion and network degradation to a certain extent, thereby effectively capturing global and local feature information at all levels.

[0068] Step 5: Input the F2-F5 features into the hybrid weighted decoder, and perform reverse optimization in combination with the channel attention mechanism of the DSE module. Generate and upsample the prediction map layer by layer, and finally output the camouflaged target detection result. The input of the last two decoders not only contains the fused features of the current level, but also the feature maps output by the first two levels of decoders.

[0069] In this embodiment, the hybrid weighted decoder integrates different information through cross-layer aggregation, dynamic weighting and reverse optimization, and can mine and utilize potential important information with different characteristics. The previous fusion features and the segmentation features output by the adjacent decoder are concatenated, convolved, reshaped and reversed, and then the feature map obtained by the addition and multiplication operation is added to the feature of the previous segmentation feature map after the dynamic weight allocation of the DSE module to finally output the segmentation map D i The DSE module of the hybrid weighted decoder includes a dual attention branch and a weight fusion part. Compared with the traditional SE module, the strong channel attention of DSE enables the model to capture the dependency between feature channels from the perspectives of average pooling and maximum pooling, and suppress irrelevant channels. The dynamic weight mechanism is introduced so that features at different levels can be adaptively weighted according to their importance during the fusion process, and the degree of dependence on different features can be flexibly adjusted.

[0070] Figure 6 A flow chart of the implementation of the hybrid weighted decoder according to an embodiment of the present invention;

[0071] like Figure 6 As shown in the figure, the features G5 and L5 are first input into the hierarchical fusion module, and a coarse prediction map D6 with one channel is generated through a series of dimensionality reduction convolutions. Then the hybrid weighted decoder receives F i (features from the previous fusion module), D i+1 and D i+2 (features of adjacent levels) as input, and F i With D i+1 , D i+2 The concatenation operation is performed, followed by the convolution operation. The processed features are reshaped and reversed, and then added and multiplied with the original concatenated features. The multiplication operation here involves the weights generated by the DSE module.

[0072] Figure 7 This is a flow chart for implementing the dynamic compression and excitation DSE according to an embodiment of the present invention;

[0073] like Figure 7 As shown, the module receives input features D i, passing through the average pooling and maximum pooling branches respectively. Each branch then passes through the FC (fully connected layer), RELU activation function, FC and Conv (convolutional layer) again, and finally generates weights w1 and w2 through the Sigmoid function. These two weights are multiplied with the original input feature D respectively, and then the weighted results are added to output the weighted feature D' for subsequent calculations in the hybrid weighted decoder. After a series of processing, it undergoes another convolution operation and is finally added to the features processed by the DSE module to output the final decoded feature D i .

[0074] Step 6: Perform multi-level supervision on the segmentation map and the final prediction map output by the hybrid weighted decoder, and jointly optimize the model parameters through the loss function;

[0075] In this embodiment, the segmentation map D6 is replaced by the output of the previous stage fusion module, and the output of each decoder is also an input of the next stage decoder. D6 and the four decoders are connected adjacently to generate the camouflaged target segmentation map D i,i=2,3,4,5 The loss is calculated with the ground truth graph GT of the training set. The selected structured loss function is the weighted binary cross entropy loss. Weighted Intersection-over-Union Loss And tensor loss Lc. To handle multi-scale outputs, all outputs D k=2,3,4,5,6 Upsampling is performed to match the resolution of the segmentation truth map GT. The formula of the total loss function L of the disguised target detection model is as follows:

[0076]

[0077] Where k represents different network layers; D i Represents the segmentation map of the disguised target at different layers; GT represents the true value label of the disguised target;

[0078] Weighted Binary Cross-Entry Loss Function It has been widely used in the field of image segmentation tasks and is defined as follows:

[0079]

[0080] Among them, w ij Represents the weight value of pixel (i, j). ij and G ij The method prediction value and GT true value of the camouflaged object pixel (i, j). In order to optimize the global structure, the weighted intersection loss is introduced It can calculate the complete structural similarity. The calculation formula is as follows:

[0081]

[0082] Figure 8 It is a schematic diagram for qualitative comparison between six mainstream camouflage target segmentation models and the method of the present invention in comparative experiments of embodiments of the present invention.

[0083] Among them, GT represents the saliency label of the disguised target, and Our represents the prediction result of the disguised target of the present invention. By comparison, it can be seen that the present invention can effectively detect the disguised target, and the edge contour of the disguised target is clearer, which improves the accuracy of the disguised target detection.

[0084] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A camouflaged target detection method based on dynamic Top-k feature selection and cross tensor decomposition, characterized by: The method comprises: S1, obtain the disguised target image, input the disguised target image into the encoder network to extract multi-level features, and generate feature maps of different scales E1-E5; S2, input the E2-E5 layer features into the global perception module and the local optimization module respectively to extract features of different scales, screen out global key features and local sensitive features, and output global features G2-G5 and local features L2-L5; S3, for the global feature G at the same level i and local feature L i Perform cross tensor decomposition to generate complementary features through low-rank factor matrix expansion and cross merging; S4, the complementary global features, local features and the previous layer fusion features F i-1 Input hierarchical fusion module, which realizes cross-level feature interaction through gated convolution and dynamic weighting mechanism, and outputs fused features F2-F5; S5, input the F2-F5 features into the hybrid weighted decoder, combine the channel attention mechanism of the DSE module for reverse optimization, generate and upsample the prediction map layer by layer, and finally output the camouflaged target detection result. The input of the last two decoders not only includes the fusion features of the current level, but also the feature maps output by the first two levels of decoders; S6. Perform multi-level supervision on the segmentation map and the final prediction map output by the hybrid weighted decoder, and jointly optimize the model parameters through the loss function.

2. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition according to claim 1, characterized in that: The encoder network in S1 includes: The resnet-50 network is used to extract five initial features. Since the resolution of the first-layer feature E1 is greater than the resolution threshold and there is background noise, E1 is discarded and E2-E5 are used to complete the subsequent COD task.

3. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition according to claim 1, characterized in that: The global perception module and local refinement module in S2 include: The global perception module is composed of several global visual mamba blocks connected together, and is used to use the global visual mamba blocks to obtain all pixel relationships from a global perspective; in the global visual mamba blocks, a dynamic Top-k feature selection mechanism is introduced to screen the most critical global features for the detection task, and output global features G2-G5; The local refinement module is composed of several local visual mamba blocks connected together, and is used to increase the spatial local information in the initial features. In the local visual mamba block, a dynamic Top-k feature selection mechanism is also introduced to dynamically screen the key features in the local features and output the local features L2-L5, so as to filter out the background local features irrelevant to the target and improve the sensitivity of the model to the key areas. The local Top-k feature selection and the global Top-k feature selection corresponding to the local refinement module and the global perception module respectively make full use of sparsity by selecting the top k tags with the highest correlation with the query, thereby capturing the most critical camouflaged target information for extraction; Among them, k is dynamically set to a series of values ​​that satisfy k i =(i+1) / (n+2); i represents the current level index, n represents the total number of levels, and the difference rate of k values ​​of different levels is not less than 15%.

4. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition as claimed in claim 3, characterized in that: The local visual mamba block and the global visual mamba block include: The local visual mamba block adopts the introduction of a local Top-k feature selection module as a feature extractor, and the local visual mamba block includes an expansion operator to effectively ensure the spatial proximity of adjacent markers in a two-dimensional or three-dimensional array; The global visual mamba block introduces a global Top-k feature selection module as a feature extractor, which is combined with a Top-k feature selection mechanism after a deep convolution layer DWC, an activation function SiLU, and flattening to generate features with GRF, which are then forwarded to the next SSM module.

5. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition as claimed in claim 1, characterized in that: S3 includes: transforming the same-level features G i and L i By decomposing the convolutional layer, local information is selectively integrated into the global information to form complementary global features, and global information is selectively integrated into the local information to form complementary local features: For the global feature G i and local feature L i Construct low-rank factor matrices A, a∈R respectively TS×r and B,b∈R r×D1D2 ; Then expand the low-rank factor matrix dimensions to A', a'∈R TS×(r+Δr) and B',b'∈R (r+Δr)×D1D2 ; Among them, Δr = 16, S and T represent the number of input channels of the convolution layer and the number of kernels learned by the layer, respectively, D1 and D2 represent the width and height of the convolution kernel space window, respectively, and r represents the rank of the original weight matrix before decomposition; Finally, the cross-merging generates the enhanced weight matrices M'=A'×b' and m'=a'×B', where the parameter increment ΔP_fac=Δr(TS+D1D2).

6. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition as claimed in claim 5, characterized in that: In order to clearly capture L i and G i Complementary cues between layers maximize the L2 or L1 distance between feature maps output from each branch of the layer and serve as an additional term in the task training objective; According to the loss Lc formula: Lc = -||K*x G -ΔK*x G -k*x L -Δk*x L || p ; The distance Lc between the output feature maps is obtained; where p = {1, 2}, * represents convolution, x represents a vector, and the dimensions of K, k and ΔK, Δk are the same, representing a convolution filter.

7. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition according to claim 1, characterized in that: The hierarchical fusion module in S4 includes: Global weight generation path: GlobalAvgPool→1x1 convolution→BN→ReLU→1x1 convolution→Sigmoid; Local weight generation path: double 1×1 convolutional layer cascade structure with BN batch normalization and GELU loss function; Front-end feature weight generation path: dual 1×1 convolutional layer cascade structure with batch normalization and ReLU loss function; The outputs of the three paths are concatenated, and then redundant information is filtered through gated convolution, and residual connection is performed to achieve cross-level feature interaction and output fused features F2-F5.

8. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition as claimed in claim 1, characterized in that: The hybrid weighted decoder in S5 comprises: The hybrid weighted decoder is used to integrate different information using cross-layer aggregation, dynamic weighting and inverse optimization: Specifically, the previous fusion features and the output segmentation features of the adjacent decoders are concatenated, convolved, reshaped, and reversed, and then the feature map obtained by the addition and multiplication operation and the previous segmentation feature map are dynamically weighted through the DSE module. Finally, the allocated features are added and the segmentation map D is output. i ; Among them, the DSE module in the hybrid weighted decoder includes a dual attention branch and a weight fusion part; specifically, the feature map is subjected to two branches of average pooling and maximum pooling respectively, and each branch subsequently passes through a fully connected layer, a RELU activation function, and then passes through a fully connected layer and a convolutional layer again, and finally generates weights through a Sigmoid function.

9. The method for detecting camouflaged targets based on dynamic Top-k feature selection and cross tensor decomposition as claimed in claim 1, characterized in that: The S6 comprises: The segmentation map D6 is replaced by the output of the previous level fusion module, and the output of each level decoder is also an input of the next level decoder. The camouflaged target segmentation map D generated by connecting D6 and the four decoders adjacently i,i=2,3,4,5 , the loss is calculated with the ground truth graph GT of the training set, and the selected structural loss function is the weighted binary cross entropy loss Weighted Intersection-over-Union Loss and tensor loss Lc; To handle multi-scale outputs, all outputs D k=2,3,4,5,6 Upsample to match the resolution of the segmentation ground truth map GT; The formula of the total loss function L of the camouflaged target detection model is as follows: Where k represents different network layers; D i Represents the segmentation map of the disguised target at different layers; GT represents the true value label of the disguised target.

Citation Information

Patent Citations

  • Camouflage target segmentation method based on deep learning technology

    CN118038047A

  • High-resolution remote sensing image target detection method based on multi-scale network

    CN118485927A

  • Camouflage target detection method for self-attention conversion network with dense receptive fields

    CN118941772A

  • Multistage refined camouflage target detection method based on wavelet enhancement recognition

    CN119540520A

Cited By

  • Preliminary screening method and system for hyperactivity disorder based on surface information time sequence characteristics

    CN120234674A

  • ADHD initial screening method and system based on temporal characteristics of surface information

    CN120234674B