Camouflage target detection method based on Mangbar capsule routing

By using a Mamba capsule routing method in the camouflage object detection algorithm, pixel-level space capsules are converted into type-level Mamba capsules, which solves the problems of incomplete detection and routing complexity, and achieves more efficient camouflage object detection.

CN120107716APending Publication Date: 2025-06-06NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510162575.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing camouflage object detection algorithm has the problem of incomplete detection, and the EM routing algorithm used by the existing capsule network has large parameters and high computational complexity, resulting in slow model inference speed.

Method used

Using a masquerading object detection method based on Mamba capsule routing, a target detection network including an encoder module, a Mamba capsule generation module, an EM routing layer, a space capsule recovery module and a decoder module is used to convert pixel-level space capsules into type-level Mamba capsules, thereby reducing the amount of parameters and calculation complexity of EM routing.

Benefits of technology

It greatly reduces the parameter quantity and calculation complexity of capsule routing, improves detection speed, and optimizes the detection effect of camouflage targets, solving the problems of incomplete detection and routing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107716A_ABST
    Figure CN120107716A_ABST
Patent Text Reader

Abstract

The invention provides a camouflage target detection method based on Mangbar capsule routing. A target detection network based on Mangbar capsule routing is constructed, the target detection network comprises an encoder module, a Mangbar capsule generation module, an EM routing layer, a space capsule recovery module and a decoder module, the Mangbar capsule generation module adopts a brand-new capsule generation mechanism, and the space capsule recovery module adopts an EM routing mechanism. According to the method, global information construction is achieved, meanwhile, implicit retention of the context relation of the capsule sequence is completed, and type-level capsules without spatial resolution can be generated, so that subsequent EM routing is changed from a pixel level to a type level, and the algorithm complexity is greatly reduced. According to the method, the parameter quantity and the calculation complexity of the original capsule routing can be greatly reduced, and the camouflage target detection effect is optimized while the detection speed is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection in camouflage scenarios, and in particular relates to a camouflage target detection method based on Mamba capsule routing. Background Art

[0002] Object detection is a basic task in computer vision, which involves identifying and locating objects in images or videos. Unlike general object detection, camouflaged object detection is more challenging because there is a high degree of similarity between the camouflaged object and the background. The application areas of camouflaged object research are very broad. In addition to its academic value, camouflaged object detection can also help promote the search and detection of camouflaged hidden targets in the military, the diagnosis of disease in the medical field, and the invasion of locusts in agricultural remote sensing.

[0003] Early researchers extracted manually annotated features, including color, texture, and optical flow, but these methods still had difficulty distinguishing objects from backgrounds. With the development of deep learning and the public release of large-scale datasets, a lot of deep learning-based work has emerged. By simulating the hunting mechanism of the biological world, human visual mechanism designs modules that mine subtle features; by combining multiple downstream tasks to optimize the detection performance of the main task. Most of the above models are based on convolutional neural networks (CNNs) and Transformers. However, the strong intrinsic correlation between camouflaged objects and their backgrounds hinders the performance of the above networks, which can easily lead to incomplete detection.

[0004] Existing studies have shown that the inherent characteristics of capsule networks in characterizing the relationship between parts and the whole are beneficial to the integrity segmentation of camouflaged objects. In order to solve the above-mentioned problem of incomplete detection, the present invention continues to use this characteristic to explore the relationship between parts and the whole of camouflaged objects. However, the disadvantages of the original EM routing, such as large number of parameters and high computational complexity, make the model's reasoning speed relatively high, which in turn hinders this progress. The reason is that the original routing is still at the pixel level, resulting in large-scale capsule matching.

[0005] In summary, there are two main problems with existing camouflaged target detection algorithms: 1) Deep learning networks based on CNN and Transformer are prone to incomplete detection; 2) The EM routing algorithm used in existing capsule networks is still pixel-level, with a large number of parameters and high computational complexity. Summary of the invention

[0006] In order to overcome the problem of large number of parameters and high computational complexity of the existing camouflaged target detection algorithm designed by using the characteristics between components and the whole, the present invention provides a camouflaged target detection method based on Mamba capsule routing. A target detection network based on Mamba capsule routing is constructed, including an encoder module, a Mamba capsule generation module, an EM routing layer, a spatial capsule restoration module and a decoder module, wherein the Mamba capsule generation module adopts a new capsule generation mechanism, which completes the implicit retention of the contextual relationship of the capsule sequence while realizing the construction of global information, and can generate type-level capsules without spatial resolution, so that the subsequent EM routing is changed from pixel level to type level, greatly reducing the algorithm complexity. The present invention can greatly reduce the number of parameters and computational complexity of the original capsule routing, and optimize the camouflaged target detection effect while improving the detection speed.

[0007] A camouflaged target detection method based on Mamba capsule routing is characterized by the following steps:

[0008] Step 1: Prepare training dataset and perform data enhancement: construct training dataset and test dataset based on public dataset, and perform data enhancement on images in the dataset;

[0009] Step 2, construct a camouflaged target detection network based on Mamba capsule routing: including an encoder module, a Mamba capsule generation module, an EM routing layer, a spatial capsule restoration module and a decoder module, wherein the encoder module performs global attention modeling on the input RGB image to obtain an image representation sequence; the image representation sequence is input into the Mamba capsule generation module to generate a Mamba capsule; the Mamba capsule is subjected to EM routing analysis by the EM routing layer to obtain a new Mamba capsule; the new Mamba capsule is input into the spatial capsule restoration module for spatial detail restoration, and then passes through the decoder module to simultaneously generate a camouflaged target detection map and a camouflaged edge detection map;

[0010] Step 3, training the network: using the images in the training data set obtained in step 1 as input, train the object detection network constructed in step 2;

[0011] Step 4, disguised target detection: Input the test data set image into the trained target detection network, output the disguised target detection map, and complete the disguised target detection.

[0012] Furthermore, the specific method for constructing a training dataset and a test dataset based on a public dataset described in step 1 is as follows: for the public datasets CAMO, COD10K, and NC4K, they are divided into 4040 training datasets and 6397 test datasets according to the default settings. The datasets include camouflaged target RGB images, camouflaged target edge maps, and binary true value segmentation maps.

[0013] Furthermore, the specific method of the data enhancement processing described in step 1 is: first, the images in the training data set are adjusted to a size of 384*384, and then, they are adjusted to a size of 352*352 by random cropping and random arrangement to complete the data enhancement processing.

[0014] Furthermore, the encoder module adopts the image extraction module in the VSCode model published in the Github community https: / / github.com / Sssssuperior / VSCode, and the decoder module adopts the pure Transformer structure in the VSCode model published in the Github community https: / / github.com / Sssssuperior / VSCode.

[0015] Furthermore, the processing process of the Mamba capsule generation module is as follows: the representation sequence of the input image is first subjected to a conv convolution operation and a sigmoid activation function and a split slicing operation along the channel dimension to generate 32 spatial capsules F, and each capsule is respectively expanded and scanned along the positive Z-shape, reverse Z-shape, positive N-shape and reverse N-shape directions in the height H and width W planes to obtain a corresponding capsule sequence; then, in each direction, the 32 capsule sequences are respectively input into the selective SSM module, and the last hidden state in the module is taken as the Mamba vector to obtain 32 Mamba vectors; finally, the Mamba vector is subjected to a conv convolution operation and a sigmoid function and a merge operation along the channel dimension to generate a Mamba capsule; wherein the selective SSM module is published in the VMamba model of the Github community https: / / github.com / MzeroMiko / VMamba.

[0016] Furthermore, the EM routing layer performs EM routing analysis, which refers to iteratively calculating the Mamba capsule using an expectation-maximization algorithm to obtain a new Mamba capsule.

[0017] Furthermore, the specific processing process of the spatial capsule restoration module is as follows: first, the similarity of the Mamba capsule input to the EM routing layer and the Mamba capsule output by the EM routing layer is calculated in each scanning direction to obtain a 32*32 similarity matrix respectively; then, the similarity matrix is ​​activated by a sigmoid function, and the Mamba capsule output by the EM routing layer is multiplied by the activated similarity matrix in the corresponding direction to obtain the corresponding spatial information matrix; finally, the spatial information matrices in all directions are integrated together according to the scanning direction using a merge operation, and after a multi-head attention mechanism, complementary capsule features in four directions are obtained to complete the restoration of spatial details.

[0018] Furthermore, the specific method of training the network described in step 3 is: using the PyTorch framework for training under the Nvidia4090 graphics card, the encoder loads the pre-trained weights, the remaining modules are randomly initialized, the training uses the Adam optimizer and sets the initial value to e -4 , and decays tenfold at half and three-quarters of the total number of iterations. The total number of training iterations is 150,000. The total loss of the training is the sum of the weighted binary cross entropy loss and the intersection-and-union ratio loss. The weighted binary cross entropy loss is the loss value obtained by weighted summing the binary cross entropy losses of the disguised target detection map and disguised edge detection map output by the three layers of the decoder and the binary true value segmentation map and disguised target edge true value map of the corresponding size. The intersection-and-union ratio loss refers to the calculation of the intersection-and-union ratio loss of the disguised target detection map finally output by the decoder and the binary true value segmentation map corresponding to the input image.

[0019] The beneficial effects of the present invention are as follows: by utilizing the selective SSM mechanism in the Mamba model to highly generalize the pixel-level spatial capsules into type-level Mamba capsules, a new capsule generation mechanism is formed. With the continuous input of capsule tokens in the capsule sequence, the information of the capsule tokens is selectively compressed. Its unique hidden state mechanism can continuously accumulate the selected capsule token information, thereby completing the construction of global information. Compared with the traditional full connection operation, the Selective SSM mechanism can be used to generate the capsule tokens in the capsule sequence. The updating mechanism of hidden state variables in SSM makes spatial capsules more effectively summarized, reduces the amount of spatial structure calculation while implicitly retaining the contextual relationship of capsule sequences, completes the modeling of global representation, and converts pixel-level spatial capsules with resolution into type-level Mamba capsules without resolution, so that the number of parameters and computational complexity required for subsequent EM routing are greatly reduced; since the spatial capsule restoration module can restore the spatial information of the Mamba capsules after routing (characterizing the whole), a capsule feature map with full resolution that characterizes the whole can be obtained, so as to better detect camouflaged objects; the present invention completes the component-whole analysis of the feature map in a more lightweight routing manner, which is more beneficial to the integrity detection of the camouflaged target, improves the problem of incomplete detection in the existing camouflaged target detection algorithm, and also solves the problem that the existing capsule network cannot be better expanded due to the complexity of routing. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flow chart of the camouflaged target detection method based on Mamba capsule routing of the present invention;

[0021] Figure 2 This is a schematic diagram of a camouflaged target detection network based on Mamba capsule routing;

[0022] Figure 3 This is a schematic diagram of the Mamba capsule generation module;

[0023] Figure 4 This is a visual comparison of the target detection results of different methods on the test dataset. DETAILED DESCRIPTION

[0024] The present invention is further described below in conjunction with the accompanying drawings and embodiments. The present invention includes but is not limited to the following embodiments.

[0025] In order to solve the problem of incomplete detection in the existing disguised target detection algorithm and the problem that the existing capsule network cannot be better expanded due to routing complexity, the present invention provides a disguised target detection method based on Mamba capsule routing. The overall flow chart is as follows: Figure 1 As shown, the specific implementation process is as follows:

[0026] 1. Prepare training data set and perform data augmentation

[0027] For the existing public datasets CAMO, COD10K, and NC4K, they are divided into 4040 training datasets and 6397 test datasets according to the default settings. The datasets include camouflaged target RGB images, camouflaged target edge maps, and binary true value segmentation maps; each image in the training dataset is resized to 384*384, and then resized to 352*352 by random cropping and random arrangement to obtain the enhanced input image X.

[0028] 2. Construct a camouflaged target detection network based on Mamba capsule routing

[0029] like Figure 2 As shown in the figure, the entire network framework includes an encoder module, a Mamba capsule generation module, an EM routing layer, a spatial capsule restoration module and a decoder module.

[0030] The encoder module uses the public image extraction module Swin-Transformer encoder to perform global attention modeling on the input RGB image X to obtain a representation sequence of images with contextual relationships. The Swin-Transformer encoder is published in the VSCode model in the Github community https: / / github.com / Sssssuperior / VSCode.

[0031] The image representation sequence is input into the Mamba capsule generation module, first passing through the spatial capsule generation layer, and then through the conv convolution and sigmoid function and using the split slicing operation along the channel dimension to obtain 32 spatial (pixel-level) capsules F with resolution, which are composed of a pose matrix and an activation value. In order to more accurately convert the two-dimensional features into a one-dimensional sequence, such as Figure 3As shown in FIG. 1 , each capsule height H and width W plane is scanned in four different directions (positive 'Z' shape, reverse 'Z' shape, positive 'N' shape, reverse 'N' shape) to obtain a capsule sequence Γ with different contextual relationships. i , (i=1, 2, 3, 4). In each direction i, 32 capsule sequences are input into different selective SSM modules (in the VMamba model published in the Github community https: / / github.com / MzeroMiko / VMamba), expressed as:

[0032]

[0033] in, Indicates the relevant parameters in the selective SSM module, h ij (u) represents the hidden state obtained by iterative update after the u-th capsule token in the j-th capsule sequence is input under scanning direction i, h ij (u-1) represents the hidden state obtained by iterative update after the u-1th capsule token is input, Γ ij (u) represents the u-th capsule token in the j-th capsule sequence under scanning direction i.

[0034] The last hidden state in the module is taken as the Mamba vector, and a total of 32 Mamba vectors are obtained. The Mamba vector is used as the posture matrix of the Mamba capsule. At the same time, it is subjected to the conv convolution operation and the sigmoid function to obtain the activation value of the Mamba capsule. Then, the concatenate operation is used along the channel dimension to obtain the Mamba (type level) capsule M i .

[0035] The Mamba capsule generation module of the present invention utilizes the selective SSM mechanism in the Mamba model to highly generalize the pixel-level spatial capsules into type-level Mamba capsules, forming a new capsule generation mechanism. With the continuous input of capsule tokens in the capsule sequence, the information compression of the capsule tokens is selectively performed, and its unique hidden state mechanism can continuously accumulate the selected capsule token information, thereby completing the construction of global information. Compared with the traditional fully connected operation, the update mechanism of the hidden state variables in the Selective SSM more effectively generalizes the spatial capsules, realizes the implicit retention of the contextual relationship of the capsule sequence while reducing the amount of spatial structure calculation, and completes the modeling of global representation. The pixel-level spatial capsules with resolution are converted into type-level Mamba capsules without resolution, so that the number of parameters and computational complexity required for subsequent EM routing are greatly reduced.

[0036] Mamba (Type Grade) Capsule M iThe input is sent to the EM routing layer, and the expectation maximization algorithm mentioned in the paper "Matrix capsules with EMRouting" is used to iteratively calculate new higher-level capsules. Specifically, the 32 Mamba capsules in the four directions are iteratively calculated to obtain high-level Mamba capsules corresponding to the four different contexts.

[0037] For the subsequent prediction tasks, the present invention designs a spatial capsule restoration module to restore the spatial details of the high-level Mamba capsule. First, the similarity of the Mamba capsules before and after the EM routing layer is calculated in each scanning direction i to obtain a 32*32 similarity matrix E i .

[0038]

[0039] Among them, E i (m, n) represents the elements in the similarity matrix under scanning direction i, and m and n represent the Mamba capsule M respectively. i , The index is [1,…,32], x, y represent the Mamba capsule M i , The channel dimension index value is [1,…,O], where O represents the maximum dimension of the channel, which is set to 17.

[0040] Then, in order to enhance the difference and distinguishability, the activation is completed through the sigmoid function to obtain And the corresponding capsule sequence Γ i The features and similarity matrix M i Perform matrix multiplication to obtain the spatial information matrix The spatial information of the Mamba capsule representing the overall relationship is restored. Then, the spatial information matrix in each direction is combined and integrated using concatenate according to the corresponding scanning method. After the multi-head attention mechanism, the complementary capsule features in four directions are obtained and input into the decoder module.

[0041] The decoder module uses a public pure Transformer structure to complete target and edge detection tasks, which has been published in the VSCode model of the Github community https: / / github.com / Sssssuperior / VSCode.

[0042] 3. Training the network

[0043] The image in the training dataset obtained in step 1 is used as input to train the object detection network constructed in step 2. Specifically, the entire network is trained using the PyTorch framework under the Nvidia 4090 graphics card. The encoder loads the pre-trained weights, and the remaining modules are randomly initialized. The training uses the Adam optimizer and sets the initial value to e. -4 , and decayed tenfold at half and three-quarters of the total number of iterations. The total number of training iterations was 150,000. The decoder outputs the disguised target detection map and the binary true segmentation map corresponding to the input image to calculate the intersection-over-union loss. The output disguised target detection map and disguised edge detection map in the three layers of the decoder are compared with the binary true segmentation map and the disguised target edge true map of the corresponding size to calculate the binary cross entropy loss (l wbce ), and weighted sum.

[0044] The total loss for training is set as the sum of the weighted binary cross entropy loss and the intersection-over-union loss:

[0045] Loss = l wbce +l iou (3)

[0046] Among them, l wbce is the weighted binary cross entropy loss value, l iou is the intersection-over-union loss value.

[0047] Weighted binary cross entropy loss l wbce The calculation formula is as follows:

[0048]

[0049] Among them, l bce is the binary cross entropy loss, They are the target detection prediction map and target detection truth map corresponding to the kth layer of the encoder, respectively. are the target edge prediction map and target edge truth map of the corresponding encoder k layer, respectively, k is [1,…,4], w k is the corresponding weight, [w 1 、w 2 、w 3 、w 4 ]The corresponding values ​​are [0.5,0.5,0.8,1].

[0050] Binary cross entropy loss bce The calculation formula is as follows:

[0051]

[0052] Among them, p and n represent the pixel-level index and total number of pixels of the prediction image respectively, G prepresents the value at index p in the truth graph, Represents the value at index p in the prediction graph.

[0053] Intersection-over-union loss iou The calculation formula is as follows:

[0054]

[0055] in, Represents the value at index p in the prediction graph, G c (p) represents the value at index p in the truth map.

[0056] 4. Camouflage target detection

[0057] Input the test data set image into the disguised target detection network trained in step 3, output the disguised target detection map, and complete the disguised target detection.

[0058] Table 1 shows the experimental results of the present invention and the current mainstream methods on the public data sets CAMO and COD10K. The indicators for measuring model performance are the mean absolute error MAE, which measures the difference between the predicted result and the true value, and the smaller the value, the better; the F measure Fm, which measures the precision and recall of the predicted result, and the larger the value, the better; the enhanced alignment Em, which measures the similarity between the predicted result and the true value, and the larger the value, the better; the structural similarity Sm, which measures the consistency of the predicted result and the true value structure, and the larger the value, the better.

[0059] Table 1

[0060]

[0061]

[0062] It can be seen from Table 1 that the performance of the present invention has surpassed the existing mainstream methods.

[0063] Table 2 gives the data statistics obtained by routing spatial capsules and Mamba capsules respectively, where Flops represents the computational complexity in billions (G), Params represents the number of parameters in megabytes (M), and Time represents the time for reasoning an image in seconds (s).

[0064] Table 2

[0065] Flops(G) Params(M) Time(s) Space capsule routing 155.16 77.96 0.039 Mamba Capsule Router 145.74 69.11 0.028

[0066] It can be seen from Table 2 that the improvement of the present invention reduces the number of parameters and the computational complexity and increases the reasoning speed, which means that the present invention effectively solves the problem that the existing capsule network cannot be better expanded due to the routing complexity.

[0067] Figure 4 A visual comparison chart of target detection results of different methods on the test data set is given, where the first column Image is the input image, the second column GT is the true value corresponding to the input image, the third column Ours is the target detection result image obtained by the present invention, and the subsequent columns are target detection result images obtained by the existing mainstream methods. Small is a small object, Large is a large object, Uncertainty is an image with blurred boundaries, Occiusion is an occluded image, and Person is an image containing a person. It can be seen that the present invention has a good effect on the detection of small, large, occluded and other objects, successfully alleviates the problem of incomplete detection, and has a good camouflage target detection effect.

[0068] The present invention analyzes the feature graph from a more lightweight component-overall perspective, which is more beneficial to the detection of camouflaged targets, improves the problem of incomplete detection in existing camouflaged target detection algorithms, and also solves the problem that existing capsule networks cannot be better expanded due to routing complexity.

Claims

1. A camouflaged target detection method based on Mamba capsule routing, characterized in that Here are the steps: Step 1: Prepare training dataset and perform data enhancement: construct training dataset and test dataset based on public dataset, and perform data enhancement on images in the dataset; Step 2, construct a camouflaged target detection network based on Mamba capsule routing: including an encoder module, a Mamba capsule generation module, an EM routing layer, a spatial capsule restoration module and a decoder module, wherein the encoder module performs global attention modeling on the input RGB image to obtain an image representation sequence; the image representation sequence is input into the Mamba capsule generation module to generate a Mamba capsule; the Mamba capsule is subjected to EM routing analysis by the EM routing layer to obtain a new Mamba capsule; the new Mamba capsule is input into the spatial capsule restoration module for spatial detail restoration, and then passes through the decoder module to simultaneously generate a camouflaged target detection map and a camouflaged edge detection map; Step 3, training the network: using the images in the training data set obtained in step 1 as input, train the object detection network constructed in step 2; Step 4, disguised target detection: Input the test data set image into the trained target detection network, output the disguised target detection map, and complete the disguised target detection.

2. A camouflaged target detection method based on Mamba capsule routing as claimed in claim 1, characterized in that: The specific method for constructing training datasets and test datasets based on public datasets described in step 1 is as follows: for the public datasets CAMO, COD10K, and NC4K, they are divided into 4040 training datasets and 6397 test datasets according to the default settings. The datasets include camouflaged target RGB images, camouflaged target edge maps, and binary true value segmentation maps.

3. The method for detecting a disguised target based on Mamba capsule routing according to claim 1, characterized in that: The specific method of the data enhancement processing described in step 1 is: first, all images in the training data set are adjusted to 384*384 size, and then, they are adjusted to 352*352 size by random cropping and random arrangement to complete the data enhancement processing.

4. The method for detecting a disguised target based on Mamba capsule routing according to claim 1, characterized in that: The encoder module adopts the image extraction module in the VSCode model published in the Github community https: / / github.com / Sssssuperior / VSCode, and the decoder module adopts the pure Transformer structure in the VSCode model published in the Github community https: / / github.com / Sssssuperior / VSCode.

5. The method for detecting a disguised target based on Mamba capsule routing according to claim 1, characterized in that: The processing process of the Mamba capsule generation module is as follows: the representation sequence of the input image is first subjected to the conv convolution operation and the sigmoid activation function and the split slicing operation along the channel dimension to generate 32 spatial capsules F, and each capsule is respectively expanded and scanned along the positive Z-shape, reverse Z-shape, positive N-shape and reverse N-shape directions in the height H and width W planes to obtain the corresponding capsule sequence; then, in each direction, the 32 capsule sequences are respectively input into the selective SSM module, and the last hidden state in the module is taken as the Mamba vector to obtain 32 Mamba vectors; finally, the Mamba vector is subjected to the conv convolution operation and the sigmoid function and the merge operation along the channel dimension to generate the Mamba capsule; wherein the selective SSM module is published in the VMamba model of the Github community https: / / github.com / MzeroMiko / VMamba.

6. The method for detecting a disguised target based on Mamba capsule routing according to claim 1, characterized in that: The EM routing layer performs EM routing analysis, which refers to iteratively calculating the Mamba capsule using the expectation maximization algorithm to obtain a new Mamba capsule.

7. The method for detecting a disguised target based on Mamba capsule routing according to claim 1, characterized in that: The specific processing process of the spatial capsule restoration module is as follows: first, the similarity of the Mamba capsule input to the EM routing layer and the Mamba capsule output by the EM routing layer is calculated in each scanning direction to obtain a 32*32 similarity matrix respectively; then, the similarity matrix is ​​activated by a sigmoid function, and the Mamba capsule output by the EM routing layer is multiplied by the activated similarity matrix in the corresponding direction to obtain the corresponding spatial information matrix; finally, the spatial information matrices in all directions are integrated together according to the scanning direction using a merge operation, and after a multi-head attention mechanism, complementary capsule features in four directions are obtained to complete the restoration of spatial details.

8. The method for detecting a disguised target based on Mamba capsule routing according to claim 1, characterized in that: The specific method of training the network described in step 3 is: use the PyTorch framework for training under the Nvidia4090 graphics card, load the pre-trained weights on the encoder, and randomly initialize the remaining modules. The training uses the Adam optimizer and sets the initial value to e -4 , and decays tenfold at half and three-quarters of the total number of iterations. The total number of training iterations is 150,000. The total loss of the training is the sum of the weighted binary cross entropy loss and the intersection-and-union ratio loss. The weighted binary cross entropy loss is the loss value obtained by weighted summing the binary cross entropy losses of the disguised target detection map and disguised edge detection map output by the three layers of the decoder and the binary true value segmentation map and disguised target edge true value map of the corresponding size. The intersection-and-union ratio loss refers to the calculation of the intersection-and-union ratio loss of the disguised target detection map finally output by the decoder and the binary true value segmentation map corresponding to the input image.

Citation Information

Cited By

  • RNA binding site prediction method based on Mangbar and graph neural network

    CN121256334A