Remote sensing image target detection method using scale sensitive Mama network
By using the remote sensing Mamba block and sparse feature fusion module of the scale-sensitive Mamba network, combined with the area intersection-over-union loss function, the problems of high computational cost, difficult target distinction and scale change in remote sensing image target detection are solved, and efficient and accurate remote sensing image target detection is achieved.
Patent Information
- Application Number
- CN202510626082.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-10-03
AI Technical Summary
Target detection in remote sensing images faces challenges such as high computational cost, complex geographical environment making it difficult to distinguish between background and target, loss of small target features, and significant scale changes. Traditional methods have degraded performance when applied to remote sensing images.
A scale-sensitive Mamba network is adopted, and local and global feature modeling is enhanced through the remote sensing Mamba block and sparse feature fusion module. The target detection is optimized by combining the area intersection-over-union loss function, thus improving the model's effect on target detection in remote sensing images.
It effectively improves the performance of target detection in remote sensing images, especially the feature expression and scale adaptability of small targets, and improves the detection accuracy and efficiency of the model in remote sensing images.
Smart Images

Figure CN120747679A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to the technical field of a remote sensing image target detection method using a scale-sensitive Mamba network. Background Art
[0002] Remote sensing target detection is an important field of cross-integration of computer vision and remote sensing technology, which aims to automatically identify and locate specific ground objects (such as buildings, vehicles, ships, aircraft, etc.) from high-resolution remote sensing images.
[0003] The rapid development of deep learning has also promoted significant progress in the field of target detection. However, remote sensing target detection still faces huge challenges. On the one hand, remote sensing images have extremely high resolution, which leads to huge computational costs. On the other hand, the targets in remote sensing images are usually very small and only occupy a very small part of the image. To address these problems, a lot of research has been carried out. Yang et al. [1] proposed an attention-based soft threshold filtering module to filter out redundant information in high-level feature maps, thereby enhancing the semantic features of small targets. Liu et al. [2] introduced an enhanced efficient channel attention mechanism to improve the feature representation ability of the backbone network, thereby reducing the interference of complex background on foreground targets. In order to adaptively fuse multi-scale features of different channels and spatial positions, they designed an adaptive feature pyramid network to capture more discriminative features. Zhang et al. [3] explored the spatial redundancy in remote sensing images and proposed an adaptive multi-granularity routing mechanism to promote the sparsity of labels in Transformer, which can significantly reduce the computational cost without sacrificing accuracy. Zhao et al. [4] proposed a new scene contextualized detection network (SCDNet), which decouples scene context information through a dedicated scene classification subnetwork, so as to better explore the relationship between small targets and their surroundings. Since the emergence of deep convolutional neural networks, general target detectors have achieved great success. However, target detection in remote sensing images still faces major challenges. The main difficulties of target detection in remote sensing images are: (1) The complex geographical environment in remote sensing images makes it difficult for the model to distinguish between background and target objects. (2) Small target objects in remote sensing images occupy relatively few pixels, making it difficult for the model to extract their features. In deeper networks, the features of small targets are often severely lost. (3) Targets in remote sensing images usually have significant scale variations.
[0004] In recent years, several Transformer-based methods [5][6][7][8][9]
[10]
[11] have been applied to the task of object detection in remote sensing images. Transformer
[12] has the ability to model global context, so many papers
[13]
[14] have combined visual transformers (ViT) with convolutional neural networks (CNNs) to solve the problem of small object detection. However, due to its quadratic complexity, this significantly increases the complexity of the algorithm, making it extremely challenging when applied to high-resolution images. Since remote sensing images are different from natural images and may be distributed in arbitrary spatial directions, traditional ViT methods
[15] (cutting large images into small blocks) will result in a significant loss of contextual information.
[0005] Compared to the Transformer, the recent Mamba model
[16] has the same powerful global context modeling capabilities, but with lower computational cost, and has significant advantages in both training and inference. Mamba is a new selective state-space model with advantages such as obtaining global receptive field, dynamic weighting, and linear computational complexity, which motivates us to apply it to vision tasks. However, we find that directly applying Mamba to remote sensing image target detection tasks leads to a decrease in detection performance. This is because remote sensing images contain a large number of small objects, and local image features need to be emphasized. However, when Mamba flattens the image into a sequence for processing, it also destroys the local dependencies between pixels.
[0006] References:
[0007] [1]Yang Y, Zang B, Song C, et al.Small object detection in remotesensing images based on redundant feature removal and progressive regression[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024.
[0008] [2]Liu Y, Li Q, Yuan Y, et al. ABNet: Adaptive balanced network for multiscale object detection in remote sensing imagery [J]. IEEE transactions ongeoscience and remote sensing, 2021, 60: 1-14.
[0009] [3]Zhang C,Su J,Ju Y,et al.Efficient inductive vision transformer fororiented object detection in remote sensing imagery[J].IEEE Transactions onGeoscience and Remote Sensing,2023,61:1-20.
[0010] [4]Zhao Z,Du J,Li C,et al.Dense tiny object detection:A scene contextguided approach and a unified benchmark[J].IEEE Transactions on Geoscienceand Remote Sensing,2024,62:1-13.
[0011] [5]Huang Y,Jiao D,Huang X,et al.A Hybrid CNN-Transformer Network forObject Detection in Optical Remote Sensing Images:Integrating Local andGlobal Feature Fusion[J].IEEE Journal of Selected Topics in Applied EarthObservations and Remote Sensing,2024.
[0012] [6]Xue J,He D,Liu M,et al.Dual network structure with interweavedglobal-local feature hierarchy for transformer-based object detection inremote sensing image[J].IEEE Journal of Selected Topics in Applied EarthObservations and Remote Sensing,2022,15:6856-6866.
[0013] [7]Zhao J,Jia Y,Ma L,et al.Adaptive dual-stream sparse transformernetwork for salient object detection in optical remote sensing images[J].IEEEJournal of Selected Topics in Applied Earth Observations and Remote Sensing,2024,17:5173-5192.
[0014] [8]Zhang Z,Lu X,Cao G,et al.ViT-YOLO:Transformer-based YOLO forobject detection[C] / / Proceedings of the IEEE / CVF international conference oncomputer vision.2021:2799-2808.
[0015] [9]Zhang C,Su J,Ju Y,et al.Efficient inductive vision transformer fororiented object detection in remote sensing imagery[J].IEEE Transactions onGeoscience and Remote Sensing,2023,61:1-20.
[0016]
[10] Li G,Bai Z,Liu Z,et al.Salient object detection in optical remotesensing images driven by transformer[J].IEEE Transactions on ImageProcessing,2023,32:5257-5269.
[0017]
[11] Mo N,Zhu R.A Novel Transformer-Based Object Detection Method WithGeometric and Object Co-Occurrence Prior Knowledge for Remote Sensing Images[J].IEEE Journal of Selected Topics in Applied Earth Observations and RemoteSensing,2024.
[0018]
[12] VaswaniA,Shazeer N,Parmar N,et al.Attention is all you need[J].
[0019] Advances in neural information processing systems,2017,30.
[0020]
[13] Guo Z,Bi G,Lv H,et al.Semantic Information Feature AggregationNetwork for Object Detection in Remote Sensing Images[J].IEEE Geoscience andRemote Sensing Letters,2024.
[0021]
[14] Zhu X,Lyu S,Wang X,et al.TPH-YOLOv5:Improved YOLOv5 based ontransformer prediction head for object detection on drone-captured scenarios[C] / / Proceedings of the IEEE / CVF international conference on computervision.2021:2778-2788.
[0022]
[15] Dosovitskiy A,Beyer L,Kolesnikov A,et al.An image is worth
[0023] 16x16 words:Transformers for image recognition at scale[J].arXivpreprint arXiv:2010.11929,2020.
[0024]
[16] Gu A, Dao T. Mamba: Linear-time sequence modeling with selectivestate spaces[J].arXiv preprint arXiv:2312.00752,2023. Summary of the Invention
[0025] To address the aforementioned technical issues, the present invention proposes a remote sensing image target detection method using a scale-sensitive Mamba network. A dual-branch structure, called the Remote Sensing Mamba Block (RSMB), is designed to improve performance in remote sensing target detection tasks. Specifically, the Remote Sensing Mamba Block consists of two modules: the Local Key Feature Perception Block (LKFPB) and the Long-Distance Dependency Building Block (LDMB). The network splits the feature map into two parts along the channel dimension. The LKFPB captures key local features in one part, while the LDMB models long-range dependencies in the other part to capture global features. Furthermore, the sparse feature fusion module fuses the shallow features of the Remote Sensing Mamba Block with the deep semantic features of the detection head to enhance the feature representation of small targets, thereby improving the model's target detection performance in remote sensing images. The present invention also proposes a bounding box loss function based on Intersection over Union (IoU)—the Area Intersection over Union (AIoU)—which makes the Scale-Sensitive Mamba Network (SSMNet) more sensitive to target scale. AIoU can fit the true bounding boxes of small objects more quickly and accurately during training, and achieves faster convergence when training on small object datasets.
[0026] The present invention proposes a remote sensing image target detection method using a scale-sensitive Mamba network, comprising preparing a training set and a test set of a DIOR dataset, and further comprising the following steps:
[0027] Step 1: Use the images in the training set to perform SSMNet network training to generate a training model;
[0028] Step 2: Save the training model to a local folder, use the images in the test set to test the effect of the training model, and if satisfactory, save the training model as a satisfactory training model;
[0029] Step 3: Save the satisfactory training model to a local folder, and use the satisfactory training model to test unlabeled remote sensing images.
[0030] Preferably, the SSMNet network includes a remote sensing Mamba block, a backbone network, a sparse feature fusion module, and an area intersection-over-union loss.
[0031] In any of the above solutions, preferably, the remote sensing Mamba block consists of a local key feature perception block and a long-distance building block.
[0032] In any of the above solutions, preferably, the local key feature perception block first uses a 3x3 kernel and 1-padding convolution to reduce the input channel dimension, followed by a batch normalized ReLU activation function. Polarized self-attention is then used to focus on local features. Finally, the input is convolved to obtain the output of the local key feature perception block.
[0033] In any of the above solutions, preferably, the long-distance dependency building module first uses layer normalization to process the input feature map F. Then, the input undergoes depthwise separable convolution, SiLU activation function, omnidirectional selective scanning module and layer normalization to generate global feature attention. This attention is multiplied with the input F′, then passes through a linear layer, and finally adds a residual connection to the input F.
[0034] In any of the above schemes, preferably, the sparse feature fusion module includes sparse local attention and remote sensing Mamba blocks, which effectively aggregate shallow and deep features and enhance the representation ability of small targets.
[0035] In any of the above solutions, preferably, the sparse feature fusion module first uses a 1×1 convolutional layer to adjust the channel dimensions of the shallow features Fs and the deep features Fd. Sparse local attention is then used to compute the sparse attention Fa of Fs and Fd, and batch normalization and SiLU activation functions are applied to Fa. The feature representation of Fa is then enhanced using a remote sensing Mamba block. Finally, the fused features are sent to the detection head.
[0036] In any of the above solutions, preferably, the area intersection over union (AIoU) loss is used to measure the pixel area difference between the predicted box and the true box.
[0037] SSMNet refers to Scale Sensitivity Mamba Network, which is a scale-sensitive Mamba network. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The figure is a flow chart of a preferred embodiment of a method for remote sensing image target detection using a scale-sensitive Mamba network according to the present invention.
[0039] Figure 2 Schematic diagram of the overall network structure of an SSMNet network according to an embodiment of the remote sensing image target detection method using the scale-sensitive Mamba network of the present invention.
[0040] Figure 3 The figure is a schematic diagram of the remote sensing Mamba module structure of the SSMNet network of the remote sensing image target detection method using the scale-sensitive Mamba network according to the present invention.
[0041] Figure 4 Schematic diagram of polarized self-attention of the SSMNet network of the remote sensing image target detection method using the scale-sensitive Mamba network according to the present invention.
[0042] Figure 5 Schematic diagram of the omnidirectional selective scanning module of the SSMNet network of the remote sensing image target detection method using the scale-sensitive Mamba network according to the present invention.
[0043] Figure 6 The figure is a schematic diagram of a sparse feature fusion module of an SSMNet network in a remote sensing image target detection method using a scale-sensitive Mamba network according to the present invention.
[0044] Figure 7 The figure is a schematic diagram of the bounding box regression step using the intersection-over-union loss of the SSMNet network of the remote sensing image target detection method using the scale-sensitive Mamba network according to the present invention.
[0045] Figure 8 Schematic diagram of defects in the loss function based on the intersection-over-union (IoU) of the SSMNet network of the remote sensing image target detection method using the scale-sensitive Mamba network according to the present invention. DETAILED DESCRIPTION
[0046] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0047] Example 1
[0048] like Figure 1 As shown, step 100 is executed to prepare the training set and test set of the DIOR dataset.
[0049] Execute step 110 to perform SSMNet network training using the images in the training set to generate a training model.
[0050] The SSMNet network includes a remote sensing Mamba block, a backbone network, a sparse feature fusion module, and an area intersection-over-union loss.
[0051] The remote sensing Mamba block consists of a local key feature perception block and a long-range dependency building block. The local key feature perception block first uses a 3x3 kernel and a padding of 1 convolution to reduce the input channel dimension, followed by a batch normalization ReLU activation function. Polarized self-attention is then used to focus on local features. Finally, the input is convolved to obtain the output of the local key feature perception block. The long-range dependency building block first processes the input feature map F using layer normalization. The input then undergoes depthwise separable convolution, SiLU activation function, omnidirectional selective scanning module, and layer normalization to generate global feature attention. This attention is multiplied with the input F′, then passed through a linear layer, and finally added with a residual connection to the input F.
[0052] The sparse feature fusion module first uses a 1×1 convolutional layer to adjust the channel dimensions of the shallow features Fs and the deep features Fd. Sparse local attention is then used to compute the sparse attention Fa of Fs and Fd. Batch normalization and SiLU activation are then applied to Fa. The feature representation of Fa is then enhanced using the remote sensing Mamba block. Finally, the fused features are sent to the detection head.
[0053] The Area Intersection over Union (AIoU) loss is used to measure the pixel area difference between the predicted box and the true box.
[0054] Execute step 120 to save the training model to a local folder, use the images in the test set to test the effect of the training model, and if satisfactory, save the training model as a satisfactory training model.
[0055] Execute step 130 to save the satisfactory training model to a local folder, and use the satisfactory training model to test unlabeled remote sensing images.
[0056] Example 2
[0057] The present invention discloses a trainable end-to-end remote sensing image target detection network named "ScaleSensitivity MambaNetwork" (SSMNet) using a scale-sensitive Mamba network.
[0058] The content of SSMNet includes remote sensing Mamba block, backbone network, sparse feature fusion module, and area intersection-over-union loss.
[0059] SSMNet trains the dataset to generate a training model. By incorporating the network hyperparameters from the trained model into the network model, it implements remote sensing image target detection, achieving the desired target detection effect. This effectively improves the model's target detection performance for remote sensing images.
[0060] Steps to use SSMNet (implemented in Python programming language):
[0061] 1. Prepare the training set and test set of the DIOR dataset, and set the image format to ".jpg" or ".png";
[0062] 2. Use the training file "train.py" to start network training and adjust network parameters such as batch size and lr as needed.
[0063] 3. Save the trained model to a local folder and use the "vali.py" file and the test set to test the effectiveness of the network training model. If you are not satisfied, you can use the "checkpoint" technology to continue training a satisfactory model.
[0064] 4. Save the satisfactory training model to a local folder and use "test.py" to test the unlabeled remote sensing image.
[0065] Example 3
[0066] Object detection in remote sensing images remains a significant challenge due to their complex geographic environments, insufficient representation of small objects, and large object scale variations. These characteristics can significantly degrade detector performance. To address these issues, we propose an efficient detector, the Scale-Sensitive Mamba Network (SSMNet).
[0067] The framework of SSMNet is summarized as follows Figure 2 As shown. We first use the backbone network to extract multi-scale features C2 to C5. Then, the Remote Sensing Mamba Block (RSMB) is applied to model the global context of the shallow feature C2 and focus on key local features. Next, the Sparse Feature Fusion Module (SFFM) is used to fuse the enhanced features with the deep semantic feature map P3. The SSMNet network includes the Remote Sensing Mamba Block, the backbone network, the Sparse Feature Fusion Module, and the Area Intersection over Union loss. The Remote Sensing Mamba Block (RSMB) can focus on key local features while modeling the global context of the remote sensing image. Sparse Feature Fusion Module (SFFM), which can aggregate shallow and deep feature information and enhance the expressive power of the features. Area Intersection over Union (AIoU) is used to accelerate the convergence of the detector on small target datasets, while enabling the detector to adapt to the scale changes of targets in remote sensing images.
[0068] The work of each module is as follows:
[0069] 1. Remote sensing Mamba module (such as Figure 3 shown)
[0070] The remote sensing Mamba block consists of a local key feature perception block (LKFPB) and a long-distance dependency building block (LDMB). First, we split the input feature map F into two equally sized sub-inputs F1 and F2 along the channel dimension. F1 focuses on key local features, while F2 is used for global context modeling.
[0071] In the Local Key Feature Perception Block (LKFPB), we first use a convolutional layer with a kernel of 3×3 and padding of 1 to reduce the input channel dimension, while using the local receptive field of the convolution to focus on the local features of the input, and then perform batch normalization and ReLU activation function processing. Next, we use polarized self-attention (PSA) to focus on local features. Polarized self-attention effectively utilizes high-resolution information through nonlinear and linear transformations and improves the performance of the model in key point detection. Figure 4 As shown in Figure 2, polarized self-attention consists of spatial attention and channel attention, and a high internal resolution is maintained in the calculation of both channel attention and spatial attention. Finally, the input passes through a convolutional layer to obtain the output of the local key feature perception block.
[0072] In the Long-Distance Dependency Building Block (LDMB), we first process the input F2 using layer normalization. The input then undergoes a depthwise separable convolution (DWConv), a SiLU activation function, an omnidirectional selective sweep module (OSSM), and layer normalization (LN) to model its long-range dependencies, thereby obtaining global feature attention. This attention is then multiplied by the input F2', passed through a linear layer, and finally added with a residual connection to the input F2.
[0073] The input is fed to the Omnidirectional Selective Scanning Module (OSSM) to further comprehensively extract information from the remote sensing image. Figure 5 As shown in the figure, OSSM uses the state space model (SSM) technique to achieve bidirectional selective scanning of remote sensing images in the horizontal, vertical, diagonal, and anti-diagonal directions. This method aims to enhance the global effective receptive field of the image in multiple directions and extract global spatial features from different angles. Since the distribution of targets in remote sensing images is random, this scanning method is more suitable for remote sensing images. It effectively solves the problem of arbitrary distribution of target locations in remote sensing images. Finally, we merge the outputs of the two modules along the channel dimension and pass them through a 1×1 convolutional layer, while using a residual strategy to add the input and output.
[0074] 2. Sparse feature fusion module (such as Figure 6 shown)
[0075] Shallow features contain more detailed image information and are characterized by high resolution, while deep features contain more semantic information. However, deep features are insufficient for the semantic representation of small objects. Therefore, aggregating shallow and deep features can enhance the semantic representation of small objects. A sparse feature fusion module is designed based on sparse local attention (SLA) and remote sensing Mamba block (RSMB) to effectively aggregate shallow and deep features to enhance the representation of small objects. First, a 1×1 convolution is used to adjust the channel dimension of the shallow features Fs and deep features Fd. Then, sparse local attention (SLA) is used to compute the sparse attention of Fs and Fd. An asymmetric feature extractor (AFE) is then used to obtain the asymmetric mapping A. After obtaining the sparse attention Fa calculated by sparse local attention (SLA), batch normalization and SiLU activation function are applied to Fa. Then, remote sensing Mamba block (RSMB) is used to enhance the feature representation of Fa, focusing on its key local information. Finally, the fused features are sent to the detection head.
[0076] 3. Area intersection ratio (such as Figure 7 shown)
[0077] Remote sensing target datasets contain a large number of tiny objects with very small pixel areas. In the early stages of training, detectors using the distributional focus loss tend to generate large predicted boxes due to high uncertainty. This results in significant pixel area discrepancies between the predicted and ground-truth boxes. Existing Intersection-over-Union (IoU)-based losses fail to account for this, resulting in slow fitting between the predicted and ground-truth boxes, ultimately impacting model convergence and accuracy. To address this issue, a penalty term is introduced based on the CIoU to measure the pixel area discrepancy between the predicted and ground-truth boxes. Theoretically, the pixel area penalty is only effective when the bounding box is very small. When both the ground-truth box and the predicted box are large, the pixel area penalty is ineffective. Therefore, a scaling factor γ is introduced to reduce the range of the bounding box's pixel area, which enhances the penalty's effectiveness. Inspired by CIoU, a balancing factor β is introduced to adjust the weighting between IoU and pixel area discrepancy. When the bounding box's IoU is high (indicating relatively accurate positioning), the impact of the pixel area discrepancy on the overall loss is reduced. On the contrary, when IoU is low, the effect of pixel area difference will be more significant, which helps the bounding box better adjust its pixel area to make it closer to the target. Unlike CIoU, β needs to calculate gradients to participate in training and learning. Area Intersection over Union (AIoU) can reduce the pixel area difference between the anchor box and the target box more quickly. Figure 8As shown in Figure 1, when the center points of the predicted box and the target box overlap, the values of IoU, CIoU, and GIoU are exactly the same. However, due to the significant difference in aspect ratio between the predicted box and the target box, EIoU incurs a large penalty, resulting in negative values. AIoU effectively measures the difference in pixel area between the predicted box and the target box without imposing a large penalty that could result in negative IoU values. It also addresses the issue of significant object scale variations in remote sensing images.
[0078] In order to better understand the present invention, the above is described in detail in conjunction with the specific embodiments of the present invention, but it is not intended to limit the present invention. Any simple modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention. Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
Claims
1. A method for remote sensing image target detection using a scale-sensitive Mamba network, comprising preparing a training set and a test set of a DIOR dataset, characterized in that: The following steps are also included: Step 1: Use the images in the training set to perform SSMNet network training to generate a training model; Step 2: Save the training model to a local folder, use the images in the test set to test the effect of the training model, and if satisfactory, save the training model as a satisfactory training model; Step 3: Save the satisfactory training model to a local folder, and use the satisfactory training model to test unlabeled remote sensing images.
2. The method for remote sensing image target detection using a scale-sensitive Mamba network according to claim 1, wherein: The SSMNet network includes a remote sensing Mamba block, a backbone network, a sparse feature fusion module, and an area intersection-over-union loss.
3. The method for remote sensing image target detection using a scale-sensitive Mamba network according to claim 2, wherein: The remote sensing Mamba block consists of a local key feature perception block and a long-distance dependency building block.
4. The method for remote sensing image target detection using a scale-sensitive Mamba network according to claim 3, wherein: The local key feature perception block first uses a 3x3 kernel and 1-padding convolution to reduce the input channel dimension, followed by a batch normalized ReLU activation function. Polarized self-attention is then used to focus on local features. Finally, the input is convolved to produce the output of the local key feature perception block.
5. The method for remote sensing image target detection using a scale-sensitive Mamba network according to claim 4, wherein: The long-distance dependency building module first uses layer normalization to process the input feature map F. Then, the input undergoes depthwise separable convolution, SiLU activation function, omnidirectional selective scanning module and layer normalization to generate global feature attention. This attention is multiplied by the input F′, then passes through a linear layer, and finally adds a residual connection to the input F.
6. The method for remote sensing image target detection using a scale-sensitive Mamba network according to claim 5, wherein: The sparse feature fusion module contains sparse local attention and remote sensing Mamba blocks, which effectively aggregates shallow and deep features and enhances the representation ability of small objects.
7. The method for remote sensing image target detection using a scale-sensitive Mamba network according to claim 6, wherein: The sparse feature fusion module first uses a 1×1 convolutional layer to adjust the channel dimensions of the shallow features Fs and the deep features Fd. Sparse local attention is then used to compute the sparse attention Fa of Fs and Fd. Batch normalization and SiLU activation are then applied to Fa. The feature representation of Fa is then enhanced using the remote sensing Mamba block. Finally, the fused features are sent to the detection head.
8. The method for remote sensing image target detection using a scale-sensitive Mamba network according to claim 7, wherein: The area intersection over union (AIoU) loss is used to measure the pixel area difference between the predicted box and the true box.
Citation Information
Cited By
Truck connecting ball head fault detection method and device, electronic equipment and storage medium
CN121582141A