Target detection method based on Mama feature fusion

Through the Mamba-R network and MAFE encoder combined with VSSA and MTMHSA modules, the problem of insufficient feature fusion of existing object detection algorithms under complex backgrounds and multi-scale targets is solved, and more efficient and accurate object detection is achieved, suitable for applications with high-resolution images and real-time performance requirements.

CN120298667APending Publication Date: 2025-07-11CHONGQING UNIV OF TECH
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510388177.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When existing object detection algorithms deal with complex backgrounds and multi-scale targets, it is difficult to effectively integrate features of different scales, resulting in insufficient recognition capabilities. The traditional hybrid architecture fails to fully utilize the respective advantages of CNN and Transformer, which limits the deployment of computer vision technology in application scenarios with high real-time and accuracy requirements.

Method used

The object detection method based on Mamba feature fusion is adopted, and the Mamba-R network and MAFE encoder are combined with VSSA and MTMHSA modules to enhance spatial feature expression and global context information capture capabilities, and optimize the feature extraction and fusion process.

Benefits of technology

It improves the accuracy and efficiency of object detection, especially the ability to identify targets at different scales in complex contexts, optimizes the computing efficiency, and is suitable for application scenarios with high-resolution images and real-time performance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298667A_ABST
    Figure CN120298667A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method based on Mama feature fusion, and relates to the technical field of image target detection. According to the method, the innovative implementation of the VSSA module is utilized, a selective scanning mechanism of the state space model is applied to 2D visual data processing, the long-distance dependency relationship in the image is effectively captured through state space modeling in four directions, the limitation of a traditional state space model in the two-dimensional visual data processing process is solved through the multi-direction processing strategy, and the processing precision of the 2D visual data is improved. The model can comprehensively perceive spatial dependency relationships in different directions in an image, the VSSA adopts learnable state space parameters to dynamically model a feature sequence, the ability of the network to understand a complex space structure is enhanced, the method is particularly suitable for processing scenes needing long-distance context information, and in addition, the method is combined with MTMHSA, so that the complexity of the network is reduced. And the fusion capability of different levels of features in target detection is further enhanced. Through the innovation, the model can better understand the target in the image, and the positioning and classification precision of the target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image target detection, and in particular to an object detection method based on Mamba feature fusion. Background Art

[0002] Object detection is one of the core tasks in the field of computer vision, aiming to classify and locate target objects in an image simultaneously. Object detection is widely used in various vision tasks, such as autonomous driving, security monitoring, intelligent healthcare, etc. However, with the continuous improvement of application requirements, especially when facing complex backgrounds and targets of different scales, existing object detection algorithms still face many challenges.

[0003] Traditional object detection methods are mainly based on convolutional neural networks (CNNs), which use convolutional operations to extract image features. These methods perform excellently in the detection of large targets. However, due to the limitation of the local receptive field of the convolutional kernel, it is difficult to effectively capture the global context information of the image, thus affecting the detection ability of targets and complex scenes. In addition, when performing multi-scale object detection, due to its structural limitations, it is difficult for CNN methods to efficiently fuse features of different scales, resulting in insufficient recognition ability for targets of different scales.

[0004] To address these problems, in recent years, object detection methods based on Transformer (such as DETR, RT-DETR) have introduced the self-attention mechanism, which can model features globally and overcome the deficiencies of convolutional neural networks in global feature extraction. The self-attention mechanism of Transformer enables the network to perform feature interaction within a larger receptive field, improves the modeling of global context information, and enables object detection to more flexibly handle targets of different sizes and different backgrounds.

[0005] Moreover, the RT-DETR detection algorithm has the following problems in the application process: while reducing the number of model parameters, it fails to fully ensure the ability to extract global context information;

[0006] In addition, although the latest state space models (such as Mamba) provide an efficient property with linear expansion O(n) with respect to the sequence length, they are mainly designed for sequence modeling tasks, not specifically optimized for 2D visual data, and lack a dedicated mechanism for processing image spatial information. Traditional state space models usually process image data unidirectionally row by row or column by column, unable to consider the spatial relationships in both horizontal and vertical directions simultaneously, which results in poor performance when dealing with visual data with complex two-dimensional structures. Although such models perform excellently in sequence tasks such as NLP, they face adaptability challenges when directly migrated to the field of computer vision.

[0007] Current hybrid architecture attempts usually simply stack different modules, failing to effectively integrate their respective advantages, resulting in an unsatisfactory balance between model accuracy and speed. For example, some works attempt to combine CNN with Transformer, but often simply add Transformer layers on top of the CNN feature extractor. This design fails to fully utilize the respective advantages of CNN and Transformer for complementarity. Similarly, attempts to apply state space models to visual tasks mostly remain at a superficial architecture stacking level, lacking specialized design and optimization for 2D spatial information. This situation is particularly evident in scenarios such as processing high-resolution images or applications requiring real-time performance, severely restricting the deployment and use of computer vision technology in many practical applications, such as autonomous driving, video surveillance, and industrial inspection, which have high requirements for real-time performance and accuracy. Therefore, a new solution to the above problems is needed. Summary of the Invention

[0008] The purpose of the present invention is to provide an object detection method based on Mamba feature fusion. By designing the backbone network Mamba-R and the encoder MAFE, the feature extraction and fusion process is optimized, thereby improving the accuracy and efficiency of object detection to solve the technical problems proposed in the background art.

[0009] To achieve the above object, the present invention provides the following technical solution: An object detection method based on Mamba feature fusion, at least including the following steps:

[0010] Step 1: Data preprocessing. Collect image data and perform data preprocessing on the image data before inputting it into the model.

[0011] Step 2: Feature extraction. Use the innovative backbone network Mamba-R to extract features from the preprocessed image to obtain feature maps of multiple scales. The Mamba-R network includes a VSSA (Visual Selective State-Space Attention) module, that is, a visual spatial feature enhancement module. The VSSA module is used to enhance visual spatial features, further improve the spatial information expression ability, and effectively improve the recognition ability of the object detection model for objects in complex scenes.

[0012] Step 3: Feature fusion. The Multiscale Attention Fusion Encoder, i.e., the MAFE encoder, is used to perform feature fusion on the shallow semantic feature information and deep semantic feature information of the feature map. The MAFE (Multiscale Attention Fusion Encoder) encoder performs deep semantic feature interaction by combining the Mixed-Topology Multi-Head Self-Attention (MTMHSA) module, effectively improving the feature fusion ability and enabling accurate capture of spatial and context information at different scales.

[0013] Step 4: Feature decoding. Multiple stacked decoders are used to decode the features output by the MAFE encoder to obtain a feature sequence, and the feature sequence is input into the prediction head for prediction, outputting the predicted box coordinates and predicted classes.

[0014] Step 5: Loss function correction. The outputs of the multiple stacked decoders are collected, and then the loss function is used to calculate the gradients for each decoder's output head and adjust the parameters.

[0015] Step 6: Object detection. Based on the corrected object detection model based on Mamba feature fusion obtained from Step 1 to Step 5, the object detection model is used to perform object detection on the image to be detected.

[0016] Further, the Step 1 at least includes the following steps:

[0017] First, perform normalization processing to standardize the pixel values to a unified range, ensuring that the image data can be effectively transmitted to the neural network, mapping the pixel values to a unified range to avoid the impact of different image pixel value ranges on the training process.

[0018] Next, apply a series of random transformations to the image data for data augmentation. The random transformations at least include random color perturbation, size expansion, cropping, and horizontal flipping. The data augmentation method is used to increase sample diversity and improve the model's adaptability in different scenarios and angles.

[0019] Finally, perform unified size adjustment on the image, adjusting the image data to a fixed size to ensure that the image meets the input requirements of the network, so as to be input into the model for subsequent processing and help reduce unnecessary computational overhead.

[0020] Through the above processing steps, it is ensured that the input image information is more representative and diverse, thereby improving the training efficiency and accuracy of the model.

[0021] Furthermore, the Mamba-R network is stacked by multiple Residual Blocks, which can effectively extract the basic feature information in the image and generate feature maps of multiple scales. These feature maps represent the information of the image at different levels, with different spatial scales and depth features. The feature maps include three scale feature maps with the number of channels being 128, 256, and 512 respectively, namely S1, S2, and S3.

[0022] The VSSA module is embedded into the backbone network, aiming to enhance the spatial feature expression ability of the image. The VSSA module adopts a 2D selective scanning mechanism based on Mamba. Through a four-way processing strategy and state space modeling, it effectively captures the long-range spatial dependency relationship in the image and improves the feature expression ability.

[0023] By combining the Mamba-R and VSSA modules, three scale feature maps are extracted, namely S1, S2, and S3, which respectively represent the spatial information of the image at different levels. The three scale feature maps provide rich context information for subsequent feature fusion and object detection, enhancing the model's perception ability of multi-scale objects. The example steps are as follows:

[0024] After the input feature map S3 is passed to the VSSA module, it enters the SS2D layer in the VSSA module. See the following formula:

[0025] S3 = LayerNorm(S3) S4 = input + DropPath(SS2D(S3))

[0026] The SS2D layer generates the enhanced spatial feature S4 through dilated convolution operations, self-attention mechanism processing, and selective scanning. This feature map contains richer spatial information, which can help the model better capture the spatial dependency of the target. The finally generated enhanced feature S4 will be fused with the feature maps of other layers for subsequent feature fusion, thereby improving the model's recognition ability for small targets in complex scenes.

[0027] Furthermore, the MTMHSA module combines multi-scale convolution and multi-head self-attention mechanism, aiming to enhance the model's processing ability for different scale features and capture global context information. The MTMHSA module first extracts the multi-scale information of the input feature map through a multi-scale convolution module, namely the MSC (Multiscale Spatial Convolution) module, and then performs weighted calculation on the features through the multi-head self-attention mechanism to capture the global spatial dependency relationship. Finally, the output is the feature map after self-attention processing, representing the enhanced version of the input feature.

[0028] Furthermore, in the MAFE encoder, the feature map undergoes deep semantic feature interaction through the MTMHSA module to generate deep semantic information. Specifically:

[0029] The extracted feature map is fed into the MAFE encoder and then enters the MTMHSA module;

[0030] Input feature map processing: In the MTMHSA module, first, the input feature map x is processed by the MSC module using multiple convolutional branches with different dilation rates (dilation = [3, 5, 7]) to extract information at different scales, resulting in a feature map kv containing information at different scales;

[0031] kv = MSC(x)

[0032] Multi-scale convolution and feature rearrangement: Dilated convolution helps capture a wider range of spatial information, ensuring that the network can handle a broader global context. The feature map kv obtained through multi-scale convolution and the original feature map x are rearranged into a format suitable for calculation, as shown in the following formula:

[0033] Q = rearrange(x)

[0034] K, V = rearrange(kv)

[0035] Then, the adjusted feature map enters the multi-head self-attention mechanism of the MTMHSA module. Through the calculation between the query Q, key K, and value V, the spatial information and context information in the feature map are further enhanced; the dot product operation (Q·K T ) is calculated in the multi-head attention mechanism, and finally, through weighted averaging, that is, using the softmax function, the fused feature information is obtained. This process is expressed as follows:

[0036]

[0037] where d is the feature dimension of each attention head, used as a scaling factor to prevent the attention dot product from having an overly large value in high dimensions, which may lead to gradient vanishing and ensure training stability;

[0038] c_attn = sigmoid(FC(AvgPool(x)))·x

[0039] At the same time, the channel attention c_attn is calculated, and the channel dimension is weighted through global average pooling and the FC layer:

[0040] output = attn + c_attn

[0041] Add the attention result attn to the channel attention c_attn, which helps to retain the input information while introducing enhanced features processed by the attention mechanism, and obtain the fused feature map of the output:

[0042] Through the MAFE encoder, the shallow and deep features are effectively fused and strengthened, thereby improving the feature representation ability.

[0043] Furthermore, the feature decoding combines the output of the encoder with the target query, and after being processed by multiple Transformer decoder layers, generates the final target detection result. The decoding process of the decoder includes a self-attention mechanism and a cross-attention mechanism;

[0044] First, update the information of the target query through self-attention;

[0045] Then perform cross-attention calculation with the reference point to further enhance the association between the target position and the features;

[0046] Next, perform a non-linear mapping on the features through a feed-forward neural network;

[0047] Finally, obtain the updated target query features;

[0048] During the training process, the decoder will calculate the predicted box coordinates and predicted classes to obtain the bounding box position of the target and the corresponding class labels;

[0049] The output of the decoder includes the predicted box coordinates and predicted classes;

[0050] In addition, if the auxiliary loss is enabled, multiple outputs in the decoder will be used to assist the optimization during the training process. The decoder also supports denoising training, which helps to improve the robustness of the model, especially when facing noisy data in the detection task;

[0051] Finally, the output of the decoder will pass through a non-linear activation function to obtain the final bounding box and classification results.

[0052] Furthermore, the loss function in S5 is as follows:

[0053] Classification loss (VFL loss):

[0054]

[0055] where p i,c is the logits value of the predicted class, q i is the IoU value between the predicted box and the ground truth box, t i,c is the one-hot encoding of the class, w i,c = α·σ(pi ,c )γ ·(1 - t i,c ) + q i ·t i,c is the dynamic weight, α and γ are the adjustment coefficients of the focal loss, mean c represents taking the average over all classes, N is the total number of samples, and BCE represents the binary cross-entropy loss;

[0056] Bounding box regression loss (L1 Loss + GIoU Loss):

[0057]

[0058] where B i and are the coordinates of the ground truth box and the predicted box respectively, ‖.‖1 represents the L1 norm; the absolute error; is the GIoU value between the ground truth box and the predicted box;

[0059] Overall expression of the loss function:

[0060] L total = λ VFL L VFL + λ L1 L L1 + λ GIoU L GIoU

[0061] where λ VFL , λ L1 and λ GIoU are the adjustment coefficients.

[0062] Compared with the prior art, the beneficial effects of the present invention are:

[0063] The present invention realizes the innovation of the VSSA (Visual Selective State-Space Attention) module, applies the selective scanning mechanism of the state space model to 2D visual data processing, and effectively captures the long-distance dependencies in the image through state space modeling in four directions (horizontal, vertical, and their flipped directions). This multi-directional processing strategy solves the limitations of traditional state space models in processing two-dimensional visual data, enabling the model to comprehensively perceive the spatial dependencies in different directions of the image. VSSA dynamically models the feature sequence using learnable state space parameters (A, B, C, D), enhancing the network's ability to understand complex spatial structures and being particularly suitable for processing scenarios that require long-distance context information.

[0064] In addition, the present invention also incorporates MTMHSA, further enhancing the fusion ability of features at different levels in object detection. Through this innovation, the model can better understand the objects in the image, improving the localization and classification accuracy of the objects. The innovative Mixed-Topology Multi-Head Self-Attention (MTMHSA) module adopts convolutional branches with different dilation rates [3, 5, 7] to effectively expand the receptive field and capture multi-scale features. At the same time, it combines an adaptive average pooling layer and a dual attention (spatial attention and channel attention) mechanism, not only optimizing the computational structure but also significantly enhancing the model's ability to recognize objects of different scales, enabling the network to more comprehensively understand the image context information. Therefore, the object detection method based on Mamba feature fusion proposed by the present invention not only improves the performance of object detection in multi-scale and complex backgrounds but also optimizes the computational efficiency, enabling the method to provide more efficient and accurate object detection capabilities in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a structural block diagram of a small object detection model based on Mamba feature fusion in the present invention;

[0066] Figure 2 It is a structural schematic diagram of the VSSA module in the present invention;

[0067] Figure 3 It is a structural schematic diagram of the MMHSA module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0069] An object detection method based on Mamba feature fusion can effectively improve the detection accuracy of objects in an image, especially the detection performance of objects of different scales in a complex background.

[0070] Step 1, data preprocessing; collect image data and perform data preprocessing on the image data before inputting it into the model.

[0071] Step 1 at least includes the following steps:

[0072] First, perform normalization processing to standardize the pixel values to a unified range, ensuring that the image data can be effectively transmitted to the neural network. Map the pixel values to a unified range (such as [0, 1] or [-1, 1]) to avoid the impact of differences in the pixel value ranges of different images on the training process;

[0073] Next, a series of random transformations are applied to the image data for data augmentation. The random transformations at least include random color perturbations, size expansion, cropping, and horizontal flipping. The data augmentation method is used to increase sample diversity and improve the adaptability of the model in different scenarios and angles;

[0074] Finally, the images are uniformly resized to a fixed size to ensure that the image data can meet the input requirements of the network, so as to be input into the model for subsequent processing and help reduce unnecessary computational overhead;

[0075] Through the above processing steps, it is ensured that the input image information is more representative and diverse, thereby improving the training efficiency and accuracy of the model.

[0076] Step 2, feature extraction; an innovative backbone network Mamba-R is used to extract features from the preprocessed images, obtaining feature maps of multiple scales. The Mamba-R network includes a VSSA module, namely a visual spatial feature enhancement module. The VSSA module is used to enhance visual spatial features, further improving the spatial information expression ability and effectively enhancing the recognition ability of the object detection model for small objects in complex scenes.

[0077] The Mamba-R network is stacked by multiple residual blocks, effectively extracting the basic feature information in the images and generating feature maps of multiple scales. These feature maps represent the information of the images at different levels, with different spatial scales and depth features, including three-scale feature maps with channel numbers of 128, 256, and 512 respectively, namely S1, S2, and S3;

[0078] The VSSA module, as an innovative part of the present invention, is embedded in the backbone network to enhance the spatial feature expression ability of the images, especially the correlation between shallow and deep features; through the combination of Mamba-R and the VSSA module, three-scale feature maps are extracted, namely S1, S2, and S3, which represent the spatial information of the images at different levels respectively. The three-scale feature maps provide rich context information for subsequent feature fusion and object detection, enhancing the model's perception ability for multi-scale objects. The example steps are as follows:

[0079] After the input feature map S3 is passed to the VSS module, the overall process is as follows:

[0080] S3 = LayerNorm(S3) S4 = input + DropPath(SS2D(S3))

[0081] The SS2D layer generates enhanced spatial feature S4 through selective scanning. This feature map contains richer spatial information, which can help the model better capture the spatial dependence of the target. The finally generated enhanced feature S4 will be fused with the feature maps of other layers for subsequent feature fusion, thereby improving the model's recognition ability for small targets in complex scenes.

[0082] The operations of the VSSA module (refer to Figure 2 ) include layer normalization, SS2D module processing, and downsampling;

[0083] First, perform layer normalization on the input feature map, normalizing the input features of each layer so that they fluctuate within a certain range, which helps to accelerate training and improve model stability;

[0084] x norm = LayerNorm(x)

[0085] where x represents the input feature map;

[0086] Then send the feature map into the SS2D module for processing, using four-way selective scan and state space model to enhance the spatial information of the input feature map;

[0087] Processing the input feature map through the state space model can capture a larger range of spatial information in the image, especially enhancing the global context awareness in low-level features; using parameterized state space parameters (A, B, C, D), the model's attention to important regions of the image is improved in a learnable way, enabling the network to focus more on the target region and ignore background noise; selective scanning in four directions (horizontal, vertical, and their flips) is used to enhance spatial context information, which helps to capture more relevant spatial information in the feature map, enabling the model to more accurately identify and distinguish different targets in complex scenes;

[0088] x enhanced = SS2D(x norm )

[0089] Step three, feature fusion. The Multiscale Attention Fusion Encoder, namely the MAFE encoder, is used to perform feature fusion on the shallow semantic feature information and deep semantic feature information of the feature map. The MAFE encoder effectively improves the feature fusion ability by combining the multi-scale multi-head self-attention mechanism module, namely the MTMHSA module, and can accurately capture spatial and context information at different scales.

[0090] The MTMHSA module combines multi-scale convolution and multi-head self-attention mechanisms, aiming to enhance the model's ability to process features at different scales and capture global context information; the MTMHSA module first extracts multi-scale information of the input feature map through a multi-scale convolution module, namely the MSC module, and then performs weighted calculation on the features through the multi-head self-attention mechanism to capture global spatial dependencies. Finally, the output is the feature map after self-attention processing, representing an enhanced version of the input features.

[0091] In the MAFE encoder, the feature map undergoes deep semantic feature interaction through the MTMHSA module to generate deep semantic information. Specifically:

[0092] The extracted feature map is fed into the MAFE encoder and then enters the MTMHSA module;

[0093] Input feature map processing: In the MTMHSA module, first, the input feature map x is processed using multiple convolutional branches with different dilation rates to extract information at different scales, resulting in a feature map kv containing information at different scales;

[0094] kv = MSC(x)

[0095] Multi-scale convolution and feature rearrangement: Dilated convolution helps capture a larger range of spatial information, ensuring that the network can handle a wider global context. The feature map kv obtained through multi-scale convolution and the original feature map x are rearranged into a format suitable for calculation, as shown in the following formula:

[0096] Q = rearrange(x)

[0097] K, V = rearrange(kv)

[0098] Then the adjusted feature map enters the multi-head self-attention mechanism of the MTMHSA module. Through the calculation between the query Q, key K, and value V, the spatial information and context information in the feature map are further enhanced; the dot product operation (Q·K T ) is calculated in the multi-head attention mechanism, and finally, through weighted averaging, that is, using the softmax function, the fused feature information is obtained. This process is expressed as follows:

[0099]

[0100] where d is the feature dimension of each attention head, used as a scaling factor to prevent the attention dot product from having an overly large value in high dimensions, which may lead to gradient disappearance and ensure training stability;

[0101] c_attn = sigmoid(FC(AvgPool(x)))·x

[0102] Meanwhile, calculate the channel attention c_attn, and weight the channel dimension through global average pooling and the FC layer:

[0103] output = attn + c_attn

[0104] Add the attention result attn and the channel attention c_attn, which helps to retain the input information and at the same time introduce enhanced features processed by the attention mechanism to obtain the fused feature map of the output:

[0105] Through the MAFE encoder, shallow and deep features are effectively fused and enhanced, thereby improving the feature representation ability.

[0106] Generally speaking, during the processing, the feature map is further enhanced by MTMHSA. Here, the MSC module is used, and dilated convolution and pooling operations are added to further improve the feature representation ability. The purpose of the pooling operation is to perform global information aggregation on the feature map, thereby reducing the computational complexity to a certain extent while enhancing the expression ability of the feature map.

[0107] The entire process successfully fuses feature maps of different scales through the combination of the multi-layer self-attention mechanism and the convolutional module, enabling the network to capture rich spatial context information. This multi-scale fusion strategy enhances the network's perception ability of various targets, especially in object detection under complex backgrounds.

[0108] The fused feature map will be passed to the subsequent decoder module for predicting the object detection task. In the decoding stage, the network will process the fused feature map and finally predict the bounding box and class information of the target.

[0109] Step 4: Feature decoding. Use multiple stacked decoders to decode the features output by the MAFE encoder to obtain a feature sequence, and input the feature sequence into the prediction head for prediction, outputting the predicted box coordinates and predicted classes;

[0110] Feature decoding combines the output of the encoder with the target query, and after being processed by multiple layers of Transformer decoder layers, generates the final object detection result. The decoding process of the decoder includes the self-attention mechanism and the cross-attention mechanism;

[0111] First, update the information of the target query through self-attention;

[0112] Then perform cross-attention calculation with the reference point to further enhance the association between the target position and the features;

[0113] Next, the features are non-linearly mapped through a feed-forward neural network (FFN);

[0114] Finally, the updated target query features are obtained;

[0115] During the training process, the decoder calculates the predicted bounding box coordinates and predicted classes to obtain the bounding box position of the target and the corresponding class label;

[0116] The output of the decoder includes the predicted bounding box coordinates and predicted classes;

[0117] In addition, if the auxiliary loss is enabled, multiple outputs in the decoder are used to assist in the optimization during the training process. The decoder also supports denoising training, which helps to improve the robustness of the model, especially when facing noisy data in detection tasks;

[0118] Finally, the output of the decoder will pass through non-linear activation functions such as Sigmoid and Softmax to obtain the final bounding box and classification results.

[0119] Step Five, loss function correction. The outputs of multiple stacked decoders are collected, and then the gradient of each decoder's output head is calculated through the loss function to adjust the parameters;

[0120] The loss function correction optimizes the performance of the model by calculating the class loss, box regression loss, and auxiliary loss. The class loss uses Focal Loss to address the class imbalance problem and focuses on difficult-to-classify samples. The box regression loss consists of L1 loss and Generalized Intersection over Union (GIoU) loss to optimize the accuracy and coverage of the target box.

[0121] During the training process, the Hungarian algorithm is used to match the targets with the predicted bounding boxes. Then, based on the matching results, the class loss and box regression loss for each pair of target and prediction are calculated. If denoising training (CDN) is used, the denoising loss is also calculated additionally to improve the robustness of the model. All loss terms are weighted to ensure the balanced impact of different losses on the model training.

[0122] VFL loss (classification loss):

[0123]

[0124] where p i,c is the logits value of the predicted class, q i is the IoU value between the predicted bounding box and the ground truth bounding box, t i,c is the one-hot encoding of the class, w i,c = α · σ(p i,c ) γ · (1 - t i,c ) + q i · ti,c is the dynamic weight, α and γ are the adjustment coefficients of the focal loss, and mean c represents taking the average over all classes, N is the total number of samples, and BCE represents the binary cross-entropy loss.

[0125]

[0126] Among them, B i and are the coordinates of the ground truth box and the predicted box respectively, and ‖.‖1 represents the L1 norm; the absolute error; is the GIoU value between the ground truth box and the predicted box;

[0127] The overall expression of the loss function:

[0128] L total = λ VFL L VFL + λ L1 L L1 + λ GIoU L GIoU

[0129] Among them, λ VFL 、λ L1 and λ GIoU are the adjustment coefficients.

[0130] Step 6: Object detection, using the calibrated object detection model based on Mamba feature fusion (see Figure 1 ), perform object detection on the image to be detected;

[0131] When the inputs are the basic images F1, F2, and F3, the specific operations of the object detection model based on Mamba feature fusion are as follows:

[0132] Input feature map processing

[0133] S1, S2, S3 = VSSA(F1), VSSA(F2), VSSA(F3)

[0134] The feature maps S1, S2, and S3 are obtained by extracting features from F1, F2, and F3 through the VSSA module;

[0135] Feature flattening:

[0136] S1′, S2′, S3′ = Flatten(S1), Flatten(S2), Flatten(S3)

[0137] Each feature map is transformed into a one-dimensional vector through the Flatten operation to obtain the flattened feature maps S1′, S2′, and S3′;

[0138] Select query

[0139] SS = QuerySelection(Concat(S1′, S2′, S3′))

[0140] QuerySelection represents the IoU-aware query module. The flattened feature maps S1′, S2′, and S3′ are concatenated through the IoU-aware query module to generate a query vector SS for subsequent attention mechanisms;

[0141] Multi-scale feature extraction

[0142] KV = MultiscaleFeature(S1, S2, S3)

[0143] The query vector SS is flattened to obtain Q

[0144] Q = Flatten(SS)

[0145] The relationship between the query Q and K is calculated through dot product. Softmax is used to generate attention weights, and weighted summation is performed with the value V to obtain the features after multi-head self-attention processing;

[0146]

[0147] Feature fusion

[0148] SS″ = Hybrid(F MMHSA )

[0149] Among them, Hybrid is feature fusion. The feature representation is enhanced through the TransformerEncoder based on MTMHSA, and the bidirectional feature fusion is realized by combining the FPN and PAN structures. The top-down path and the bottom-up path complement each other to effectively integrate multi-scale feature information;

[0150] Object detection

[0151] Output box = Boxhead(SS″)

[0152] Output class = Classhead(SS″)

[0153] Among them, Boxhead and Classhead represent the coordinate prediction head and the class prediction head.

[0154] The innovative content in Step 2 and Step 3 is mainly reflected in:

[0155] The innovative implementation of the VSSA (Visual Selective State-Space Attention) module applies the selective scanning mechanism of the state space model to 2D visual data processing. Through state space modeling in four directions (horizontal, vertical, and their flipped directions), it effectively captures long-range dependencies in images. This multi-directional processing strategy addresses the limitations of traditional state space models in processing two-dimensional visual data, enabling the model to comprehensively perceive spatial dependencies in different directions in the image. VSSA dynamically models the feature sequence using learnable state space parameters (A, B, C, D), enhancing the network's ability to understand complex spatial structures and being particularly suitable for scenarios that require long-range context information.

[0156] The Mixed-Topology Multi-Head Self-Attention (MTMHSA) uses convolutional branches with different dilation rates [3, 5, 7] in its design, effectively expanding the receptive field to capture multi-scale features. At the same time, it combines an adaptive average pooling layer and a dual attention (spatial attention and channel attention) mechanism, not only optimizing the computational structure but also significantly enhancing the model's ability to recognize targets at different scales, enabling the network to more comprehensively understand image context information.

[0157] To sum up:

[0158] The object detection model based on Mamba feature fusion is trained using the above method. The model includes an innovative Mamba-R backbone network, a MAFE encoder, and a decoding module stacked by multiple decoders. The entire process of the object detection model is divided into three stages: feature extraction, feature fusion, feature decoding, and prediction.

[0159] In the feature extraction stage, the input image is subjected to feature extraction through the Mamba-R backbone network. The backbone network enhances the features by utilizing the VSSA (Visual Spatial Feature Enhancement) module. Based on the principle of the state space model, this module achieves efficient long-range dependence modeling through four-way selective scanning, enabling the network to have a stronger ability to capture global context information, thereby enhancing the feature representation ability. In the feature fusion stage, the input features are first processed by the MSC (Multi-Scale Convolution) component. The MSC consists of three parallel convolutional branches that use convolutional operations with different dilation rates [3, 5, 7], capable of expanding the receptive field without increasing the number of parameters and effectively capturing multi-scale context information. The processed features are dimensionally reduced through adaptive average pooling to generate key-value pairs (K, V). After calculation by the attention mechanism, the MTMHSA also introduces a channel attention branch that generates channel weights through global average pooling and non-linear transformation, and adds them to the spatial attention result to form the final output. This dual attention design focuses on both "what" (channels) and "where" (space), significantly enhancing the model's perception ability for targets at different scales and its understanding ability of complex scenes. In the feature decoding stage, the fused feature sequence is input into the stacked decoder for decoding, and finally, the target position and category are predicted through the prediction head to obtain the final output.

[0160] Experimental verification:

[0161] Based on the content of the above invention, an evaluation of the technical effects is proposed:

[0162] Performance evaluation: The present invention was experimentally evaluated on the VisiWear dataset. The VisiWear dataset contains 9,026 high-resolution images, including 7,220 training images, 902 validation images, and 904 test images. The dataset covers two object categories: people wearing reflective vests and people not wearing reflective vests. The present invention used 7,220 images from the training set and 904 images from the test set for training and testing respectively.

[0163] The present invention was compared with the baseline RT-DETR V2 algorithm. Table 1 shows the detection results on the VisiWear test set.

[0164] Compared with the baseline RT-DETR V2 algorithm, the method of the present invention not only has a more obvious improvement in accuracy under the condition of a small floating reduction in the number of parameters and computational complexity, with the mAP increased by 1% and the AP50 increased by 0.8%

[0165] Table 1 Comparative experiment results on the VisiWear dataset

[0166]

[0167] Based on the traditional object detection framework, the present invention enhances the feature extraction and fusion capabilities of the model in object detection and improves the detection accuracy and efficiency of the model by integrating innovative modules such as Mixed-Topology Multi-Head Self-Attention (MTMHSA) and Visual Selective State-Space Attention (VSSA).

[0168] As described above, it is only used to help understand the method and its core concept of the present invention, but the protection scope of the present invention is not limited thereto. For those of ordinary skill in the art in the technical field disclosed by the present invention, any equivalent replacement or change made according to the technical solution and inventive concept of the present invention should be covered within the protection scope of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A target detection method based on Mamba feature fusion, characterized in that: At least include the following steps: Step 1: Data preprocessing. Collect image data and perform data preprocessing on the image data before inputting it into the model. Step 2: Feature extraction. Use the innovative backbone network Mamba-R to extract features from the preprocessed image, obtaining feature maps of multiple scales. The Mamba-R network includes a VSSA block, namely the Visual Spatial Feature Enhancement Module. The VSSA module is used to enhance visual spatial features, further improving the spatial information expression ability and effectively enhancing the target detection model's recognition ability for targets in complex scenes. Step 3: Feature fusion. Use the Multiscale Attention Fusion Encoder, i.e., the MAFE encoder, to perform feature fusion on the shallow semantic feature information and deep semantic feature information of the feature maps. The MAFE encoder performs deep semantic feature interaction by combining the Mixed Topology Multi-Head Self-Attention Module, i.e., the MTMHSA module, effectively enhancing the feature fusion ability and being able to accurately capture spatial and context information at different scales. Step 4: Feature decoding. Use multiple stacked decoders to perform decoding operations on the features output by the MAFE encoder to obtain a feature sequence, and input the feature sequence into the prediction head for prediction, outputting the predicted box coordinates and predicted categories. Step 5: Loss function correction. Collect the outputs of the multiple stacked decoders, and then calculate the gradients of each decoder's output head through the loss function and adjust the parameters. Step 6: Object detection. Based on the corrected object detection model based on Mamba feature fusion obtained from Step 1 to Step 5, perform object detection on the image to be detected through the object detection model.

2. The object detection method based on Mamba feature fusion according to claim 1, characterized in that: The Step 1 at least includes the following steps: First, perform normalization processing to standardize the pixel values to a unified range, ensuring that the image data can be effectively transmitted to the neural network, mapping the pixel values to a unified range to avoid the impact of differences in the pixel value ranges of different images on the training process. Next, apply a series of random transformations to the image data for data augmentation. The random transformations at least include random color perturbation, size expansion, cropping, and horizontal flipping. Use the data augmentation method to increase sample diversity and improve the model's adaptability in different scenes and angles. Finally, perform unified size adjustment on the image, adjusting the image data to a fixed size to ensure that the image meets the input requirements of the network, so as to be input into the model for subsequent processing and help reduce unnecessary computational overhead. Through the above processing steps, ensure that the input image information is more representative and diverse, thereby improving the training efficiency and accuracy of the model.

3. A target detection method based on Mamba feature fusion according to claim 1, characterized in that: The Mamba-R network is stacked by multiple residual blocks, effectively extracting the basic feature information in the image and generating feature maps of multiple scales. These feature maps respectively represent the information of the image at different levels, having different spatial scales and depth features. The feature maps include three scale feature maps with channel numbers of 128, 256, and 512, namely S1, S2, and S3. The VSSA module is embedded in the backbone network to enhance the spatial feature expression ability of images; The VSSA module adopts a 2D selective scanning mechanism based on Mamba. Through a four-way processing strategy and state space modeling, it effectively captures the long-range spatial dependence relationships in images and improves the feature expression ability; Feature maps of three scales, namely S1, S2, and S3, are extracted through the combination of Mamba-R and the VSSA module. They represent the spatial information of the image at different levels. The feature maps of the three scales provide rich context information for subsequent feature fusion and object detection, enhancing the model's perception ability of multi-scale objects. The example steps are as follows: After the input feature map S3 is passed to the VSSA module, it enters the SS2D layer in the VSSA module. See the following formula: S3 = LayerNorm(S3) S4 = input + DropPath(SS2D(S3)) The SS2D layer generates an enhanced spatial feature S4 through dilated convolution operations, self-attention mechanism processing, and selective scanning. This feature map contains richer spatial information, which can help the model better capture the spatial dependence of the target. The finally generated enhanced feature S4 will be fused with the feature maps of other layers for subsequent feature fusion, thereby improving the model's recognition ability for small targets in complex scenes.

4. A target detection method based on Mamba feature fusion according to claim 1, characterized in that: The MTMHSA module combines multi-scale convolution and multi-head self-attention mechanism to enhance the model's processing ability for features of different scales and capture global context information; the MTMHSA module first extracts multi-scale information of the input feature map through a multi-scale convolution module, namely the MSC (Multiscale Spatial Convolution) module, and then calculates the weights of the features through the multi-head self-attention mechanism to capture global spatial dependence relationships. Finally, the output is the feature map after self-attention processing, representing an enhanced version of the input features.

5. The object detection method based on Mamba feature fusion according to claim 4, characterized in that: In the MAFE encoder, the feature map undergoes deep semantic feature interaction through the MTMHSA module to generate deep semantic information. Specifically: The extracted feature map is fed into the MAFE encoder and then enters the MTMHSA module; Input feature map processing: In the MTMHSA module, first, the input feature map x is processed by the MSC module using multiple convolutional branches with different dilation rates (dilation = [3, 5, 7]) to extract information of different scales, obtaining a feature map kv containing information of different scales; kv = MSC(x) Multi-scale convolution and feature rearrangement: Dilated convolution helps capture a wider range of spatial information, ensuring that the network can process a broader global context. The feature map kv obtained through multi-scale convolution and the original feature map x are rearranged into a format suitable for calculation. See the following formula: Q = rearrange(x) K, V = rearrange(kv) Then the adjusted feature map enters the multi-head self-attention mechanism of the MTMHSA module. Through the calculation between the query Q, key K, and value V, the spatial information and context information in the feature map are further enhanced; the dot product operation (Q·K T ) is calculated in the multi-head attention mechanism. Finally, through weighted averaging, that is, by using the softmax function, the fused feature information is obtained. This process is described as follows: Among them, d is the feature dimension of each attention head, which is used as a scaling factor to prevent the attention dot product from having an overly large value in high dimensions, resulting in vanishing gradients and ensuring training stability; c_attn = sigmoid(FC(AvgPool(x)))·x At the same time, calculate the channel attention c_attn, and weight the channel dimension through global average pooling and the FC layer: output = attn + c_attn Add the attention result attn and the channel attention c_attn, which helps to retain the input information and at the same time introduce enhanced features processed by the attention mechanism to obtain the fused feature map of the output: Through the MAFE encoder, shallow and deep features are effectively fused and enhanced, thereby improving the feature representation ability.

6. The object detection method based on Mamba feature fusion according to claim 1, characterized in that: The feature decoding combines the output of the encoder with the target query, and after being processed by multiple Transformer decoder layers, generates the final object detection result. The decoding process of the decoder includes a self-attention mechanism and a cross-attention mechanism; First, update the information of the target query through self-attention; Then perform cross-attention calculation with the reference point to further enhance the association between the target position and the features; Next, perform a non-linear mapping on the features through a feed-forward neural network; Finally, obtain the updated target query features; During the training process, the decoder calculates the predicted box coordinates and predicted classes to obtain the bounding box position of the target and the corresponding class labels; The output of the decoder includes the predicted box coordinates and predicted classes; In addition, if the auxiliary loss is enabled, multiple outputs in the decoder will be used to assist the optimization during the training process. The decoder also supports denoising training, which helps to improve the robustness of the model, especially when facing noisy data in the detection task; Finally, the output of the decoder will pass through a non-linear activation function to obtain the final bounding box and classification results.

7. A target detection method based on Mamba feature fusion according to claim 1, characterized in that: The loss function in S5 is as follows: The classification loss is the VFL loss: Among them, p i,c is the logits value of the predicted category, q i is the IoU value between the predicted bounding box and the ground truth bounding box, t i,c is the one-hot encoding of the category, w i,c = α·σ(p i,c ) γ ·(1 - t i,c ) + q i ·t i,c is the dynamic weight, α and γ are the adjustment coefficients of the focal loss, mean c represents taking the average over all categories, N is the total number of samples, and BCE represents the binary cross-entropy loss; The bounding box regression loss is L1Loss + GIoU Loss: Among them, B i and are the coordinates of the ground truth box and the predicted box respectively, and ‖.‖1 represents the L1 norm; the absolute error; is the GIoU value of the ground truth box and the predicted box; The overall expression of the loss function: L total = λ VFL L VFL + λ L1 L L1 + λ GIoU L GIoU Among them, λ VFL , λ L1 and λ GIoU are adjustment coefficients.

Citation Information

Cited By

  • PRNU anonymity method based on multi-scale and hierarchical feature fusion

    CN120599058A

  • Single-flow RGB-D target tracking method based on state space model

    CN120599236A

  • Single-stream rgb-d object tracking method based on state space model

    CN120599236B

  • Video target detection method based on motion modulation attention spatial-temporal feature pyramid

    CN120635791A

  • Dynamic adaptive image semantic transmission method based on state space model

    CN120769302A