Multi-modal target detection method based on dual-backbone YOLO architecture

Through the multimodal object detection method of the dual-backbone YOLO architecture, the redundancy and mutual interference problems in cross-modal information fusion are solved, and efficient and accurate object detection effects are achieved.

CN120495638APending Publication Date: 2025-08-15ZHEJIANG NORMAL UNIV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510925248.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-05
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing multimodal object detection methods have problems such as redundant model structure increasing training costs, not fully considering the object detection requirements, mutual interference between modals and insufficient feature utilization, resulting in insufficient detection performance.

Method used

The multimodal object detection method based on the dual-backbone YOLO architecture is adopted, and RGB and IR images are processed separately through the dual-stream feature extraction network, and feature fusion is performed by combining the Z-scan state space channel fusion module and the dual-Z-scan state space fusion module. A windmill convolution is introduced to replace traditional convolution, enhancing feature interaction and context information capture.

Benefits of technology

It improves the efficiency and accuracy of cross-modal object detection, significantly improves detection performance, especially in complex scenarios and small-objective dense scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495638A_ABST
    Figure CN120495638A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal target detection method based on a dual-backbone YOLO architecture, which is applied to the technical field of multi-modal target detection, and comprises the following steps: constructing a multi-modal target detection model based on the dual-backbone YOLO architecture; wherein the detection trunk comprises a double-flow feature extraction network and three fusion Mama blocks, the detection network is composed of a neck module and a head module which are used for multi-modal target detection, and the input of the detection network is the output of the three fusion Mama blocks; the double-flow feature extraction network is used for extracting local features from the RGB image and the IR image respectively; each fusion Mama block comprises a Z scanning state space channel fusion module for performing shallow feature fusion on local features and a double Z scanning state space fusion module for performing deep feature fusion on a shallow feature fusion result; and inputting a to-be-detected image to the multi-modal target detection model to obtain a target detection result. According to the invention, the target detection precision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal target detection, and in particular to a multimodal target detection method based on a dual-backbone YOLO architecture. Background Art

[0002] Object detection in complex scenes faces numerous challenges. Due to the limited wavelength range of visible light, it is difficult to effectively acquire object information in poor lighting conditions (such as those in dense smoke). While the inclusion of infrared information can improve this problem, infrared image quality is low, making it difficult to meet detection requirements when used alone. The cross-modal complementary information between visible and infrared images can significantly improve detection performance.

[0003] However, commonly used fusion detection strategies have some flaws: bimodal image fusion does not focus on the target detection task, and existing fusion methods often fail to fully consider the specific needs of target detection, resulting in redundant model structures that increase training costs; they also fail to fully consider the mutual interference between the two modal images. For example, infrared images may weaken the quality of visible light imaging, and directly fusing image pairs without cross-modal enhancement makes it difficult to improve performance. Existing RGB-IR detection models lack multimodal information fusion strategies. Although these models have improved in multimodal information fusion, they are still insufficient in terms of interaction. The clear boundary between single-modal image processing and feature fusion leads to insufficient utilization of cross-modal information, and a lack of complex interaction in the channel and spatial dimensions, ignoring the potential relationship between semantic and structural information.

[0004] With the rapid development of single-modality detectors (such as the YOLO series), multimodal object detectors have emerged, aiming to fully leverage image information from different modalities. To date, research in multimodal object detection has primarily focused on two directions: pixel-level fusion and feature-level fusion. However, both approaches still have limitations in modeling modal differences and fusion complexity. Pixel-level fusion generates a new input image by simply superimposing or weighted averaging visible and infrared images at the pixel level. While this approach offers the advantage of simplicity, it also fails to fully exploit the deep connections between the two modalities, is susceptible to noise and interference, and struggles to effectively improve detection performance. Feature-level fusion fuses the features of visible and infrared images during the feature extraction phase. While this approach can better capture complementary information between modalities, it also faces challenges, such as effectively modeling modal differences and handling the spatial and semantic consistency of features across modalities. Furthermore, feature-level fusion typically requires a more complex network structure, increasing computational cost and training difficulty.

[0005] Therefore, how to provide a multimodal target detection method based on a dual-backbone YOLO architecture that can effectively solve the above problems and achieve efficient cross-modal target detection is an issue that technicians in this field urgently need to solve. Summary of the Invention

[0006] In view of this, the present invention proposes a multimodal target detection method based on a dual-backbone YOLO architecture.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A multimodal target detection method based on a dual-backbone YOLO architecture, comprising:

[0009] Step 1: Build a multimodal object detection model based on the dual-backbone YOLO architecture. The detection backbone of the dual-backbone YOLO architecture includes a two-stream feature extraction network and three fusion Mamba blocks. The detection network consists of a neck module and a head module for multimodal object detection. The inputs of the neck module and the head module are the outputs of the three fusion Mamba blocks. The two-stream feature extraction network is used to extract local features from the RGB image and the IR image respectively. Each fusion Mamba block includes a Z-scan state space channel fusion module for shallow feature fusion of local features, and a dual Z-scan state space fusion module for deep feature fusion of the shallow feature fusion results.

[0010] Step 2: Input the image to be tested into the multimodal object detection model to obtain the object detection result.

[0011] Optionally, a Z-scan state space channel fusion module for shallow feature fusion of local features is provided, specifically:

[0012] First, the local features F of RGB image and IR image are Ri 、F IRi The channel fusion operation is used to generate new local features, and the Conv 1×1 Restore to the original number of channels to get F Mi , and then respectively with F Ri 、F IRi Perform element-by-element multiplication to obtain M Ri 、M IRi ; Apply two Z-scan state space modules to M Ri 、M IRi , and enhance the feature map again so that F Ri 、F IRi Two Z-scan state space modules are applied to M Ri 、M IRi The corresponding feature weights obtained are multiplied to obtain the output after shallow fusion features The specific process is as follows:

[0013]

[0014] Among them, CZSSBlock is the Z-scan state space module.

[0015] Optionally, a dual Z-scan state space fusion module is used to perform deep feature fusion on the shallow feature fusion results, specifically:

[0016] First, the shallow fusion features Projected into the hidden state space through an ungated Z-SSM block, we get X Ri 、X IRi , and Projection to obtain the gate parameter Y Ri 、Y IRi , use Y Ri 、Y IRi The gated output modulates X Ri 、X IRi , so that the hidden state features are fused into Finally, the projection is back to the original space and the complementary features are obtained through residual connection. The specific process is as follows:

[0017]

[0018] Among them, Project is the operation of projecting features into the hidden state space; Project_Linear is the projection operation with linear transformation; and X IRi is the hidden state feature; and They are respectively the two streams with parameters θ i and ω i Gating operation; and are the hidden states of RGB and IR after feature interaction; This is element-wise multiplication.

[0019] The optional dual-backbone YOLO architecture detection backbone also includes: introducing the C2f_PC module, which is used to replace the ordinary convolution in C2f with windmill convolution, and dynamically adjust the shape, size, and parameters of the convolution kernel according to the characteristics of the input image and task requirements.

[0020] Optional, windmill convolution, specifically:

[0021] The pinwheel convolution creates horizontal and vertical convolution kernels for different areas of the image through asymmetric padding. The convolution kernels diffuse outward, and batch normalization and Sigmoid linear units are applied after each convolution. The first layer of the pinwheel convolution performs parallel convolution as follows:

[0022]

[0023] in, is the convolution operator; is a 1×3 convolution kernel with an output channel of c′; BN is batch normalization; SiLU is a Sigmoid linear unit; P(1,0,0,3) is a padding parameter, indicating the number of padded pixels on the left, right, top, and bottom sides, respectively; after the first layer of interleaved convolution, the relationship between the height h′, width w′, and number of channels c′ of the output feature map and the input feature map is as follows:

[0024]

[0025]

[0026] Where h1 and w1 are the height and width of the input tensor X respectively; c2 is the number of channels of the final output feature map of the windmill convolution; s is the stride; the results of the first layer of interleaved convolution are connected and output as follows:

[0027]

[0028] Finally, the concatenated tensor is passed through a convolution kernel with no padding Normalization is performed; the height and width of the output feature map are adjusted to the preset values h2 and w2, so that the windmill convolution can be used interchangeably with the Conv layer as a channel attention mechanism to analyze the contribution of different convolution directions; the final output is as follows:

[0029]

[0030] The parameters of the windmill convolution are calculated as follows:

[0031]

[0032] The present invention also provides a multimodal target detection system based on a dual-backbone YOLO architecture using a multimodal target detection method based on a dual-backbone YOLO architecture, comprising:

[0033] Multimodal object detection model construction module: This module is used to build a multimodal object detection model based on the dual-backbone YOLO architecture. The detection backbone of the dual-backbone YOLO architecture includes a two-stream feature extraction network and three fusion Mamba blocks. The detection network consists of a neck module and a head module for multimodal object detection. The inputs of the neck module and the head module are the outputs of the three fusion Mamba blocks. The two-stream feature extraction network is used to extract local features from RGB images and IR images respectively. Each fusion Mamba block includes a Z-scan state space channel fusion module for shallow feature fusion of local features, and a dual Z-scan state space fusion module for deep feature fusion of the shallow feature fusion results.

[0034] Object detection module: used to input the image to be tested into the multimodal object detection model to obtain the object detection result.

[0035] The present invention further provides an electronic device, comprising:

[0036] memory for storing computer programs;

[0037] A processor is configured to implement steps of a multimodal target detection method based on a dual-backbone YOLO architecture when executing a computer program.

[0038] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the steps of a multimodal target detection method based on a dual-backbone YOLO architecture.

[0039] As can be seen from the above technical solution, compared with the existing technology, the present invention proposes a multimodal object detection method based on the dual-backbone YOLO architecture. This invention achieves efficient cross-modal object detection through multimodal data fusion and innovative convolution operations. Specifically, a two-stream feature extraction network is adopted to process RGB and IR images separately. This design enables the extraction of local features from different input types, thereby better capturing scene information. After preliminary feature extraction, shallow feature fusion is performed using ZSSCFB to generate interactive features. This process not only reduces the differences between the two modalities but also enhances feature complementarity. The generated interactive features are then input into DZSSFB to achieve deep feature fusion. The DZSSFB module further explores the deep relationships between the features of the two modalities, generating richer complementary features. This stage of fusion not only enhances the expressive power of features but also provides a more solid foundation for subsequent object detection tasks. In addition, the Cf_PC module is introduced in the backbone network, replacing traditional convolution operations with pinwheel convolutions to better capture contextual information and increase the receptive field. This not only improves the efficiency of feature extraction but also ensures that the original information is not lost. Experimental results show that the performance of this model surpasses several current object detection methods based on image fusion, demonstrating its superior performance in cross-modal object detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0041] Figure 1 Schematic diagram of the method of the present invention.

[0042] Figure 2 Schematic diagram of the fused Mamba block structure of the present invention.

[0043] Figure 3 Schematic diagram of the Z-scan state space channel fusion module structure of the present invention.

[0044] Figure 4 Schematic diagram of the dual Z scanning state space fusion module structure of the present invention.

[0045] Figure 5 It is a schematic diagram of the windmill convolution structure of the present invention.

[0046] Figure 6 Schematic diagram for comparing the detection effects of the present invention on M3FD. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0048] Example 1:

[0049] Embodiment 1 of the present invention discloses a multimodal target detection method based on a dual-backbone YOLO architecture, such as Figure 1 As shown, including:

[0050] Step 1: Construct a multimodal target detection model (CMOD-YOLO) based on the dual-backbone YOLO architecture; the detection backbone of the dual-backbone YOLO architecture includes a two-stream feature extraction network and three fusion Mamba blocks (FusionMambaBlock, FMB). The detection network consists of a neck module and a head module for multimodal target detection. The input of the YOLOv8 neck module and head module are the outputs P3, P4, and P5 of the three fusion Mamba blocks (FMB is only added to the last three stages to generate fusion features P3, P4, and P5). The two-stream feature extraction network is used to extract local features from RGB images and IR images respectively, denoted as F Ri and F IRi .

[0051] Existing methods mainly focus on the integration of spatial features, but fail to fully consider the feature differences between different modalities. Therefore, these fusion models are insufficient in modeling the correlation between different modal targets, thus limiting the representation ability of the model. Figure 2 As shown in FIG, in order to reduce the differences between multimodal features and enhance the expression consistency of fused features and realize the effective fusion of multimodal features, the present invention designs a fusion Mamba block (FMB). This module promotes the interaction and association of multimodal features by constructing a hidden state space, thereby improving the expression ability and detection performance of the model. Each fusion Mamba block includes a Z-scan State Space Channel Fusion Block (ZSSCFB) for shallow feature fusion of local features to generate interactive features. and And the Dual Z-scan State Space Fusion Block (DZSSFB) is used to perform deep feature fusion on the shallow feature fusion results to generate complementary features These two modules enhance the representation consistency of fused features by reducing the differences between multimodal features, ensuring efficient interaction and consistent expression of multimodal features.

[0052] Images of different modalities typically have different feature distributions and intensities across channels. Channel fusion allows for the direct and organic combination of features from different modalities, allowing the fused feature map to simultaneously incorporate information from multiple modalities, resulting in a richer and more expressive feature representation. For example, in the fusion of infrared and visible light images, channel fusion can combine the thermal radiation features of the infrared image with the texture details of the visible light image, allowing the target object's temperature characteristics to be highlighted in the fused image while retaining clear outlines and texture information.

[0053] Traditional SS2D methods mainly focus on processing within the two-dimensional plane, and pay less attention to the spatial continuity between different areas in the image, which may lead to the loss or blurring of image details and affect the utilization of detail information in cross-modal target detection. In order to solve this problem, the present invention adopts a Z-shaped scanning method to enhance the spatial consistency of adjacent pixels in the image by simulating the natural scanning path of the human visual system for the image. Ensure the coherence of each adjacent pixel in the image during the scanning process, so as to better preserve the spatial details and structural information of the image. In cross-modal target detection, this method can fully consider the spatial relationship between pixels, more effectively capture local features in the image, and improve the accuracy and completeness of target detection.

[0054] ZSSCFB aims to enhance the cross-modal feature interaction of shallow feature fusion through channel fusion operations and Z-SSM blocks. It integrates information from different channels through a dual enhancement mechanism to construct cross-modal feature correlation, which enriches the diversity of channel features to improve fusion performance.

[0055] The Z-scan state space channel fusion module performs shallow feature fusion on local features, such as Figure 3 As shown, specifically:

[0056] First, the local features F of RGB image and IR image are Ri 、F IRi The channel fusion operation is used to generate new local features, and the Conv 1×1 Restore to the original number of channels to get F Mi , and then respectively with F Ri 、F IRi Perform element-by-element multiplication to obtain M Ri 、M IRi ; Apply two Z-scan state space modules to M Ri 、M IRi, enhances the cross-modal interaction from shallow features and enhances the feature map again, this time in order to make each feature map of RGB and IR take full advantage of the other modality, so that F Ri 、F IRi Two Z-scan state space modules are applied to M Ri 、M IRi The corresponding feature weights obtained are multiplied to obtain semantic and texture information from another modality, and finally the output after shallow fusion features is obtained The specific process is as follows:

[0057]

[0058] Among them, CZSSBlock is the Z-scan state space module.

[0059] In order to further reduce the modal differences, the present invention constructs a hidden state space for the association and complementation of cross-modal features. A dual Z-scan state space fusion module (DZSSFB) is designed to model cross-modal object correlation to promote deep feature fusion. Specifically, by projecting the features of the two modalities into a hidden state space, and using the Z-SSM block combined with the gating mechanism for dual hidden state conversion. The gating mechanism can selectively allow key features to pass through and suppress less important features based on the importance and relevance of different modal features, so that the model can focus more on key information and improve the fusion effect.

[0060] The double Z-scan state space fusion module performs deep feature fusion on the shallow feature fusion results, such as Figure 4 As shown, specifically:

[0061] First, the shallow fusion features Projected into the hidden state space through an ungated Z-SSM block, we get X Ri 、X IRi , and Projection to obtain the gate parameter Y Ri 、Y IRi , use Y Ri 、Y IRi The gated output modulates X Ri 、X IRi , so that the hidden state features are fused into Finally, the projection is back to the original space and the complementary features are obtained through residual connection. The specific process is as follows:

[0062]

[0063] Among them, Project is the operation of projecting features into the hidden state space; Project_Linear is the projection operation with linear transformation; and X IRi is the hidden state feature; and They are respectively the two streams with parameters θ i and ω i Gating operation; and are the hidden states of RGB and IR after feature interaction; This is element-wise multiplication.

[0064] Furthermore, to better extract and preserve useful features from both modalities, the dual-backbone YOLO architecture's detection backbone also includes the introduction of the C2f_PC module, which replaces the standard convolution in C2f with a pinwheel convolution (PConv). This module dynamically adjusts the shape, size, and parameters of the convolution kernel based on the characteristics of the input image and task requirements, thereby enhancing the flexibility and accuracy of feature extraction. Pinwheel convolution increases the backbone network's receptive field in different directions while ensuring that original information is not lost.

[0065] Pinwheel convolution is an innovative convolution operation method. Unlike traditional convolution, which performs the same convolution operation at every location in an image, PConv creates horizontal and vertical convolution kernels in different regions through asymmetric padding. This method can adjust the weights and shapes of the convolution kernels in different directions based on the distribution characteristics of the target, thereby more effectively capturing the features of small and faint targets. In addition, PConv's convolution kernels are not fixed but can change dynamically. It can adaptively adjust the shape and parameters of the convolution kernel based on the characteristics of the target in the input image, thereby extracting features more flexibly. This dynamism makes PConv more adaptable and robust when processing complex and changing images, significantly improving feature extraction.

[0066] Windmill convolution, such as Figure 5 As shown, specifically:

[0067] The pinwheel convolution creates horizontal and vertical convolution kernels for different areas of the image through asymmetric padding. The convolution kernels diffuse outward, and batch normalization and Sigmoid linear units are applied after each convolution. The first layer of the pinwheel convolution performs parallel convolution as follows:

[0068]

[0069] in, is the convolution operator; is a 1×3 convolution kernel with an output channel of c′; BN is batch normalization; SiLU is a Sigmoid linear unit; P(1,0,0,3) is a padding parameter, indicating the number of padded pixels on the left, right, top, and bottom sides, respectively; after the first layer of interleaved convolution, the relationship between the height h′, width w′, and number of channels c′ of the output feature map and the input feature map is as follows:

[0070]

[0071]

[0072] Where h1 and w1 are the height and width of the input tensor X respectively; c2 is the number of channels of the final output feature map of the windmill convolution; s is the stride; the results of the first layer of interleaved convolution are connected and output as follows:

[0073]

[0074] Finally, the concatenated tensor is passed through a convolution kernel with no padding Normalization is performed; the height and width of the output feature map are adjusted to the preset values h2 and w2, so that the windmill convolution can be used interchangeably with the Conv layer as a channel attention mechanism to analyze the contribution of different convolution directions; the final output is as follows:

[0075]

[0076] like Figure 5 The upper right corner shows that the receptive field of PConv (k=3) is 25. As the number of convolutions decreases from the center outward, the effectiveness of the receptive field also decreases, showing characteristics similar to Gaussian distribution. PConv utilizes grouped convolution to significantly increase the receptive field while minimizing the number of parameters. The parameters of the pinwheel convolution are calculated as follows:

[0077]

[0078] The C2f module typically focuses on feature extraction and fusion in local regions, while neglecting the importance of global contextual information. In cross-modal object detection, global information is crucial for understanding the relationship between the target and the background, as well as the overall distribution of the target. PConv, through techniques such as asymmetric padding and dynamic convolution kernels, can better adapt to the characteristics of data from different modalities and significantly expand the receptive field without increasing the computational load. This improvement enables the C2f module, after replacing the pinwheel convolution, to capture a wider range of contextual information, helping to reduce information loss.

[0079] Step 2: Input the image to be tested into the multimodal object detection model to obtain the object detection result.

[0080] To verify the effectiveness of the CMOD-YOLO method proposed in this paper, two backbone networks based on YOLOv8s and YOLOv8l were used as baseline models for fair comparison with the latest methods. Comparative experiments were conducted on the LLVIP dataset and the M3FD dataset. The M3FD dataset is a widely used infrared and visible light image fusion and target detection dataset, containing 4,200 pairs of images, totaling 8,400 images. This dataset covers a variety of scenes, such as urban streets and natural landscapes, and can effectively reflect the characteristics of infrared and visible light images in different environments. Secondly, the LLVIP dataset is larger, including 16,836 pairs of images, totaling 33,672 images. The characteristic of this dataset is its diverse scene settings, from indoors to outdoors, from daytime to nighttime, covering almost all possible application scenarios, providing rich samples for model training and testing. The results of different methods are shown in Tables 1 and 2, respectively.

[0081] Table 1 Comparative experiments on LLVIP

[0082]

[0083] Table 2 Comparative experiments on M3FD

[0084]

[0085] This paper uses two different backbone network approaches and compares them with four state-of-the-art multispectral object detection methods and five single-modal detection methods. For single-modal detection, the experimental results in Table 1 show that detection performance using only infrared images outperforms that using only RGB images. This is primarily due to the ability of infrared images to better capture thermal radiation information in low-light conditions. After feature fusion of RGB and infrared images, detection performance is significantly improved, generally surpassing single-modal detection methods. For example, the IRFS method outperforms the Cascade R-CNN method using only RGB modality in mAP50 by 8.9% and outperforms the infrared method in mAP50 by 2.6%. The YOLOv8s detection framework using infrared modality input achieves 93.7% mAP50 and 60.8% mAP50:95, significantly outperforming the fusion method DIVFusion (89.4% and 52.7%). The CMOD-YOLO-s model outperforms the previous best fusion method IRFS in mAP50 and mAP50:95 by 2.8% and 4.4%, respectively. In addition, the proposed method also performs well on the YOLOv8-l backbone network, achieving state-of-the-art performance with mAP50 reaching 97.8% and mAP50:95 reaching 65.3%.

[0086] On the M3FD dataset, the proposed model was compared in detail with five cross-modal object detectors. As shown in Table 2, the proposed model based on YOLOv8-s and YOLOv8-1 outperformed the other five cross-modal object detectors across all evaluation metrics. These metrics include mAP50 and mAP50.95, as well as detection performance for specific categories such as people, buses, cars, motorcycles, lamps, and trucks. In particular, the method based on the YOLOv8-1 backbone achieved the best results across all metrics. Specifically, the mAP50 reached 86.9%, demonstrating the model's excellent detection accuracy at high IoU thresholds. The mAP50.95 reached 58.0%, demonstrating the model's robustness across different IoU thresholds. For the people category, the detection accuracy reached 86.8%, which is particularly important for pedestrian detection, as pedestrians' morphology and pose vary greatly, making detection more challenging. The detection accuracy for the bus category reached 89.2%, demonstrating the model's superiority in detecting large vehicles. The detection accuracy for the Car category reached 95.7%, demonstrating the model's excellent performance in detecting common vehicles. The detection accuracy for the Motorcycle category reached 76.5%, demonstrating the model's ability to effectively identify motorcycles despite their small size and ease of occlusion. The detection accuracy for the Lamp category reached 91.0%, demonstrating the model's robustness in detecting fixed objects. The detection accuracy for the Truck category reached 82.1%, further validating the model's reliability in detecting diverse vehicle types.

[0087] like Figure 6Figure 2 shows the detection performance of different models on the M3FD dataset. The results demonstrate that image fusion-based methods offer limited improvements in object detection performance. These methods primarily focus on improving the quality of image fusion, for example by optimizing brightness, detail, and texture to produce clearer composite images. Despite significant improvements in image quality, their detection performance remains unsatisfactory when dealing with dense scenes. Specifically, when faced with a large number of small targets or densely populated objects, image fusion-based methods often exhibit poor detection accuracy and suffer from serious false and missed detection issues. These issues not only affect the accuracy of detection results but can also lead to misjudgments in subsequent analysis. In contrast, the model proposed in this paper does not rely on traditional image fusion techniques, but instead employs a more advanced cross-modal object detection approach. This approach constructs a hidden state space using Multi-Factor fusion (MFB), enhancing the interaction and correlation between deep and shallow features. Specifically, MFB effectively captures and integrates data features from different modalities, fully accounting for differences between the modalities. This results in a model with greater robustness and accuracy when dealing with complex scenes. This significantly reduces false and missed detections, particularly in scenes with densely populated small targets. In addition, PConv effectively solves the problem of information loss in the edge area of the network caused by traditional convolution, allowing the model to maintain high detection accuracy when dealing with incomplete or occluded targets.

[0088] This paper uses the M3FD dataset to conduct ablation studies to validate the effectiveness of ZSSCFB, DZSSFB, and PConv. All ablation experiments were conducted on YOLOv8-1. First, the impact of the ZSSCFB and DZSSFB modules on model performance was verified. The experimental results, shown in Table 3, detail the changes in detection accuracy under different configurations. The data shows that both modules significantly improve the model's detection accuracy. Compared to the baseline model, using the ZSSCFB module alone improves mAP50 by 4.2% and mAP50:95 by 3.4%. Using the DZSSFB module alone improves mAP50 by 6.1% and mAP50:95 by 7.0%. These results demonstrate that both the ZSSCFB and DZSSFB modules effectively enhance the model's detection capabilities, particularly when handling object detection tasks in complex scenarios. When both modules are used simultaneously, the initial exchange and shallow mapping fuse the features of the two modalities, significantly reducing feature differences and fully leveraging the advantages of each modality and their complementary information during deep fusion. This combination not only enhances the model's understanding of different types of features, but also enables the model to maintain high detection accuracy in more complex environments. Because the deep feature map contains stronger semantic information, the target detection performance is greatly improved, with mAP50 and mAP50:95 increasing by 7.8% and 11.6% respectively. In addition, after adding the PConv module alone, the detection performance is also improved due to its ability to capture a wider range of contextual information and reduce information loss. In addition, the effect of using all three modules (ZSSCFB, DZSSFB and PConv) at the same time was tested. The results show that this combination further improves the accuracy of target detection, with mAP50 and mAP50:95 increasing by 9.7% and 15.8% respectively. This shows that the synergy between multiple modules can bring more significant performance improvements, especially when dealing with difficult target detection tasks. This combination can give full play to the advantages of each module to achieve higher detection accuracy.

[0089] Table 3 Ablation experiments on M3FD

[0090]

[0091] Example 2:

[0092] Embodiment 2 of the present invention discloses a multimodal target detection system based on a dual-backbone YOLO architecture using a multimodal target detection method based on a dual-backbone YOLO architecture, including:

[0093] Multimodal object detection model construction module: This module is used to build a multimodal object detection model based on the dual-backbone YOLO architecture. The detection backbone of the dual-backbone YOLO architecture includes a two-stream feature extraction network and three fusion Mamba blocks. The detection network consists of a neck module and a head module for multimodal object detection. The inputs of the neck module and the head module are the outputs of the three fusion Mamba blocks. The two-stream feature extraction network is used to extract local features from RGB images and IR images respectively. Each fusion Mamba block includes a Z-scan state space channel fusion module for shallow feature fusion of local features, and a dual Z-scan state space fusion module for deep feature fusion of the shallow feature fusion results.

[0094] Object detection module: used to input the image to be tested into the multimodal object detection model to obtain the object detection result.

[0095] Example 3:

[0096] Embodiment 3 of the present invention discloses an electronic device, including:

[0097] memory for storing computer programs;

[0098] A processor is configured to implement steps of a multimodal target detection method based on a dual-backbone YOLO architecture when executing a computer program.

[0099] Example 4:

[0100] Embodiment 4 of the present invention discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a multimodal target detection method based on a dual-backbone YOLO architecture are implemented.

[0101] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0102] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal target detection method based on a dual-backbone YOLO architecture, characterized in that: include: Step 1: Construct a multimodal target detection model based on a dual-backbone YOLO architecture; wherein the detection backbone of the dual-backbone YOLO architecture includes a dual-stream feature extraction network and three fusion Mamba blocks. The detection network consists of a neck module and a head module for multimodal target detection. The inputs of the neck module and the head module are the outputs of the three fusion Mamba blocks. The dual-stream feature extraction network is used to extract local features from RGB images and IR images respectively. Each fusion Mamba block includes a Z-scan state space channel fusion module for performing shallow feature fusion on the local features, and a dual Z-scan state space fusion module for performing deep feature fusion on the shallow feature fusion results. Step 2: Input the image to be tested into the multimodal target detection model to obtain the target detection result.

2. A multimodal target detection method based on a dual-backbone YOLO architecture according to claim 1, characterized in that The Z-scan state space channel fusion module for performing shallow feature fusion on the local features is specifically: First, the local features F of RGB image and IR image are Ri 、F IRi The channel fusion operation is used to generate new local features, and the Conv 1×1 Restore to the original number of channels to get F Mi , and then respectively with F Ri 、F IRi Perform element-by-element multiplication to obtain M Ri 、M IRi ; Apply two Z-scan state space modules to M Ri 、M IRi , and enhance the feature map again so that F Ri 、F IRi Separately with two Z-scan state space modules applied to M Ri 、M IRi The corresponding feature weights obtained are multiplied to obtain the output after shallow fusion features The specific process is as follows: Among them, CZSSBlock is the Z-scan state space module.

3. A multimodal target detection method based on a dual-backbone YOLO architecture according to claim 1, characterized in that The dual Z-scan state space fusion module performs deep feature fusion on the shallow feature fusion results, specifically: First, the shallow fusion features Projected into the hidden state space through an ungated Z-SSM block, we get X Ri 、X IRi , and Projection to obtain the gate parameter Y Ri 、Y IRi , use Y Ri 、Y IRi The gated output modulates X Ri 、X IRi , so that the hidden state features are fused into Finally, the projection is back to the original space and the complementary features are obtained through residual connection. The specific process is as follows: Among them, Project is the operation of projecting features into the hidden state space; Project_Linear is the projection operation with linear transformation; and X IRi is the hidden state feature; and They are respectively the two streams with parameters θ i and ω i Gating operation; and are the hidden states of RGB and IR after feature interaction; This is element-wise multiplication.

4. The multimodal target detection method based on the dual-backbone YOLO architecture according to claim 1, characterized in that: The detection backbone of the dual-backbone YOLO architecture also includes: introducing a C2f_PC module for replacing ordinary convolution in C2f with windmill convolution, and dynamically adjusting the shape, size, and parameters of the convolution kernel according to the characteristics of the input image and task requirements.

5. A multimodal target detection method based on a dual-backbone YOLO architecture according to claim 4, characterized in that: The windmill convolution is specifically: The windmill convolution creates horizontal and vertical convolution kernels for different regions of the image through asymmetric padding. The convolution kernels spread outward, and batch normalization and Sigmoid linear units are applied after each convolution. The first layer of the windmill convolution performs parallel convolution as follows: in, is the convolution operator; is a 1×3 convolution kernel with an output channel of c′; BN is batch normalization; SiLU is a Sigmoid linear unit; P(1,0,0,3) is a padding parameter, indicating the number of padded pixels on the left, right, top, and bottom sides, respectively; after the first layer of interleaved convolution, the relationship between the height h′, width w′, and number of channels c′ of the output feature map and the input feature map is as follows: Where h1 and w1 are the height and width of the input tensor X respectively; c2 is the number of channels of the final output feature map of the windmill convolution; s is the stride; the results of the first layer of interleaved convolution are concatenated and output as follows: Finally, the concatenated tensor is passed through a convolution kernel with no padding Normalization is performed; the height and width of the output feature map are adjusted to the preset values h2 and w2, so that the windmill convolution can be used interchangeably with the Conv layer as a channel attention mechanism for analyzing the contribution of different convolution directions; the final output is as follows: The parameters of the windmill convolution are calculated as follows:

6. A multimodal target detection system based on a dual-backbone YOLO architecture using a multimodal target detection method based on a dual-backbone YOLO architecture according to any one of claims 1 to 5, characterized in that: include: Multimodal target detection model construction module: used to build a multimodal target detection model based on the dual-backbone YOLO architecture; wherein the detection backbone of the dual-backbone YOLO architecture includes a dual-stream feature extraction network and three fusion Mamba blocks. The detection network consists of a neck module and a head module for multimodal target detection. The inputs of the neck module and the head module are the outputs of the three fusion Mamba blocks. The dual-stream feature extraction network is used to extract local features from RGB images and IR images respectively. Each fusion Mamba block includes a Z-scan state space channel fusion module for performing shallow feature fusion on the local features, and a dual Z-scan state space fusion module for performing deep feature fusion on the shallow feature fusion results. Target detection module: used to input the image to be tested into the multimodal target detection model to obtain the target detection result.

7. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of a multimodal target detection method based on a dual-backbone YOLO architecture as described in any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the multimodal target detection method based on the dual-backbone YOLO architecture according to any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • Cross-modal smoke shielding human body identification method and device based on double-engine cooperation

    CN121170846A

  • Vehicle target detection system and method based on Mama and double-domain interaction

    CN122135319A

  • Vehicle target detection system and method based on mamba and dual domain interaction

    CN122135319B