Multi-modal image target detection method based on heterogeneity perception attention fusion network

Through the dual-stream network backbone and heterogeneous perception enhancement module, refined channel attention fusion module and elliptical dynamic CIoU loss function, the modal heterogeneity and bounding box regression problems in the multi-spectral fusion method are solved, and high-precision object detection in complex environments is achieved.

CN120431316APending Publication Date: 2025-08-05NORTHEASTERN UNIV CHINA +1

Patent Information

Application Number
CN202510518677.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing multispectral fusion method ignores the heterogeneity between modes and the feature differences of abnormal targets under different modes, resulting in unsatisfactory model feature extraction and excessive computational overhead. The existing bounding box regression method converges slowly and positioning in multispectral target detection. The performance of traditional detectors significantly declined under low light or extreme weather conditions.

Method used

The dual-stream network backbone is used for multi-scale feature extraction, a heterogeneous perception enhancement module and a refined channel attention fusion module are introduced, a channel-space bidirectional attention mechanism is built, and the elliptical dynamic CIoU loss function is combined to optimize the geometric alignment of the anchor box and the target, and the cross-modal feature fusion and bounding box regression are enhanced.

Benefits of technology

The accuracy and robustness of multimodal image object detection are improved, especially in low light or extreme weather conditions, the target detection capability in complex scenes is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431316A_ABST
    Figure CN120431316A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous perceptual attention fusion network-based multi-modal image target detection method, which aims at the target detection problem in a complex scene, utilizes a double-backbone network to respectively perform multi-scale feature extraction on visible light and infrared images, generates a feature pyramid containing different levels of details and semantic information, and performs feature extraction on the visible light and infrared images. Through a heterogeneous perception enhancement module, a channel-space bidirectional attention mechanism is constructed, cross-modal heterogeneous enhancement is carried out on features of different levels, key region information is highlighted, by means of a refined channel attention fusion module, discriminative features are effectively amplified, modal specific noise is suppressed at the same time, and in a target detection link, an elliptical dynamic CIoU loss function is adopted, so that the target detection accuracy is improved. According to the method, the external connection ellipse or the inconnection ellipse is dynamically selected according to the proportion of the target to calculate the intersection-to-union ratio, geometric alignment of the anchor frame and the target is optimized, the detection precision and robustness of the multi-modal image target are effectively improved, and the method has wide application prospects in the fields of automatic driving, anomaly detection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a multimodal image fusion method and target detection method based on deep learning, which can be applied to scenarios such as autonomous driving, intelligent monitoring, and driving anomaly detection that require reliable environmental perception under various lighting and weather conditions. Background Art

[0002] With the rapid development of autonomous driving and intelligent transportation systems, detecting sudden anomalies in road environments is becoming increasingly important. For example, sudden road collapses, obstacles, and scattered objects at night or in inclement weather can pose serious safety hazards. With the rapid evolution of Level 4 autonomous driving technology and the development of V2X vehicle-to-everything (V2X) systems, detecting road anomalies has become a key research area for ensuring intelligent transportation safety. According to statistics, a large proportion of traffic accidents caused annually by sudden road anomalies (including but not limited to collapsed potholes, scattered obstacles, and illegally intruded objects) occur in low-light or extreme weather conditions. While current mainstream detection algorithms based on visible light cameras, such as YOLO and Faster R-CNN, can achieve high mean average average precision (mAP) accuracy under normal lighting conditions, their detection performance plummets in nighttime environments with illumination below 10 lux or in rainy and foggy conditions with visibility less than 50 meters. This significant degradation in perception is primarily due to the strong dependence of the visible light spectrum on lighting conditions and the limitations of single-modality data in terms of feature representation. Object detection is a core function in detecting road anomalies, requiring robustness across a wide range of lighting and weather conditions. Efficient and reliable object detection algorithms can assist autonomous vehicles in better detecting road anomalies and planning paths. However, the performance of traditional detectors based on visible light (RGB) cameras often degrades significantly in low-light conditions or extreme weather conditions. To address this issue, cross-modal fusion of infrared (IR) and visible light (RGB) spectra has emerged as a promising solution, effectively addressing the limitations of traditional visible light detectors in low-light environments. By incorporating multispectral data into deep learning models, the accuracy and stability of object detection can be further improved, providing autonomous vehicles with more comprehensive environmental perception capabilities. This cross-modal fusion detection solution not only maintains sensitivity to anomalies in complex weather conditions and during day-night cycles, but also provides more comprehensive and reliable environmental perception in real-time detection scenarios.

[0003] Patent CN119445306A, "Multispectral Target Detection Method, Apparatus, and Electronic Device Based on Feature Fusion," leverages the strengths of each image type, using the strengths of one modality to compensate for the limitations of another. This effectively addresses the blurriness of color images in harsh weather conditions or changing lighting environments. This patented method provides rich additional information for target recognition, thereby enhancing the robustness and accuracy of target detection.

[0004] Patent CN118426736A, "Bionic Compound Eye Multispectral Target Detection System and Method for Harsh Environments," utilizes a bionic compound eye multispectral imaging device to capture a target area in a harsh environment. The captured multi-view heterospectral image sequence is processed to obtain a pixel-aligned heterospectral image sequence, which is then fused. Furthermore, the patent utilizes an adaptive point sampling transformation model to capture sparse supervisory signals from the target boundary area of the image obtained in step three, using the coded features. This allows for rapid location of hidden targets.

[0005] Existing multispectral fusion methods usually ignore the heterogeneity between modalities and the feature differences of abnormal targets in different modalities, resulting in suboptimal feature extraction of the model and excessive computational overhead, resulting in insufficient detection accuracy of the model for unstructured abnormal objects.

[0006] Furthermore, existing bounding box regression methods for object detection, such as those based on the IoU loss function, suffer from slow convergence and inaccurate positioning in multispectral object detection scenarios. Furthermore, anomalous objects are often not regular rectangles. For example, a suddenly scattered object may be a long, thin strip; a pothole may appear irregularly curved but resemble an ellipse when viewed from above; and animals may exhibit different postures when entering the road. These limitations hinder the effective deployment of multispectral fusion technology in key applications such as autonomous driving and road anomaly detection.

[0007] Although the patent "CN119445306A Multispectral target detection method, device and electronic equipment based on feature fusion" and the patent "CN117115513A Personnel detection method based on cross-modal fusion model of multispectral target detection" have high detection accuracy, the position of their predicted anchor frames does not fit the real target detection objects well, and there is still room for improvement. Summary of the Invention

[0008] To solve the above problems, the present invention discloses a multimodal image target detection method based on heterogeneity-aware attention fusion network.

[0009] The specific technical solutions are as follows:

[0010] A multimodal image target detection method based on heterogeneity-aware attention fusion network includes the following steps:

[0011] S1. Visible light modality (X V ) and infrared mode (X I ) input data to perform hierarchical and progressive multi-scale feature extraction, establish a hierarchical feature pyramid, and generate multimodal image features;

[0012] S2. Introducing the CMHAP (Channel-Spatial Attention Module) at the shallow level of the feature pyramid. This module builds a channel-spatial bidirectional attention mechanism, dynamically generates a heterogeneous feature weight matrix, and enhances key sub-feature regions in the visible and infrared modalities.

[0013] S3. Through the refined channel attention fusion module (RCAF), the cross-modal features after interaction enhancement are adaptively fused to generate fused features;

[0014] S4. The fused multimodal features obtained in S3 and the multimodal image features output by S1 are jointly input into the detection head, and the ellipse dynamic CIoU loss function is used for target classification and bounding box regression. The ellipse dynamic CIoU loss function replaces the traditional rectangular anchor box with the ellipse geometric characteristics, adaptively selects the circumscribed ellipse or the inscribed ellipse to calculate the intersection-over-union ratio according to the target size, and introduces a dynamic scaling parameter to adaptively adjust the major and minor axes of the ellipse anchor box to optimize the geometric alignment between the anchor box and the real target.

[0015] The multi-scale feature extraction process of the dual-stream network backbone includes:

[0016] Initial feature extraction: At the shallow feature pyramid level, by stacking basic units consisting of convolutional layers (Conv), batch normalization (BN), and SiLU activation functions, preliminary texture and low-level semantic information are extracted to achieve preliminary feature extraction;

[0017] Multi-layer feature extraction: YOLOv5's C3 module is used to enhance basic features, and cross-layer feature interaction is strengthened through channel splitting, multi-path convolution, and residual connections;

[0018] Multi-scale pooling and attention enhancement: After concatenating the features extracted from the dual backbones, multi-scale pooling is performed using the SPPF (Spatial Pyramid Pooling-Fast) structure to obtain contextual information under different receptive fields. Channel-spatial attention is then enhanced through the feature refinement module and the C2PSA (Channel Spatial Attention) module. The C2PSA module assigns more discriminative channel weights through channel attention and captures key target areas through spatial attention.

[0019] The channel-space bidirectional attention mechanism includes the following formula:

[0020]

[0021] in and It means that after the visible light modality is grouped, the kth group of sub-features is split into two parts along the channel branch. The same is true for the infrared modality. and are the weighting coefficients learned through the cross-modal heterogeneous channel / spatial attention path, Indicates element-by-element multiplication enhancement at a channel or spatial position. shuffle[·] indicates channel shuffling after concatenating each group; It is the heterogeneous perception enhancement module for visible light and infrared modalities The sub-features are globally averaged pooled and a cross-modal channel attention descriptor is generated through a fully connected layer and a sigmoid activation function. The heterogeneous perception enhancement module performs channel normalization on another part of the sub-features and constructs a spatial attention descriptor through a fully connected layer and activation function.

[0022] The fusion process of the refined channel attention fusion module (RCAF) includes the following formula:

[0023]

[0024] and is the channel importance distribution of visible light and infrared modalities, that is, the cross-modal features output by CMHAP are spliced, and the channel attention weights are generated through global average pooling, fully connected layers, and Softmax activation function. ⊙ represents the channel-level multiplication modulation based on the corresponding attention weights. It represents channel-by-channel or element-by-element feature superposition, which is used to introduce residual feature preservation and enhancement.

[0025] The calculation steps of the ellipse dynamic CIoU loss function include:

[0026] Large object processing: When the object occupies more than half of the image area, the IoU is calculated using a circumscribed ellipse. The semi-axis of the ellipse is determined by the width and height of the rectangular anchor box. Dynamic parameters are introduced to adjust the size of the anchor box. The equation of the circumscribed ellipse is:

[0027]

[0028] Among them, x and y represent the center coordinates of the rectangular anchor box, and the upper left and lower right vertices of the rectangular anchor box are located at (x tl ,y tl ) and (x br ,y br), α∈(0,1.5) is a dynamic parameter, a and b represent the sizes of the major and minor axes of the external ellipse respectively;

[0029] Small target processing: When the real target is completely contained in the anchor box, the inscribed ellipse is used to calculate the IoU. The equation of the inscribed ellipse is:

[0030]

[0031] Among them, w and h represent the width and height of the rectangle anchor,

[0032] Finally, the ellipse CIoU can be calculated as follows:

[0033]

[0034] The heterogeneous perception enhancement module (CMHAP) exists at the P2 / P3 level of the backbone network.

[0035] The advantages of the present invention are: deep fusion of cross-modal complementary features: targeting the heterogeneous differences between visible light and infrared modalities, a channel-spatial bidirectional attention mechanism (CMHAP module) is designed to dynamically enhance cross-modal key area features, avoid information loss in traditional fusion methods, effectively integrate visible light texture details with infrared thermal radiation semantics, and improve the detection capability of low-contrast targets (such as pedestrians at night and hidden heat sources);

[0036] Adaptive Geometric Matching Strategy: This framework introduces an ellipse dynamic CIoU loss function, automatically selecting a circumscribed / inscribed ellipse fitting strategy based on the target scale. This framework, combined with dynamic scaling parameters and aspect ratio constraints, provides a more meaningful overlap measurement compared to rectangular baselines. This framework effectively improves detection accuracy in multimodal fusion scenarios while maintaining robust convergence properties.

[0037] The refined channel attention fusion module optimizes the fused representation through channel-adaptive weighting, effectively amplifying discriminative features while suppressing modality-specific noise. This refinement process outputs the fused feature representation, which, in the subsequent detection and recognition stages, better combines the strengths of visible light and infrared modalities, enabling accurate localization and classification of targets in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Schematic diagram of the target detection model structure based on heterogeneity-aware attention fusion network;

[0039] Figure 2 This is the logic block diagram of the heterogeneous perception module;

[0040] Figure 3 Schematic diagram of the refined channel attention fusion module. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0042] This paper proposes a heterogeneity perception attention fusion network, such as Figure 1 As shown in Figure 2, the overall network consists of three parts: a multimodal feature extraction backbone, a cross-modal heterogeneous interaction fusion module, and a detection head. Dynamic cross-modal feature enhancement is achieved through a bidirectional heterogeneous attention mechanism. The following describes the multimodal feature extraction dual-backbone fusion network and the novel elliptical dynamic CIoU loss function designed in this paper.

[0043] Multimodal feature extraction dual backbone network:

[0044] In order to allow the model to learn the visible light modality (X V ) and infrared mode (X I ) in terms of data distribution and feature representation, while taking into account the complementarity of the two modalities in the detection task, we constructed a two-stream network backbone. This backbone processes the input heterogeneous bimodal data {X V ,X I} to extract multi-scale features, and the obtained representations are recorded as Where b∈{P1,…,P5} corresponds to different levels in the feature pyramid. The specific process can be expressed as follows:

[0045]

[0046] The entire backbone network downsamples the image layer by layer in a bottom-up order and extracts multi-scale features to adapt to the scale and appearance changes of objects in complex scenes. To ensure that different levels can effectively represent the gradual transition from local texture to semantic information, we set up five feature pyramid levels (denoted as P1, P2, P3, P4, and P5). Among them, the high-level layers (P4 and P5) focus more on global semantic information and fused feature information, while the shallow layers (P1, P2, and P3) focus on finer-grained local structures. Features at different levels are also fused through horizontal or vertical paths to improve the overall detection effect and robustness.

[0047] At the shallow feature pyramid level, we first use several basic convolutional units to initially extract texture and low-level semantic information. Each unit consists of a convolution (Conv), batch normalization (BN), and an activation function SiLU, as shown in the following formula: After stacking this module n times, a preliminary feature representation is formed:

[0048]

[0049] Where n represents the number of times each unit is stacked.

[0050] After initial extraction of basic features, we employ the cascaded C3 module in the YOLOv5 model as a feature enhancement module to strengthen cross-layer feature interaction. The C3 module utilizes a branching structure to split input features along the channel dimension, extracting multi-scale features through different convolutional paths before fusing them. This results in a more discriminative representation while maintaining a manageable computational footprint. This cascaded approach further integrates semantic and spatial information at different scales, improving the performance of feature extraction from multimodal data.

[0051] After completing multi-layer feature extraction, we concatenate the multimodal image features extracted by the dual-backbone network and perform further fusion operations. Specifically, we use the SPPF (Spatial Pyramid Pooling-Fast) structure in the YOLOv11 model to perform multi-scale pooling of input features, obtaining contextual information under different receptive fields. The resulting multimodal image features are then passed through a feature refinement module to further enhance the focus on the target area and integrate multi-scale information. Next, we use C2PSA (Channel Spatial Attention) to implement channel-spatial attention enhancement. Specifically, this module assigns higher weight to channels with stronger discriminative power through channel attention (Channel Attention) and, on the other hand, uses spatial attention (Spatial Attention) to capture key areas related to the target on a two-dimensional plane. The synergistic effect of these two methods can greatly improve the network's target capture and discrimination capabilities in complex scenarios, thereby improving the problem of feature redundancy or weak representation that may arise when fusing heterogeneous modal data. The fused multimodal image features will be input into the neck and detection head together with the multimodal image fusion features obtained by the cross-modal heterogeneous interactive fusion module for efficient target detection.

[0052] Cross-modal heterogeneous interaction fusion module:

[0053] like Figure 1 Architecture diagram. We introduce a heterogeneous perception enhancement module at the P2 / P3 layer of the backbone network to build bidirectional cross-modal interaction:

[0054]

[0055] We use the Cross-Modality Heterogeneous Attention Perception (CMHAP) module, such as Figure 2 , establish a channel-space bidirectional attention mechanism and dynamically generate heterogeneous feature weight matrices and This module can achieve sub-feature key area enhancement in the strategic exchange of channel and spatial attention maps. The module can be expressed by the following formula:

[0056]

[0057] in and It represents the two parts obtained by grouping the k-th group of sub-features of the visible light modality and then splitting them along the channel branch (the same applies to the infrared modality). and are the weighting coefficients learned through the cross-modal heterogeneous channel / spatial attention path, Represents element-by-element multiplication enhancement at a channel or spatial position. shuffle[·] represents channel shuffling after concatenating each group to further promote cross-channel feature interaction.

[0058] Specifically, in the channel-heterogeneous attention path, the module pays attention to each semantic sub-feature Perform global average pooling (GAP) to extract channel global statistical features s V and s I , and then generate a cross-modal channel attention descriptor through a fully connected layer and activation function (sigmoid) and This process explicitly compares the global features of different modalities, thereby assigning differentiated weights in the channel dimension.

[0059] In the spatial heterogeneity attention path, this module performs channel normalization (GroupNormalization, GN) on the other half of the features and then uses the fully connected layer and activation function to construct the spatial attention descriptor and This highlights key heterogeneous regions in the spatial dimension. This process can be regarded as measuring the difference between visible light and infrared features at the local region level and dynamically enhancing them through attention weights.

[0060] Finally, the cross-modal heterogeneous attention perception module applies the attention weights learned from the above two paths, channel and space, to the original sub-feature map, and obtains the enhanced cross-modal heterogeneous information. and This process ensures feature diversity while allowing the two modalities to influence each other in attention allocation through a weight exchange mechanism, thereby effectively capturing the complementary and heterogeneous features across modalities.

[0061] To further align cross-modal features and eliminate possible geometric or semantic offsets, we use the Refined Channel Attentive Fusion (RCAF) module after the cross-modal heterogeneity attention perception module to build a more robust fusion representation. This module can be expressed as follows:

[0062]

[0063] After splicing and After global average pooling, full connection layer and Softmax activation function, the channel attention weight is obtained and (corresponding to the channel importance distribution of visible light and infrared modalities respectively). ⊙ represents the channel-level multiplicative modulation based on the corresponding attention weight, It represents channel-by-channel or element-by-element feature superposition, which is used to introduce residual feature preservation and enhancement.

[0064] Traditional fusion methods such as channel concatenation and element summation often introduce noise. The refined channel attention fusion module optimizes the fusion representation through channel adaptive weighting, effectively amplifying the discriminative features while suppressing modality-specific noise. The fused feature representation F is output through the refinement process. fuse ,In the subsequent neck-to-detection and recognition stage, it can better combine the advantages of visible light and infrared modalities to achieve accurate positioning and classification of targets in complex scenes.

[0065] Ellipse dynamic CIoU loss function:

[0066] Bounding box regression is a key component in object detection and plays a crucial role in target localization. As shown in Figure 4, the proposed elliptical dynamic CIoU loss function replaces the traditional rectangular bounding box with an elliptical representation, accelerating the slow convergence of bounding box regression and achieving more accurate predicted anchor box localization.

[0067] Existing IoU calculation methods rely primarily on the overlapping areas of rectangular anchor boxes, which suffer from inaccurate annotation. Rectangular anchor boxes cannot precisely conform to the object contour. We believe that the optimal anchor design should better align with the detection target to achieve more accurate loss calculation.

[0068] Inspired by InnerIoU and CIoU, this paper proposes a dynamic IoU calculation method based on ellipse, which uses the geometric properties of ellipse. The calculation method of CIoU is as follows:

[0069]

[0070] where ρ(·) represents the Euclidean distance and C represents the diagonal length of the minimum circumscribed rectangle.

[0071] We use the ellipse's major axis a and minor axis b to calculate the width w and height h of a traditional rectangular anchor, respectively, to construct an elliptical anchor that better matches the detection target. Furthermore, this paper introduces a dynamic scaling parameter α to adaptively adjust the anchor size based on the target size and shape, facilitating the flexible adaptation of the loss function in various detection scenarios.

[0072] Large target case (using circumscribed ellipse):

[0073] When the object occupies more than half of the image size, the IoU is calculated using the circumscribed ellipse. First, calculate the coordinates of the upper left and lower right corners of the anchor box:

[0074]

[0075] According to the ellipse circumscribed rectangle anchor theorem we proposed, a given vertex is located at (x tl ,y tl )(upper left) and (x br ,y br )(lower right) has a unique circumscribed ellipse whose semi-axes satisfy:

[0076]

[0077] where Δx = x br -x tl and Δy=y tl -y br Represent the width and height of the rectangular anchor respectively. Let α∈(0,1.5) be a dynamic parameter, then the equation of the circumscribed ellipse becomes:

[0078]

[0079] Small target case (using an inscribed ellipse):

[0080] For small-scale targets where the ground truth object is completely contained within the anchor box, we use the inscribed ellipse IoU calculation. The inscribed ellipse configuration is defined by aligning its major / minor axis with the anchor size:

[0081]

[0082] Where w and h represent the width and height of the anchor respectively. Then we can get the inscribed ellipse:

[0083]

[0084] Finally, combining the advantages of CIoU, integrating the center distance and aspect ratio constraints, the ellipse CIoU can be calculated as follows:

[0085]

[0086] where ρ(·) represents the Euclidean distance, and C represents the diagonal length of the minimum bounding rectangle. The remaining parameters are calculated in the same way as for CIoU. Elliptical CIoU enhances the geometric alignment between anchor boxes and objects through elliptical approximation, providing a more meaningful overlap measure compared to rectangular baselines. This framework effectively improves detection accuracy in multimodal fusion scenarios while maintaining robust convergence properties.

[0087] Beneficial effects of the present invention

[0088] The experiment of this invention selects the autonomous driving scene and the multispectral target detection dataset under low light conditions:

[0089] M3FD: A dataset for object detection in thermal infrared and visible light images. It provides 4,200 road scene image pairs at a resolution of 1024×768, all of which are registered. It contains 34,407 annotated instances covering six object categories: people, cars, buses, motorcycles, lamps, and trucks. It specifically includes extreme scenes such as tunnel entrances and exits.

[0090] FLIR: This dataset includes both daytime and nighttime scenes. Since the images in the original dataset are not aligned, the FLIR-aligned version was used for comparison in the experiments. It contains 5142 sets of aligned daytime and nighttime images with a resolution of 640×512. Three target categories, pedestrians, cars, and bicycles, are annotated. Motion blur and small object detection are the primary challenges. As shown in the table, all datasets are divided into training / validation / test sets, and the samples in the table are image pairs. The input size is uniformly resized to 1024×1024. Multi-scale scaling (±20%), HSV color perturbation (hue ±0.1, saturation / value ±0.5), and mosaic enhancement strategies are used during the training phase.

[0091] Table 1: Model comparison experiments on the M3FD multispectral dataset

[0092]

[0093] In order to verify the effectiveness of the framework proposed in this paper, we compared it with SeaDATE, EI2Det, CFT, and Scene-Adaptive CBAM methods on the M3FD dataset. Table 1 shows the mAP performance of our YOLO-HAFN and the above methods. 50 、mAP 50-95 Comparison results of the two evaluation indicators: the higher the two indicators, the better the performance. Scene-AdaptiveCBAM uses a lightweight scene adaptive fusion module, combined with a pre-trained single-modal model and scene classifier, to solve the drawback of traditional methods that require repeated training; EI2Det proposes a light-aware weighted and edge-guided fusion module to dynamically adjust the feature contribution of visible light and infrared to solve the problem of modal complementarity and positioning in dynamic lighting. Our YOLO-HAFN has a mAP of 1. 50 The evaluation index is 91.1%, which exceeds Scene-AdaptiveCBAM by 9.64%, EI2Det by 4.9%, and mAP by 1. 50-95 The results are 61.4%, surpassing EI2Det by 5.9% and CFT by 3.2%, respectively. Among the M3FD comparison methods, the smallest parameter-heavy model, Scene-Adaptive CBAM, has 26.7M parameters, while our method has only 3.4M parameters. This demonstrates that our method achieves both detection accuracy and model lightweightness, reducing computational effort.

[0094] Table 2: Model comparison experiments on the FLIR multispectral dataset

[0095]

[0096] In Table 2, the performance advantage on the FLIR dataset is also very significant. Compared with MMPedestron, which builds a unified encoder and adaptive fusion module to support multi-modal input and dynamic combination, although the proposed method has a lower mAP 50 The mAP is only 0.2% higher than that of MMPedestron, but the number of parameters in this paper is dozens of times less than that of MMPedestron, achieving high accuracy while meeting the requirements of lightweight model. For CMX, which has a larger number of parameters, the generalization problem of multimodal fusion is solved through bidirectional feature calibration and long-range context exchange. 50 There is still a difference of 4.4% compared with the method in this paper.

Claims

1. A multimodal image target detection method based on heterogeneity-aware attention fusion network, characterized in that: The steps include: S1. Visible light modality (X V ) and infrared mode (X I ) input data to perform hierarchical and progressive multi-scale feature extraction, establish a hierarchical feature pyramid, and generate multimodal image features; S2. Introducing the CMHAP (Channel-Spatial Attention Module) at the shallow level of the feature pyramid. This module builds a channel-spatial bidirectional attention mechanism, dynamically generates a heterogeneous feature weight matrix, and enhances key sub-feature regions in the visible and infrared modalities. S3. Through the refined channel attention fusion module (RCAF), the cross-modal features after interaction enhancement are adaptively fused to generate fused features; S4. The fused multimodal features obtained in S3 and the multimodal image features output by S1 are jointly input into the detection head, and the ellipse dynamic CIoU loss function is used for target classification and bounding box regression. The ellipse dynamic CIoU loss function replaces the traditional rectangular anchor box with the ellipse geometric characteristics, adaptively selects the circumscribed ellipse or the inscribed ellipse to calculate the intersection-over-union ratio according to the target size, and introduces a dynamic scaling parameter to adaptively adjust the major and minor axes of the ellipse anchor box to optimize the geometric alignment between the anchor box and the real target.

2. The method for multimodal image target detection based on heterogeneity-aware attention fusion network according to claim 1 is characterized in that: The multi-scale feature extraction process of the dual-stream backbone network includes: Initial feature extraction: At the shallow feature pyramid level, by stacking basic units consisting of convolutional layers (Conv), batch normalization (BN), and SiLU activation functions, texture and low-level semantic information are initially extracted to achieve shallow feature extraction; Multi-layer feature extraction: YOLOv5's C3 module is used to enhance basic features, and cross-layer feature interaction is strengthened through channel splitting, multi-path convolution, and residual connections; Multi-scale pooling and attention enhancement: After concatenating the features extracted from the dual backbones, multi-scale pooling is performed using the SPPF (Spatial Pyramid Pooling-Fast) structure to obtain contextual information under different receptive fields. Channel-spatial attention is then enhanced through the feature refinement module and the C2PSA (Channel Spatial Attention) module. The C2PSA module assigns more discriminative channel weights through channel attention and captures the key target areas through spatial attention.

3. The method for multimodal image target detection based on heterogeneity-aware attention fusion network according to claim 1 is characterized in that: The channel-space bidirectional attention mechanism Includes the following formulas: in and It means that after the visible light modality is grouped, the kth group of sub-features is split into two parts along the channel branch. The same is true for the infrared modality. and are the weighting coefficients learned through the cross-modal heterogeneous channel / spatial attention path, Indicates element-by-element multiplication enhancement at a channel or spatial position. shuffle[·] indicates channel shuffling after concatenating each group; It is the heterogeneous perception enhancement module for visible light and infrared modalities The sub-features are globally averaged pooled and a cross-modal channel attention descriptor is generated through a fully connected layer and a sigmoid activation function. The heterogeneous perception enhancement module is used to enhance the other part Channel normalization is performed and a spatial attention descriptor is constructed through a fully connected layer and an activation function.

4. The method for multimodal image target detection based on heterogeneity-aware attention fusion network according to claim 1 is characterized in that: The fusion process of the refined channel attention fusion module (RCAF) includes the following formula: and is the channel importance distribution of visible light and infrared modalities, that is, the cross-modal features output by CMHAP are spliced, and the channel attention weights are generated through global average pooling, fully connected layers, and Softmax activation function. ⊙ represents the channel-level multiplication modulation based on the corresponding attention weights. It represents channel-by-channel or element-by-element feature superposition, which is used to introduce residual feature preservation and enhancement.

5. The multimodal image target detection method according to claim 1, characterized in that: The calculation steps of the ellipse dynamic CIoU loss function include: Large object processing: When the object occupies more than half of the image area, the IoU is calculated using a circumscribed ellipse. The semi-axis of the ellipse is determined by the width and height of the rectangular anchor box. Dynamic parameters are introduced to adjust the size of the anchor box. The equation of the circumscribed ellipse is: Among them, x and y represent the center coordinates of the rectangular anchor box, and the upper left and lower right vertices of the rectangular anchor box are located at (x tl ,y tl ) and (x br ,y br ), α∈(0,1.5) is a dynamic parameter, a and b represent the sizes of the major and minor axes of the external ellipse respectively; Small target processing: When the real target is completely contained in the anchor box, the inscribed ellipse is used to calculate the IoU. The equation of the inscribed ellipse is: Among them, w and h represent the width and height of the rectangle anchor, Finally, the ellipse dynamic CIoU can be calculated as follows: where ρ(·) represents the Euclidean distance and C represents the diagonal length of the minimum circumscribed rectangle. IECIoU is the ellipse intersection over union ratio, ρ(·) represents the Euclidean distance, B gt and B pred is the distance between the true rectangle and the predicted rectangle.

6. The method for multimodal image target detection based on heterogeneity-aware attention fusion network according to claim 1, characterized in that: The heterogeneous perception enhancement module (CMHAP) exists at the P2 / P3 level of the backbone network.

Citation Information

Patent Citations

  • Personnel detection method of cross-modal fusion model based on multispectral target detection

    CN117115513A

  • Multispectral target detection method and device based on feature fusion and electronic equipment

    CN119445306A

Cited By

  • TOF image target detection model training method based on structure guide item

    CN120997485A

  • High-resolution remote sensing image road vehicle detection method based on feature fusion

    CN121033683A

  • Target detection method based on adaptive fusion of visible light and infrared features

    CN121214139A

  • A target detection method based on adaptive fusion of visible light and infrared features

    CN121214139B

  • Construction site dynamic adaptive target detection method, device and equipment and storage medium

    CN121582877A