Multi-modal target detection method and device based on attention self-modulation fusion
By employing an attention-based self-modulation fusion method, shallow features from visible light and infrared modal images are extracted and fused, solving the problem of high computational cost in multimodal target detection and achieving efficient and lightweight multimodal target detection.
Patent Information
- Application Number
- CN202510932508.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-14
AI Technical Summary
Multimodal target detection is computationally expensive and has high model complexity, which limits its application in various scenarios.
A lightweight multimodal target detection method based on attention self-modulation fusion is achieved by acquiring shallow features from visible light and infrared modal images, performing feature fusion and enhancement.
It achieves efficient and lightweight multimodal target detection, reducing computational costs and improving detection accuracy.
Smart Images

Figure CN120953629A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of remote sensing image processing technology, and in particular to a multimodal target detection method and apparatus based on attention self-modulation fusion. Background Technology
[0002] In recent years, object detection technologies based on visible light or infrared imaging have made significant progress, leading to their widespread application in remote sensing, autonomous driving, medical diagnostics, traffic management, and industrial production. However, the performance of visible light object detection is adversely affected by insufficient lighting conditions and object occlusion. Similarly, infrared object detection suffers from issues related to temperature sensitivity and a lack of color information. To address these challenges, researchers have leveraged the complementarity of infrared and visible light modes, integrating information from both to enhance multimodal object detection, thereby overcoming the limitations of single-modal detection. This has yielded positive responses and rapid development.
[0003] In multimodal target detection tasks, common sensor data sources include optical and infrared images. To compensate for the limitations of single-modal data, data fusion techniques are typically employed to combine image information from different sensors. Based on the stage at which the fusion occurs, existing multimodal target detection methods are generally classified into three categories: pixel-level fusion, feature-level fusion, and decision-level fusion. Pixel-level fusion aims to combine image data from different modalities at the pixel level to generate richer, higher-quality fusion information, thereby improving detection accuracy. Decision-level fusion improves accuracy by independently detecting data from various sensors and then combining the results. In contrast, feature-level fusion integrates features from different modalities during the feature extraction stage to generate a comprehensive representation, allowing for a deeper exploration of the differences and relationships between modalities, thus improving detection performance. Therefore, feature-level fusion has been widely used in this field. However, this popularity often overlooks the significant computational overhead associated with it, thus limiting the wider applicability of multimodal detection algorithms.
[0004] In recent years, Vision Transformer and Vision Mamba have demonstrated outstanding performance in multimodal object detection because they overcome the limitations of convolutional neural networks in local perception and achieve global self-modulation. These models have proven the potential of self-modulation in effectively adjusting multiple modal features to improve object detection accuracy. However, despite the promising detection results achieved using these models, their high computational cost limits their application in various scenarios. Summary of the Invention
[0005] This application provides a multimodal target detection method and apparatus based on attention self-modulation fusion to solve the problems of high computational cost and high model complexity in multimodal target detection in the background art, and realizes efficient and lightweight multimodal target detection.
[0006] The first aspect of this application provides a multimodal target detection method based on attention self-modulation fusion, including the following steps:
[0007] Acquire visible light modal images and infrared modal images with the same labels;
[0008] Feature extraction is performed on the visible light modal image to obtain shallow features of the visible light modal, and feature extraction is performed on the infrared modal image to obtain shallow features of the infrared modal;
[0009] The visible light mode shallow features and the infrared mode shallow features are fused to obtain target fused features, and feature extraction is performed on the target fused features to obtain the deep semantic features of the target fused features;
[0010] The deep semantic features are aggregated at multiple scales, and the feature aggregation results are enhanced based on a preset feature enhancement mechanism to obtain multi-scale fused features. The multi-scale fused features are then detected to obtain detection results.
[0011] According to one embodiment of this application, the fusion processing of the visible light mode shallow features and the infrared mode shallow features to obtain the target fused features includes:
[0012] Based on a preset attention fusion mechanism, the shallow features of the visible light mode and the shallow features of the infrared mode are fused to obtain initial fused features;
[0013] Based on a preset feature modulation fusion mechanism, the initial fusion features are fused to obtain modulation fusion features;
[0014] Based on a preset channel shuffling and fusion mechanism, the modulation fusion features are fused to obtain the target fusion features;
[0015] The preset attention fusion mechanism is as follows:
[0016] F a =PAM(M rgb +M ir );
[0017] Among them, F a The output is the result of attention fusion, PAM is the position attention operation, and M is the output. rgb M represents the visible light modal characteristics after channel attention enhancement.ir These are the infrared modal features after channel attention enhancement.
[0018] According to one embodiment of this application, the step of fusing the initial fusion features based on a preset feature modulation fusion mechanism to obtain modulation fusion features includes:
[0019] The initial fusion feature is adjusted globally to obtain a globally adjusted initial fusion feature, and the initial fusion feature is adjusted locally to obtain a locally adjusted initial fusion feature.
[0020] The features after global and local adjustments are obtained based on the initial fusion features adjusted globally and the initial fusion features adjusted locally.
[0021] The residual connection result is obtained based on the initial fusion feature and the global and local adjusted fusion feature. The residual connection result is then subjected to feature mapping operation to obtain the modulation fusion feature.
[0022] The features after the fusion of global and local adjustments are as follows:
[0023] F m =Conv1×1(X g +Y l );
[0024] Among them, F m X represents the features resulting from the fusion of global and local adjustments. g Y is the initial fusion feature after global adjustment. l This represents the initial fusion features after local adjustments.
[0025] According to one embodiment of this application, the preset channel shuffling and fusion mechanism is as follows:
[0026] F b =reshape(F b ,B,G,C / G,H,W);
[0027] F b =transpose(F b ',1,2);
[0028] F c =reshape(F b (B,C,H,W);
[0029] Among them, F b ′ represents the reshaped feature, reshape represents the feature reshaping operation, B is the batch size, G is the number of channel groups, C / G is the number of channels per group, H is the feature height, W is the feature width, and F is the feature height.b "" represents the result of feature dimension swapping after reshaping, "transpose" represents feature dimension swapping, and F c This is the output of the channel shuffling and merging, where C is the number of channels.
[0030] According to one embodiment of this application, the preset feature enhancement mechanism is as follows:
[0031]
[0032] Among them, F LC The features are enhanced by channel attention, and x is the input for feature enhancement. For element-wise multiplication, LCAM() extracts attention information from the global average and local strongest responses of the input features in a two-branch manner and fuses them to obtain the final channel attention weights. LPAM() performs global average pooling and global max pooling along the channel dimension in a two-branch manner and fuses them to obtain the final positional attention weights. F LP The features enhanced by positional attention are used as the final feature enhancement output.
[0033] According to the multimodal target detection method based on attention self-modulation fusion in this application, shallow features of the visible light modality of the visible light modality image and shallow features of the infrared modality image are extracted. The shallow features of the visible light modality and the shallow features of the infrared modality are fused to obtain target fusion features, and feature extraction is performed to obtain deep semantic features. Multi-scale feature aggregation is performed on the deep semantic features, and the feature aggregation results are enhanced based on a preset feature enhancement mechanism to obtain multi-scale fusion features. The multi-scale fusion features are then detected to obtain the detection result. This solves the problems of high computational cost and high model complexity in multimodal target detection in the prior art, achieving efficient and lightweight multimodal target detection.
[0034] A second aspect of this application provides a multimodal target detection device based on attention self-modulation fusion, comprising:
[0035] The acquisition module is used to acquire visible light modal images and infrared modal images with the same label;
[0036] The extraction module is used to extract features from the visible light modal image to obtain shallow features of the visible light modal, and to extract features from the infrared modal image to obtain shallow features of the infrared modal;
[0037] The fusion module is used to fuse the shallow features of the visible light mode and the shallow features of the infrared mode to obtain the target fused features, and to extract features from the target fused features to obtain the deep semantic features of the target fused features;
[0038] The detection module is used to perform multi-scale feature aggregation on the deep semantic features, and to enhance the feature aggregation results based on a preset feature enhancement mechanism to obtain multi-scale fused features, and to detect the multi-scale fused features to obtain detection results.
[0039] According to one embodiment of this application, the fusion module is used for:
[0040] Based on a preset attention fusion mechanism, the shallow features of the visible light mode and the shallow features of the infrared mode are fused to obtain initial fused features;
[0041] Based on a preset feature modulation fusion mechanism, the initial fusion features are fused to obtain modulation fusion features;
[0042] Based on a preset channel shuffling and fusion mechanism, the modulation fusion features are fused to obtain the target fusion features;
[0043] The preset attention fusion mechanism is as follows:
[0044] F a =PAM(M rgb +M ir );
[0045] Among them, F a The output is the result of attention fusion, PAM is the position attention operation, and M is the output. rgb M represents the visible light modal characteristics after channel attention enhancement. ir These are the infrared modal features after channel attention enhancement.
[0046] According to one embodiment of this application, the fusion module is used for:
[0047] The initial fusion feature is adjusted globally to obtain a globally adjusted initial fusion feature, and the initial fusion feature is adjusted locally to obtain a locally adjusted initial fusion feature.
[0048] The features after global and local adjustments are obtained based on the initial fusion features adjusted globally and the initial fusion features adjusted locally.
[0049] The residual connection result is obtained based on the initial fusion feature and the global and local adjusted fusion feature. The residual connection result is then subjected to feature mapping operation to obtain the modulation fusion feature.
[0050] The features after the fusion of global and local adjustments are as follows:
[0051] F m=Conv1×1(X g +Y l );
[0052] Among them, F m X represents the features resulting from the fusion of global and local adjustments. g Y is the initial fusion feature after global adjustment. l This represents the initial fusion features after local adjustments.
[0053] According to one embodiment of this application, the preset channel shuffling and fusion mechanism is as follows:
[0054] F b =reshape(F b ,B,G,C / G,H,W);
[0055] F b =transpose(F b ',1,2);
[0056] F c =reshape(F b (B,C,H,W);
[0057] Among them, F b ′ represents the reshaped feature, reshape represents the feature reshaping operation, B is the batch size, G is the number of channel groups, C / G is the number of channels per group, H is the feature height, W is the feature width, and F is the feature height. b "" represents the result of feature dimension swapping after reshaping, "transpose" represents feature dimension swapping, and F c This is the output of the channel shuffling and merging, where C is the number of channels.
[0058] According to one embodiment of this application, the preset feature enhancement mechanism is as follows:
[0059]
[0060] Among them, F LC The features are enhanced by channel attention, and x is the input for feature enhancement. For element-wise multiplication, LCAM() extracts attention information from the global average and local strongest responses of the input features in a two-branch manner and fuses them to obtain the final channel attention weights. LPAM() performs global average pooling and global max pooling along the channel dimension in a two-branch manner and fuses them to obtain the final positional attention weights. F LP The features enhanced by positional attention are used as the final feature enhancement output.
[0061] According to the multimodal target detection device based on attention self-modulation fusion according to the embodiments of this application, the visible light modal shallow features of the visible light modal image and the infrared modal shallow features of the infrared modal image are extracted; the visible light modal shallow features and the infrared modal shallow features are fused to obtain target fusion features, and feature extraction is performed to obtain deep semantic features; the deep semantic features are aggregated at multiple scales, and the feature aggregation results are enhanced based on a preset feature enhancement mechanism to obtain multi-scale fusion features; the multi-scale fusion features are then detected to obtain the detection result. Thus, the problems of high computational cost and high model complexity in multimodal target detection in the prior art are solved, and efficient and lightweight multimodal target detection is achieved.
[0062] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the attention-based self-modulation fusion-based multimodal target detection method as described in the above embodiments.
[0063] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the attention-based self-modulation fusion-based multimodal target detection method as described in the above embodiments.
[0064] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0065] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0066] Figure 1 This is a schematic diagram of the structure of a lightweight self-modulation feature fusion network for multimodal target detection according to an embodiment of this application;
[0067] Figure 2 This is a flowchart of a multimodal target detection method based on attention self-modulation fusion provided according to an embodiment of this application;
[0068] Figure 3 This is a block diagram of a multimodal target detection device based on attention self-modulation fusion according to an embodiment of this application;
[0069] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0070] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0071] The following describes a multimodal target detection method and apparatus based on attention self-modulation fusion according to embodiments of this application, with reference to the accompanying drawings.
[0072] Before introducing the multimodal target detection method based on attention self-modulation fusion according to the embodiments of this application, we will first introduce the lightweight self-modulation feature fusion network for multimodal target detection involved in the multimodal target detection method based on attention self-modulation fusion of this application.
[0073] Those skilled in the art will understand that while models in related technologies can produce promising detection results, the significant computational cost of these models remains a concern. Therefore, a key challenge lies in developing low-power feature self-modulation techniques to facilitate lightweight development of multimodal target detection. To this end, this application provides a lightweight attention-guided self-modulation feature fusion network for multimodal target detection, which achieves high-performance inference with low computational cost. Specifically, to achieve efficient feature fusion, this invention introduces an innovative attention-guided self-modulation feature fusion module. Guided by an attention mechanism, this module promotes the fusion of deep features based on self-modulation of global and local information. Furthermore, a lightweight and flexible feature enhancement module is designed at the network neck to enhance attention to multimodal fusion features and reduce information loss.
[0074] Specifically, such as Figure 1 As shown, the input portion of this lightweight self-modulation feature fusion network includes a visible light image I. rgb and infrared image I irFirst, a symmetrical convolutional network structure is used to extract shallow texture features from visible light and infrared modalities. Then, the extracted shallow features from both modalities are input into an attention modulation fusion module. Specifically, key information is first extracted and fused through attention, then the feature representation is adjusted through global adaptation and local enhancement via feature modulation fusion, and finally, channel shuffling is used to improve feature interaction, further promoting the fused representation. In this way, the model obtains rich fused features adjusted from channel information, positional information, global information, and local information. The fused features are then subjected to deep extraction and input into the network neck for feature aggregation to improve the model's ability to detect multi-scale targets. During this process, a feature enhancement module reduces redundant information, increases attention to key information, and promotes feature aggregation. Finally, target detection is performed on the multi-scale features obtained from the neck to obtain the final output.
[0075] Therefore, the lightweight self-modulation feature fusion network for multimodal object detection proposed in this application significantly reduces the model burden while maintaining accurate detection performance by using a single feature fusion unit. The attention-guided self-modulation feature fusion module in this application utilizes attention information from different modalities to adaptively and deeply fuse global and local features based on response levels. This module generates rich and robust representations by progressively exploring the information in the fused features from coarse to fine. Furthermore, this application also proposes a feature enhancement module located at the network neck. This module emphasizes the key representations of multimodal fused features to reduce redundant information interference, improve multi-scale feature fusion, and thus enhance the detection capability for objects at different scales.
[0076] It should be noted that this specific implementation is written in Python and implemented using the classic deep learning framework PyTORCH. The Dataset and DataLoader functions in PyTORCH are used as data reading methods to read multimodal image data. The PyTORCH data reading and loading functions are well-known technologies in this field and will not be elaborated upon here.
[0077] The following section details the attention-based self-modulation fusion-based multimodal target detection method proposed in this application, which is applied to the aforementioned lightweight self-modulation feature fusion network.
[0078] Specifically, Figure 2 This is a flowchart illustrating a high-precision and high-efficiency multimodal target detection method based on attention self-modulation fusion, provided in an embodiment of this application.
[0079] like Figure 2 As shown, this attention-based self-modulation fusion-based multimodal target detection method includes the following steps:
[0080] In step S201, visible light modal images and infrared modal images with the same label are acquired.
[0081] Among them, visible light modal images and infrared modal images with the same label refer to pairs of images of the same scene or target captured in different imaging modalities, and these images are labeled as the same category or target in the dataset.
[0082] Furthermore, visible light modal images capture the reflected light information of the target within the visible spectrum, presenting color and texture details, while infrared modal images capture the thermal radiation information of the target, reflecting temperature distribution.
[0083] Optionally, embodiments of this application may simultaneously acquire visible light modal images and infrared modal images with the same label using a multimodal sensor, or acquire visible light modal images and infrared modal images with the same label using other methods in the prior art, without specific limitations.
[0084] In step S202, feature extraction is performed on the visible light modal image to obtain shallow features of the visible light modal, and feature extraction is performed on the infrared modal image to obtain shallow features of the infrared modal.
[0085] Specifically, pairs of multimodal images with the same label I rgb and I ir The input is fed into a lightweight attention-guided self-modulation feature fusion network model, where I rgb Represents a visible light modal image, I ir This represents an infrared modal image.
[0086] Furthermore, shallow feature information of visible light modal images and infrared modal images is extracted by a dual-branch symmetrical convolutional network structure, resulting in shallow features of visible light modal images and shallow features of infrared modal images.
[0087] In step S203, the shallow features of the visible light mode and the shallow features of the infrared mode are fused to obtain the target fused features, and the target fused features are extracted to obtain the deep semantic features of the target fused features.
[0088] Furthermore, in some embodiments, the shallow features of the visible light mode and the shallow features of the infrared mode are fused to obtain the target fused features, including: fusing the shallow features of the visible light mode and the shallow features of the infrared mode based on a preset attention fusion mechanism to obtain the initial fused features; fusing the initial fused features based on a preset feature modulation fusion mechanism to obtain the modulation fused features; and fusing the modulation fused features based on a preset channel shuffling fusion mechanism to obtain the target fused features.
[0089] Specifically, the obtained shallow features P of the visible light mode rgb and infrared mode shallow layer features P ir The data is fed into the attention modulation fusion module for fusion processing to obtain a comprehensive and rich fusion feature representation.
[0090] Furthermore, a pre-defined attention fusion mechanism is first used to fuse key information from different modalities. This pre-defined attention fusion mechanism is as follows:
[0091] F a =PAM(M rgb +M ir (1)
[0092] Among them, F a The output is the result of attention fusion, PAM is the position attention operation, and M is the output. rgb M represents the visible light modal characteristics after channel attention enhancement. ir These are the infrared modal features after channel attention enhancement.
[0093] Furthermore, in some embodiments, based on a preset feature modulation fusion mechanism, the initial fusion features are fused to obtain modulation fusion features, including: adjusting the initial fusion features globally to obtain globally adjusted initial fusion features, and adjusting the initial fusion features locally to obtain locally adjusted initial fusion features; obtaining globally and locally adjusted fused features based on the globally adjusted initial fusion features and the locally adjusted initial fusion features; obtaining residual connection results based on the initial fusion features and the globally and locally adjusted fused features, and performing feature mapping operations on the residual connection results to obtain modulation fusion features.
[0094] Specifically, the initial fused features after attention fusion are fed into the feature modulation fusion stage. In this stage, the initial fused features are first effectively adjusted through global adaptation and local enhancement. The features after global and local adjustment fusion are as follows:
[0095] F m =Conv1×1(X g +Y l (2)
[0096] Among them, F m X represents the features resulting from the fusion of global and local adjustments. g Y is the initial fusion feature after global adjustment. l This represents the initial fusion features after local adjustments.
[0097] To embed global feature information and adaptively adjust global structural information, embodiments of this application introduce variance to provide spatial variability information, helping the model focus on feature changes in different regions. Two learnable parameters, α and β, enable the model to dynamically adjust the contributions of variance and structural features, thereby adaptively fusing information from different spatial regions. This process can be represented as:
[0098] X g =X⊙U(φ(Conv1×1(α·X) s +β·σ 2 (X))))(3)
[0099] Where X represents F a The global branch input obtained after channel splitting, X s σ represents the global structure information obtained through adaptive max pooling. 2 To calculate the variance, φ is the GELU activation function, U is the nearest neighbor upsampling operation, and ⊙ is the element-wise multiplication operation.
[0100] Complementing local detail features with global features can yield more contextual information and enhance model representation. This application's embodiment designs a simple, near-bottleneck design for local feature capture using a local branch. Specifically, dimensionality is first increased by encoding local information using a 3x3 depthwise convolution on the local branch input feature Y. Then, a 1x1 convolution operation with GELU activation is used to obtain the dimensionality-reduced and enhanced local feature Y. l The formula is as follows:
[0101]
[0102] Where Y represents F a The local branch input obtained after channel splitting.
[0103] In the final stage of feature modulation fusion, this embodiment refines the features through feature mapping and residual connection operations while retaining richer representations to obtain the output of the feature modulation fusion stage:
[0104] F b =FM(F dr )+F dr (5)
[0105] Here, FM represents the feature mapping operation.
[0106] Furthermore, channel shuffling fusion, as the final fusion stage in this step, breaks the independence between channels by exchanging the channel order of different groups, enabling the network to exchange information between multiple channel groups. This approach can promote complementarity between different features with low computational cost, enhance feature diversity, and thus improve the model's expressive power. It can be represented by the following process:
[0107] F b = reshape(F) b ,B,G,C / G,H,W)(6)
[0108] F b " = transpose(F b ′,1,2)(7)
[0109] F c =reshape(F b ″,B,C,H,W)(8)
[0110] Among them, F b ′ represents the reshaped feature, reshape represents the feature reshaping operation, B is the batch size, G is the number of channel groups, C / G is the number of channels per group, C is the number of channels, and F is the number of channels per group. b After reshaping, the shape becomes (B, G, C / G, H, W), which can be understood as dividing C channels into G groups, with each group containing C / G channels; H is the feature height, W is the feature width, and F... b "This refers to the result of feature dimension exchange after reshaping. Transpose is the feature dimension exchange. By exchanging the group dimension G and the channel dimension C / G, the element distribution of each channel group is no longer limited to the original structure, but spans different groups, thereby enhancing the interactivity between different channels. F" c For the channel shuffle and fusion output, the reshape operation in formula (8) reshapes the channel back to its original shape to obtain the final channel shuffle and fusion output F. c .
[0111] Furthermore, deep feature extraction is performed on the target fusion features to obtain the deep semantic features of the fusion features. This embodiment of the application uses a convolutional network structure for deep fusion feature extraction, specifically the deep feature extraction structure of the CSPDarknet53 network.
[0112] In step S204, multi-scale feature aggregation is performed on the deep semantic features, and the feature aggregation result is enhanced based on a preset feature enhancement mechanism to obtain multi-scale fused features. The multi-scale fused features are then detected to obtain the detection result.
[0113] Specifically, the obtained deep semantic features are fed into the network neck for multi-scale feature aggregation to improve the model's ability to detect targets at different scales. At the same time, during the feature aggregation process, a feature enhancement module is used to reduce information loss and improve the salient feature representation of the fused information.
[0114] Furthermore, the fused multi-scale deep semantic features extracted from the deep learning process are aggregated at the network neck to improve the network's sensitivity to multi-scale targets. Each feature aggregation operation accepts two inputs for channel concatenation. During feature aggregation, this embodiment designs a lightweight feature enhancement module to remove redundant features and highlight key information at the network neck. The feature enhancement module effectively suppresses unimportant channel information and unimportant location regions by weighting channel and location features, which better aggregates significant features at different scales, thereby improving detection performance. The feature enhancement operation can be represented by the following process:
[0115]
[0116] Among them, F LC The features are enhanced by channel attention, and x is the input for feature enhancement. For element-wise multiplication, LCAM() extracts attention information from the global average and local strongest responses of the input features in a two-branch manner and fuses them to obtain the final channel attention weights. LPAM() performs global average pooling and global max pooling along the channel dimension in a two-branch manner and fuses them to obtain the final positional attention weights. F LP The features enhanced by positional attention are used as the final feature enhancement output.
[0117] Furthermore, the multi-scale fusion features obtained from the network neck are detected, and the final detection results are output.
[0118] To help those skilled in the art to understand more clearly and intuitively the beneficial effects of the multimodal target detection method based on attention self-modulation fusion proposed in this application, comparative experiments are conducted below to verify the beneficial effects of the present invention.
[0119] As shown in Table 1, the multimodal datasets used in this experiment are the DroneVehicle dataset and the VTUAV_det dataset from the perspective of UAVs, and the LLVIP dataset from the perspective of multispectral cameras.
[0120] The DroneVehicle dataset is primarily used for visible and infrared modal aerial vehicle detection, containing imagery data from different time periods throughout the day. This dataset features relatively small target scales and presents challenges such as poor target visibility in the visible light modal at night and thermal artifacts resembling vehicles in the infrared modal. The dataset contains 5 target classes and a total of 19,459 RGB-IR image pairs, with 17,990 pairs used as the training set and the remaining 1,469 pairs as the test set. Due to the significant differences between the infrared and visible light modalities, this dataset effectively reflects the model's performance in data complementarity and fusion. The LLVIP dataset is a large-scale pedestrian detection dataset acquired by a binocular camera in low-light conditions. It contains 15,488 strictly aligned visible-infrared modal images, with 12,025 pairs used as the training set and 3,463 pairs as the test set. This dataset features strict registration and alignment, and the target scale is more uniform and larger than the other two datasets, effectively reflecting the model's overall detection performance. The VTUAV_det dataset is a multimodal object detection dataset obtained by re-annotating the large-scale multimodal drone tracking dataset VTUAV. The dataset contains 11,392 image pairs used for training and 5,378 image pairs used as the test set. This dataset includes targets of three different scales: small, medium, and large. Furthermore, it does not have strict alignment and suffers from significant target occlusion and image blurring. Therefore, this dataset effectively reflects the model's detection performance in challenging scenarios such as multi-scale variations, target occlusion, and image alignment.
[0121] Table 1
[0122]
[0123] Furthermore, CFT (Method 1), SuperYOLO (Method 2), GHOST (Method 3), ADCNet (Method 4), ICAFusion (Method 5), GM-DETR (Method 6), CDC-YOLOFusion (Method 7), and the method of this invention were used for target detection tasks.
[0124] The classification evaluation index and model lightweighting evaluation index in this application embodiment adopt a quantitative evaluation method, and the evaluation indexes are as follows:
[0125] 1) mAP metric:
[0126] The mAP (middle accuracy) metric is an authoritative evaluation indicator for object detection problems. A higher mAP metric generally indicates higher accuracy. The mAP metric is calculated as follows:
[0127] First, calculate the precision and recall for each category:
[0128]
[0129]
[0130] Where, N TP A true positive result, N FP False positive, N FN The number of false negatives. A true positive indicates that the object was correctly identified, a false positive indicates that the object was incorrectly detected, and a false negative is an instance of the object that was missed in the detection results. The specific formula for Average Precision is as follows:
[0131]
[0132]
[0133] The mean precision (AP) is determined by averaging the precision of a set of recalls S, using 11 numbers S = {0, 0.1, ..., 1}. C represents the categories of the dataset. When calculating AP, thresholds range from 0.5 to 0.95 with a step size of 0.05 to calculate AP at different thresholds. For example, if the intersection of predicted labels exceeds 0.5, it is classified as TP; otherwise, it is classified as FP. Furthermore, when multiple objects have high Intersection over Union (IOU) values with the ground truth, the object with the highest detection confidence is typically designated as TP. Subsequently, AP is calculated using the average precision within the equally spaced recall range represented by the set S. Finally, AP results for multiple categories are obtained, and the average AP for each category yields the evaluation metric mAP.
[0134] 2) Number of model parameters and GFLOPs metric:
[0135] The number of parameters measures a model's storage requirements and is a key metric for evaluating its availability and deployment difficulty on storage-constrained devices. GFLOPs (gigaflops per second) is an important metric for evaluating a model's computational requirements and efficiency during inference. The calculation methods for model parameter count and GFLOPs differ for different network architectures, therefore there is no single, universally accepted formula.
[0136] The comparative experimental results of different methods used in the embodiments of this application for target detection tasks are shown in Table 2:
[0137] Table 2
[0138]
[0139] As can be seen from Table 2, the method of the present invention can achieve higher mAP index and lower model parameter number and GFLOPs index, indicating that the method of the present invention has better detection capability while achieving model lightweighting.
[0140] Therefore, it can be concluded that, compared with traditional multimodal target detection methods, the method of this invention has a lighter model structure and higher detection accuracy. This invention fully considers the significant computational cost generated during multimodal fusion, achieving a lighter and more efficient multimodal detection method.
[0141] According to the multimodal target detection method based on attention self-modulation fusion in this application, shallow features of the visible light modality of the visible light modality image and shallow features of the infrared modality image are extracted. The shallow features of the visible light modality and the shallow features of the infrared modality are fused to obtain target fusion features, and feature extraction is performed to obtain deep semantic features. Multi-scale feature aggregation is performed on the deep semantic features, and the feature aggregation results are enhanced based on a preset feature enhancement mechanism to obtain multi-scale fusion features. The multi-scale fusion features are then detected to obtain the detection result. This solves the problems of high computational cost and high model complexity in multimodal target detection in the prior art, achieving efficient and lightweight multimodal target detection.
[0142] Next, referring to the accompanying drawings, a multimodal target detection device based on attention self-modulation fusion proposed according to an embodiment of this application is described.
[0143] Figure 3 This is a block diagram of a multimodal target detection device based on attention self-modulation fusion according to an embodiment of this application.
[0144] like Figure 3 As shown, the multimodal target detection device 10 based on attention self-modulation fusion includes: an acquisition module 100, an extraction module 200, a fusion module 300, and a detection module 400.
[0145] The system comprises the following modules: an acquisition module 100, which acquires visible light modal images and infrared modal images with the same label; an extraction module 200, which extracts features from the visible light modal images to obtain shallow features of the visible light modal images and extracts features from the infrared modal images to obtain shallow features of the infrared modal images; a fusion module 300, which fuses the shallow features of the visible light modal images and the shallow features of the infrared modal images to obtain target fusion features, and extracts features from the target fusion features to obtain deep semantic features of the target fusion features; and a detection module 400, which performs multi-scale feature aggregation on the deep semantic features, enhances the feature aggregation results based on a preset feature enhancement mechanism to obtain multi-scale fusion features, and detects the multi-scale fusion features to obtain detection results.
[0146] Further, in some embodiments, the fusion module 300 is used to: fuse the shallow features of the visible light mode and the shallow features of the infrared mode based on a preset attention fusion mechanism to obtain initial fused features; fuse the initial fused features based on a preset feature modulation fusion mechanism to obtain modulation fused features; and fuse the modulation fused features based on a preset channel shuffling fusion mechanism to obtain target fused features; wherein the preset attention fusion mechanism is:
[0147] F a =PAM(M rgb +M ir );
[0148] Among them, F a The output is the result of attention fusion, PAM is the position attention operation, and M is the output. rgb M represents the visible light modal characteristics after channel attention enhancement. ir These are the infrared modal features after channel attention enhancement.
[0149] Further, in some embodiments, the fusion module 300 is configured to: adjust the initial fusion features globally to obtain globally adjusted initial fusion features, and adjust the initial fusion features locally to obtain locally adjusted initial fusion features; obtain globally and locally adjusted fused features based on the globally adjusted initial fusion features and the locally adjusted initial fusion features; obtain residual connection results based on the initial fusion features and the globally and locally adjusted fused features; and perform feature mapping operations on the residual connection results to obtain modulation fusion features; wherein the globally and locally adjusted fused features are:
[0150] F m =Conv1×1(X g +Y l );
[0151] Among them, F m X represents the features resulting from the fusion of global and local adjustments. g Y is the initial fusion feature after global adjustment. l This represents the initial fusion features after local adjustments.
[0152] Furthermore, in some embodiments, the preset channel shuffling and fusion mechanism is as follows:
[0153] F b =reshape(F b ,B,G,C / G,H,W);
[0154] F b =transpose(Fb ',1,2);
[0155] F c =reshape(F b (B,C,H,W);
[0156] Among them, F b ′ represents the reshaped feature, reshape represents the feature reshaping operation, B is the batch size, G is the number of channel groups, C / G is the number of channels per group, H is the feature height, W is the feature width, and F is the feature height. b "" represents the result of feature dimension swapping after reshaping, "transpose" represents feature dimension swapping, and F c This is the output of the channel shuffling and merging, where C is the number of channels.
[0157] Furthermore, in some embodiments, the preset feature enhancement mechanism is as follows:
[0158]
[0159] Among them, F LC The features are enhanced by channel attention, and x is the input for feature enhancement. For element-wise multiplication, LCAM() extracts attention information from the global average and local strongest responses of the input features in a two-branch manner and fuses them to obtain the final channel attention weights. LPAM() performs global average pooling and global max pooling along the channel dimension in a two-branch manner and fuses them to obtain the final positional attention weights. F LP The features enhanced by positional attention are used as the final feature enhancement output.
[0160] It should be noted that the foregoing explanation of the multimodal target detection method based on attention self-modulation fusion also applies to the multimodal target detection device based on attention self-modulation fusion in this embodiment, and will not be repeated here.
[0161] According to the multimodal target detection device based on attention self-modulation fusion according to the embodiments of this application, the visible light modal shallow features of the visible light modal image and the infrared modal shallow features of the infrared modal image are extracted; the visible light modal shallow features and the infrared modal shallow features are fused to obtain target fusion features, and feature extraction is performed to obtain deep semantic features; the deep semantic features are aggregated at multiple scales, and the feature aggregation results are enhanced based on a preset feature enhancement mechanism to obtain multi-scale fusion features; the multi-scale fusion features are then detected to obtain the detection result. Thus, the problems of high computational cost and high model complexity in multimodal target detection in the prior art are solved, and efficient and lightweight multimodal target detection is achieved.
[0162] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0163] The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0164] When the processor 402 executes the program, it implements the multimodal target detection method based on attention self-modulation fusion provided in the above embodiments.
[0165] Furthermore, electronic devices also include:
[0166] Communication interface 403 is used for communication between memory 401 and processor 402.
[0167] The memory 401 is used to store computer programs that can run on the processor 402.
[0168] The memory 401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0169] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0170] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0171] Processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0172] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described multimodal target detection method based on attention self-modulation fusion.
[0173] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0174] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0175] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A multimodal target detection method based on attention self-modulation fusion, characterized in that, Includes the following steps: Acquire visible light modal images and infrared modal images with the same labels; Feature extraction is performed on the visible light modal image to obtain shallow features of the visible light modal, and feature extraction is performed on the infrared modal image to obtain shallow features of the infrared modal; The visible light mode shallow features and the infrared mode shallow features are fused to obtain target fused features, and feature extraction is performed on the target fused features to obtain the deep semantic features of the target fused features; The deep semantic features are aggregated at multiple scales, and the feature aggregation results are enhanced based on a preset feature enhancement mechanism to obtain multi-scale fused features. The multi-scale fused features are then detected to obtain detection results.
2. The method according to claim 1, characterized in that, The process of fusing the visible light mode shallow features and the infrared mode shallow features to obtain the target fused features includes: Based on a preset attention fusion mechanism, the shallow features of the visible light mode and the shallow features of the infrared mode are fused to obtain initial fused features; Based on a preset feature modulation fusion mechanism, the initial fusion features are fused to obtain modulation fusion features; Based on a preset channel shuffling and fusion mechanism, the modulation fusion features are fused to obtain the target fusion features; The preset attention fusion mechanism is as follows: F a =PAM(M rgb +M ir ); Among them, F a The output is the result of attention fusion, PAM is the position attention operation, and M is the output. rgb M represents the visible light modal characteristics after channel attention enhancement. ir These are the infrared modal features after channel attention enhancement.
3. The method according to claim 2, characterized in that, The preset feature modulation fusion mechanism performs fusion processing on the initial fusion features to obtain modulation fusion features, including: The initial fusion feature is adjusted globally to obtain a globally adjusted initial fusion feature, and the initial fusion feature is adjusted locally to obtain a locally adjusted initial fusion feature. The features after global and local adjustments are obtained based on the initial fusion features adjusted globally and the initial fusion features adjusted locally. The residual connection result is obtained based on the initial fusion feature and the global and local adjusted fusion feature. The residual connection result is then subjected to feature mapping operation to obtain the modulation fusion feature. The features after the fusion of global and local adjustments are as follows: F m =Conv1×1(X g +Y l ); Among them, F m X represents the features resulting from the fusion of global and local adjustments. g Y is the initial fusion feature after global adjustment. l This represents the initial fusion features after local adjustments.
4. The method according to claim 2, characterized in that, The preset channel shuffling and fusion mechanism is as follows: F b ′=reshape(F b ,B,G,C / G,H,W); F b ″=transpose(F b ′,1,2); F c =reshape(F b ″,B,C,H,W); Among them, F b ′ represents the reshaped feature, reshape represents the feature reshaping operation, B is the batch size, G is the number of channel groups, C / G is the number of channels per group, H is the feature height, W is the feature width, and F is the feature height. b "" represents the result of feature dimension swapping after reshaping, "transpose" represents feature dimension swapping, and F c This is the output of the channel shuffling and merging, where C is the number of channels.
5. The method according to claim 1, characterized in that, The preset feature enhancement mechanism is as follows: Among them, F LC The features are enhanced by channel attention, and x is the input for feature enhancement. For element-wise multiplication, LCAM() extracts attention information from the global average and local strongest responses of the input features in a two-branch manner and fuses them to obtain the final channel attention weights. LPAM() performs global average pooling and global max pooling along the channel dimension in a two-branch manner and fuses them to obtain the final positional attention weights. F LP The features enhanced by positional attention are used as the final feature enhancement output.
6. A multimodal target detection device based on attention self-modulation fusion, characterized in that, include: The acquisition module is used to acquire visible light modal images and infrared modal images with the same label; The extraction module is used to extract features from the visible light modal image to obtain shallow features of the visible light modal, and to extract features from the infrared modal image to obtain shallow features of the infrared modal; The fusion module is used to fuse the shallow features of the visible light mode and the shallow features of the infrared mode to obtain the target fused features, and to extract features from the target fused features to obtain the deep semantic features of the target fused features; The detection module is used to perform multi-scale feature aggregation on the deep semantic features, and to enhance the feature aggregation results based on a preset feature enhancement mechanism to obtain multi-scale fused features, and to detect the multi-scale fused features to obtain detection results.
7. The apparatus according to claim 6, characterized in that, The fusion module is used for: Based on a preset attention fusion mechanism, the shallow features of the visible light mode and the shallow features of the infrared mode are fused to obtain initial fused features; Based on a preset feature modulation fusion mechanism, the initial fusion features are fused to obtain modulation fusion features; Based on a preset channel shuffling and fusion mechanism, the modulation fusion features are fused to obtain the target fusion features; The preset attention fusion mechanism is as follows: F a =PAM(M rgb +M ir ); Among them, F a The output is the result of attention fusion, PAM is the position attention operation, and M is the output. rgb M represents the visible light modal characteristics after channel attention enhancement. ir These are the infrared modal features after channel attention enhancement.
8. The apparatus according to claim 7, characterized in that, The fusion module is used for: The initial fusion feature is adjusted globally to obtain a globally adjusted initial fusion feature, and the initial fusion feature is adjusted locally to obtain a locally adjusted initial fusion feature. The features after global and local adjustments are obtained based on the initial fusion features adjusted globally and the initial fusion features adjusted locally. The residual connection result is obtained based on the initial fusion feature and the global and local adjusted fusion feature. The residual connection result is then subjected to feature mapping operation to obtain the modulation fusion feature. The features after the fusion of global and local adjustments are as follows: F m =Conv1×1(X g +Y l ); Among them, F m X represents the features resulting from the fusion of global and local adjustments. g Y is the initial fusion feature after global adjustment. l This represents the initial fusion features after local adjustments.
9. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal target detection method based on attention self-modulation fusion as described in any one of claims 1-5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor to implement the attention-based self-modulation fusion-based multimodal target detection method as described in any one of claims 1-5.
Citation Information
Cited By
Cross-modal semantic segmentation system and method
CN121482790A
Cross-modal semantic segmentation system and method
CN121482790B
Multi-modal fusion target detection method, system and device and storage medium
CN122049545A