Multi-modal data-based sparse enhancement dynamic fusion expert network target detection method

Through the sparse enhanced dynamic fusion expert network, the asymmetric multimodal sparse attention and light intensity expert-guided multimodal fusion sub-model are used to solve the problem of insufficient detection accuracy caused by sensor modality differences, and achieve efficient multimodal data fusion and target detection.

CN120808079APending Publication Date: 2025-10-17BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510723293.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In earthquake ruins search and rescue scenarios, the imaging principles, detection ranges, and environmental adaptability of different sensors vary significantly, making it difficult for a single modality to meet the target detection needs in complex scenarios. Existing multimodal fusion methods fail to effectively cope with illumination changes and target diversity.

Method used

A sparse enhanced dynamic fusion expert network is adopted to enhance the visible light and infrared image features respectively through an asymmetric multimodal sparse attention sub-model and a light intensity expert-guided multimodal fusion sub-model, and the fusion weights are dynamically adjusted according to the light intensity to achieve efficient fusion of multimodal data.

Benefits of technology

It improves the target detection accuracy in complex environments, enhances the single-modal feature representation, optimizes the multimodal fusion process, and improves detection accuracy and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808079A_ABST
    Figure CN120808079A_ABST
Patent Text Reader

Abstract

The invention provides a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data. The method comprises the following steps: acquiring a visible light image of a target area acquired by a visible light sensor and an infrared image of the target area acquired by an infrared sensor; inputting the visible light image and the infrared image into a sparse enhancement dynamic fusion model to obtain fused multi-modal features; the sparse enhancement dynamic fusion model is used for enhancing and fusing sensor data in different modes; and performing target detection according to the fused multi-modal features. According to the method provided by the embodiment of the invention, the detection precision in a complex environment is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target detection, and in particular to a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data. BACKGROUND

[0002] In the earthquake debris search and rescue scene, the unmanned aerial vehicle needs to rely on visible light, infrared, thermal imaging and other multi-modal sensors to cope with challenges such as light mutation, target diversity and environmental interference. However, the imaging principles, detection ranges and environmental adaptabilities of different sensors are significantly different. For example, visible light sensors are easily disturbed by night or smoke, and infrared imaging is easily disabled in high-temperature environments. Therefore, how to efficiently and accurately detect targets based on sensor data of different modalities is a technical problem that needs to be solved by those skilled in the art. SUMMARY

[0003] The present application provides a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data. Considering the significant differences in spectral characteristics, imaging principles and perception mechanisms between visible light images and infrared images, the visible light images and infrared images are input into a sparse enhancement dynamic fusion model to achieve enhancement and fusion of sensor data of different modalities, enhance single-modal feature representation, improve the feature extraction capability of the model for different modal data, optimize the multi-modal fusion process, and effectively improve the detection accuracy in complex environments.

[0004] The present application provides a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data, comprising the following steps.

[0005] Obtaining a visible light image of a target area collected by a visible light sensor and an infrared image of a target area collected by an infrared sensor; Inputting the visible light image and the infrared image into a sparse enhancement dynamic fusion model to obtain fused multi-modal features; the sparse enhancement dynamic fusion model is used for enhancing and fusing sensor data of different modalities; Performing target detection according to the fused multi-modal features.

[0006] According to the sparse enhancement dynamic fusion expert network target detection method based on multi-modal data provided by the present application, the sparse enhancement dynamic fusion model comprises an asymmetric multi-modal sparse attention sub-model and a light intensity expert guided multi-modal fusion sub-model; wherein, The asymmetric multi-modal sparse attention sub-model is used for enhancing the extracted visible light features in the visible light image and the extracted infrared features in the infrared image, filtering out important information in the visible light image and enhancing the context information of the infrared data, to obtain target visible light features and target infrared features; The light intensity expert guided multi-modal fusion sub-model is used for determining a weight value in a fusion process of the target visible light feature and the target infrared feature, and obtaining a fused multi-modal feature.

[0007] According to the application, a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data is provided. A plurality of sparse operators are used to perform sparse operations on the attention matrix corresponding to the visible light feature, to obtain a plurality of sparse operation results; different sparse operators in the plurality of sparse operators represent different sparse degrees. The plurality of sparse operation results are added to obtain the target visible light feature.

[0008] According to the application, a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data is provided. The infrared features of different sparse levels are combined to obtain the target infrared feature; wherein the infrared features of different sparse levels include detail features and semantic features.

[0009] According to the application, a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data is provided. According to the light intensity of the visible light image, a weight value in a fusion process of the target visible light feature and the target infrared feature is determined.

[0010] According to the application, a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data is provided. If the light intensity of the visible light image is less than the first threshold value, in the fusion process of the target visible light feature and the target infrared feature, the weight value of the target visible light feature is less than the weight value of the target infrared feature.

[0011] The application further provides a sparse enhancement dynamic fusion expert network target detection device based on multi-modal data, comprising the following modules. The acquisition module is used to acquire a visible light image of a target region collected by a visible light sensor and an infrared image of the target region collected by an infrared sensor. The fusion module is used to input the visible light image and the infrared image into a sparse enhancement dynamic fusion model to obtain a fused multi-modal feature; the sparse enhancement dynamic fusion model is used to enhance and fuse sensor data of different modalities. The detection module is used to perform target detection based on the fused multimodal features.

[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the sparse enhanced dynamic fusion expert network target detection method based on multimodal data as described above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described sparse enhanced dynamic fusion expert network target detection methods based on multimodal data.

[0014] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described sparse enhanced dynamic fusion expert network target detection methods based on multimodal data.

[0015] The sparse enhanced dynamic fusion expert network target detection method based on multimodal data provided by the present invention takes into account the significant differences between visible light images and infrared images in spectral characteristics, imaging principles and perception mechanisms. By inputting visible light images and infrared images into a sparse enhanced dynamic fusion model, it realizes the enhancement and fusion of sensor data of different modalities, enhances the single-modal feature representation, improves the model's feature extraction capability for different modal data, optimizes the multimodal fusion process, and effectively improves the detection accuracy in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 It is a flow chart of the sparse enhanced dynamic fusion expert network target detection method based on multimodal data provided by the present invention.

[0018] Figure 2 It is a structural diagram of the sparse enhanced dynamic fusion model provided by the present invention.

[0019] Figure 3 It is a structural diagram of the asymmetric multimodal sparse attention sub-model provided by the present invention.

[0020] Figure 4A structure schematic diagram of a multimodal fusion submodel of light intensity expert guidance provided by the application.

[0021] Figure 5 A structure schematic diagram of a multimodal fusion submodel of light intensity expert guidance provided by the application.

[0022] Figure 6 A structure schematic diagram of a multimodal fusion submodel of light intensity expert guidance provided by the application.

[0023] Figure 7 A structure schematic diagram of a multimodal fusion submodel of light intensity expert guidance provided by the application. DETAILED DESCRIPTION

[0024] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described below in connection with the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0025] The multimodal fusion submodel of light intensity expert guidance provided by the application will be described below in connection with the drawings in the present application. Figures 1-7 The multimodal fusion submodel of light intensity expert guidance provided by the application will be described below in connection with the drawings in the present application.

[0026] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described below in connection with the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0027] In the scene of earthquake debris rescue, the unmanned aerial vehicle needs to rely on visible light, infrared, thermal imaging and other multi-modal sensors to cope with challenges such as light mutation, target diversity and environmental interference. However, the imaging principles, detection ranges and environmental adaptabilities of different sensors are significantly different. For example, visible light sensors are easily disturbed at night or by smoke, and infrared imaging is easily invalid in high-temperature environments, which makes it difficult for a single mode to meet the needs of complex scenes. Therefore, multi-modal fusion emerges as the times require. By fusing sensor data of different modalities such as visible light and infrared, the advantages of each sensor can be fully utilized and the shortcomings of each other can be made up. According to the fusion strategy, multi-modal fusion is usually divided into three categories: early fusion (pixel-level fusion), late fusion (decision-level fusion) and mid-level fusion (feature-level fusion).

[0028] Early fusion is the most intuitive fusion method by directly aligning the original pixel data of different modalities and inputting a unified model. Commonly used pixel-level fusion methods include local guided fusion, bidirectional symmetric fusion, color normalization spectral sharpening (CNSS), principal component spectral sharpening transform (PCSST), and Gram-schmidt transform (GST). Overall, early fusion can retain rich underlying information, but it faces difficulties in modality alignment, high computational overhead, and insufficient dynamic adaptability to the environment. Late fusion inputs the inputs of different modalities into independent detection models, and then integrates the predicted results through statistics or rules. Chen et al. proposed a method of inputting the inputs of two modalities into single-modality target detection models to generate bounding boxes, and then integrating the predicted bounding boxes through probability integration, which allows the model to handle unaligned data. This method is more efficient, but it relies on high-precision single-modality target detection models. Therefore, late fusion is difficult to deal with complex background interference and target local feature loss due to its excessive reliance on high-level semantic information.

[0029] Compared with early and late fusion, mid-fusion extracts modality features through a double backbone network, and then designs a flexible feature interaction module, which has the potential to mine modality independence and deep correlation. It is the most commonly used multi-modal fusion method. Although existing research has made progress in intra-modality feature optimization and inter-modality redundancy suppression, there are still two limitations: first, environmental factors such as light intensity and smoke interference are not explicitly modeled to affect the difference in multi-modal feature representation, resulting in insufficient dynamic adaptation of the fusion strategy to the scene. For example, the MoE-Fusion dynamic image fusion framework proposed by Cao et al. realizes dynamic learning of multi-modal features through local expert mixing (MoLE) and global expert mixing (MoGE), but does not explicitly include environmental factors such as light intensity and smoke interference in the model, so the expert network cannot dynamically adjust the processing of multi-modal features according to environmental changes; second, there is a lack of cross-modal feature enhancement mechanism for diverse targets such as partially buried personnel and deformed wreckage, making it difficult to balance the fusion weights of global semantics and local details. For example, the CSSA module designed by Cao et al. considers both feature-level and spatial-level attention, but lacks the ability to dynamically adjust the fusion weight when fusing features, which cannot flexibly allocate the weights of global semantics and local details according to the characteristics of the target, thereby affecting the detection effect of diverse targets.

[0030] Figure 1 is one of the process schematic diagrams of the sparse enhancement dynamic fusion expert network target detection method based on multi-modal data provided by the application, as shown in Figure 1 The method comprises the following steps: Step 101, acquiring a visible light image of a target area collected by a visible light sensor and an infrared image of the target area collected by an infrared sensor.

[0031] Specifically, in the embodiments of the present application, first, the visible light image of the target area collected by the visible light sensor and the infrared image of the target area collected by the infrared sensor are acquired. Optionally, the target area can be a search and rescue area of earthquake ruins, or can be an area in other scenes, which is not limited in the embodiments of the present application.

[0032] Step 102, input the visible light image and the infrared image into a sparse enhancement dynamic fusion model to obtain a fused multi-modal feature; the sparse enhancement dynamic fusion model is used for enhancing and fusing sensor data of different modalities.

[0033] Specifically, after acquiring the visible light image of the target area collected by the visible light sensor and the infrared image of the target area collected by the infrared sensor, in the embodiments of the present application, the visible light image and the infrared image are input into a sparse enhancement dynamic fusion model to obtain a fused multi-modal feature. The sparse enhancement dynamic fusion model is used for enhancing and fusing sensor data of different modalities. That is, in the embodiments of the present application, considering that the visible light image and the infrared image have significant differences in spectral characteristics, imaging principles and perception mechanisms, through an asymmetric multi-modal sparse attention mechanism, the features of each modality are enhanced respectively, thereby enhancing the single-modal feature representation, improving the feature extraction capability of the model for different modal data, optimizing the multi-modal fusion process, solving the problems of current multi-modal fusion methods such as inconsistent multi-modal features and limited adaptability to light changes, realizing efficient fusion and accurate use of different modal features, and improving the detection accuracy in complex environments.

[0034] Step 103, performing target detection according to the fused multi-modal feature.

[0035] Specifically, after enhancing and fusing sensor data of different modalities based on the sparse enhancement dynamic fusion model, target detection can be accurately and efficiently realized through the fused multi-modal feature, thereby improving the detection accuracy in complex environments.

[0036] The method of the above embodiments considers that the visible light image and the infrared image have significant differences in spectral characteristics, imaging principles and perception mechanisms, and through inputting the visible light image and the infrared image into a sparse enhancement dynamic fusion model, the enhancement and fusion of sensor data of different modalities are realized, the single-modal feature representation is enhanced, the feature extraction capability of the model for different modal data is improved, the multi-modal fusion process is optimized, and the detection accuracy in complex environments is effectively improved.

[0037] In some embodiments, the sparse enhancement dynamic fusion model includes an asymmetric multi-modal sparse attention sub-model and a light intensity expert guided multi-modal fusion sub-model; wherein, The asymmetric multimodal sparse attention sub-model is used to enhance the visible light features extracted from the visible light image and the infrared features extracted from the infrared image, filtering out important information in the visible light image and enhancing the contextual information of the infrared data to obtain the target visible light features and target infrared features; The multimodal fusion sub-model guided by light intensity experts is used to determine the weight values ​​in the fusion process of target visible light features and target infrared features to obtain the fused multimodal features.

[0038] Specifically, the sparse enhanced dynamic fusion model in the embodiment of the present application includes an asymmetric multimodal sparse attention sub-model and a light intensity expert-guided multimodal fusion sub-model; wherein, the asymmetric multimodal sparse attention sub-model is used to enhance the visible light features extracted from the visible light image and the infrared features extracted from the infrared image, filter out important information in the visible light image, enhance the contextual information of the infrared data, and obtain the target visible light features and target infrared features, thereby enhancing the single-modal feature representation and improving the model's feature extraction capability for different modal data. That is, in the embodiment of the present application, taking into account the significant differences in spectral characteristics, imaging principles, and perception mechanisms between visible light images and infrared images, the features of each modality are enhanced separately through the asymmetric multimodal sparse attention mechanism. Optionally, the light intensity expert-guided multimodal fusion sub-model is used to determine the optimal contribution of each modality, thereby obtaining the fused multimodal features, realizing dynamic multimodal feature fusion, and improving detection accuracy in complex environments.

[0039] For example, the model structure of the sparse enhanced dynamic fusion model in the embodiment of the present application is as follows: Figure 2 As shown in the figure, in the sparse enhanced dynamic fusion expert network for multimodal target detection, the backbone network first extracts features from the input visible light image and infrared image respectively to obtain basic features; then the asymmetric multimodal sparse attention module enhances the characteristics of the two modal features to improve the single modal feature expression; then the intensity-guided multimodal hybrid expert module determines the weight according to the intensity of the visible light image, and adaptively fuses the enhanced two modal features; finally, target detection is completed based on the fused features, and the detection results are output.

[0040] In the method of the above embodiment, the asymmetric multimodal sparse attention sub-model in the sparse enhanced dynamic fusion model is used to enhance the visible light features extracted from the visible light image and the infrared features extracted from the infrared image, filter out important information in the visible light image and enhance the contextual information of the infrared data, thereby enhancing the single-modal feature representation and improving the model's feature extraction capability for different modal data; the light intensity expert-guided multimodal fusion sub-model in the sparse enhanced dynamic fusion model is used to determine the optimal contribution of each modality, realize dynamic multimodal feature fusion, and improve detection accuracy in complex environments.

[0041] In some embodiments, the asymmetric multi-modal sparse attention sub-model is specifically used for: performing sparse operations on the attention matrix corresponding to the visible light feature by using a plurality of sparse operators to obtain a plurality of sparse operation results; different sparse operators in the plurality of sparse operators represent different sparse degrees; adding the plurality of sparse operation results to obtain a target visible light feature.

[0042] Specifically, in the embodiments of the present application, the asymmetric multi-modal sparse attention sub-model is specifically used for performing sparse operations on the attention matrix corresponding to the visible light feature by using a plurality of sparse operators to obtain a plurality of sparse operation results; different sparse operators in the plurality of sparse operators represent different sparse degrees; and the plurality of sparse operation results are added to obtain a target visible light feature. That is, in the embodiments of the present application, considering the significant differences in spectral characteristics, imaging principles and perception mechanisms between visible light images and infrared images, the features of each modality are enhanced respectively through the asymmetric multi-modal sparse attention mechanism. Optionally, for the visible light feature, a learnable sparse-k operator is created to filter important information from the image.

[0043] For example, given a visible light feature , the visible light feature map is first processed by a dense connection layer. The specific operation of this step is to first perform dimension reduction, using a 1x1 convolution layer to reduce the channel number of the input visible light feature map to an intermediate channel number, then perform dense connection operation, and construct a dense block containing k 3x3 convolution layers. In the dense block, the input of each layer is obtained by concatenating the output features of the previous layers and the reduced features in the channel dimension, and after the convolution layer processing, the output channel number is always maintained as the intermediate channel number. Finally, the features are fused by concatenating the original visible light feature map and the features output by the last layer of the dense block in the channel dimension, and then adjusting the channel number to the output channel number through a convolution layer. In this way, the feature information at different levels can be fully utilized, avoiding the loss of feature information and ensuring maximum information transmission between network layers.

[0044] Then, the feature processed by the dense connection layer is input into a 3x3 convolution layer, and then linearly projected. This process generates queries (Q), keys (K) and values (V) for the feature map, as follows: wherein , , .

[0045] The standard Transformer computes self-attention by considering all tokens, which operates on the query, key and value matrices to get the dot-product attention, and the operation is defined as follows: However, one key limitation of this approach is that it computes the matching degree between the query and the key at all spatial positions, which ensures the integrity of the information, but contains a large amount of irrelevant information, which may involve noise interactions between irrelevant features, which is not conducive to feature extraction.

[0046] To solve this problem, the present application uses a sparse-k operator to perform sparse operations on the attention matrix obtained by the dot product of Q and The specific operation is to only retain the values of the first k key components in each row. Different k values represent different degrees of sparsity, and a larger k value can retain more detailed information, while a smaller k value enhances the ability to eliminate irrelevant information. Therefore, multiple different k values are taken to perform sparse-k operations on the attention matrix to obtain matrices with different degrees of sparsity, is a dynamically adjustable parameter, where H and W represent the height and width of the feature map, respectively. This dynamic selection makes the attention from dense to sparse. The specific process is as follows: where, represents retaining the first k values in each row of the matrix and setting the values at other positions to 0.

[0047] The weights of attention with different degrees of sparsity are multiplied by a learnable parameter to dynamically adjust: where, and are learnable parameters used to dynamically adjust the attention matrices with different degrees of sparsity.

[0048] Finally, the sum of the obtained attention matrices is added to the original feature after 1x1 convolution, and the final visible light feature is This way, both long-range dependencies and key semantic information in the image can be captured, while the basic and general feature representation of the image is retained, avoiding the loss of important basic information when emphasizing key features.

[0049] The method of the above embodiment uses multiple sparse operators to perform sparse operations on the attention matrix corresponding to the visible light features to obtain multiple sparse operation results; different sparse operators among the multiple sparse operators represent different degrees of sparsity; the multiple sparse operation results are added together to obtain the target visible light features, thereby not only capturing the long-distance dependencies and key semantic information in the visible light image, but also retaining the basic and general feature expression of the image, avoiding the loss of important basic information when emphasizing key features, and improving the accuracy of the detection results.

[0050] In some embodiments, the asymmetric multimodal sparse attention sub-model is specifically used to: The infrared features of different sparse levels are combined to obtain the target infrared features; among them, the infrared features of different sparse levels include detail features and semantic features.

[0051] Specifically, the asymmetric multimodal sparse attention mechanism structure is as follows: Figure 3 As shown, for infrared features, the symmetric multimodal sparse attention sub-model combines features of different sparse levels (low-level features and high-level features), so that the rich detail information contained in the low-level features, such as local information such as edges and textures, and the semantic information contained in the high-level features can be used at the same time to reflect global information such as the overall attributes of the target and enrich the expression of infrared features. Unlike the YOLO model, which is a gradual abstraction process from low to high-level feature extraction, the core operation of this application is to emphasize the combination of texture details and semantic information. Optionally, the detail features and semantic features in the infrared features can be extracted based on a neural network, which is not specifically limited in the embodiments of this application. Optionally, the spliced ​​features can be convolved and a rectified linear unit (RLU) can be applied. ) activation function, introducing nonlinear factors to enhance the model's learning ability. It also sets values ​​less than 0 to 0, achieving sparsification and reducing computational complexity. The resulting data is then added to the high-level features obtained through a separate 1×1 convolution operation to further enhance the high-level semantic features. This fully enhances the global and local contextual information of the infrared data. The processing of infrared features can be expressed as follows: The method of the above embodiment obtains the target infrared features by combining infrared features of different sparse levels; wherein, the infrared features of different sparse levels include detail features and semantic features, so that in the process of target detection, not only the rich detail information contained in the low-level features, such as local information such as edges and textures, can be utilized, but also the semantic information contained in the high-level features can be combined to enrich the expression of infrared features and improve the accuracy of target detection results.

[0052] In some embodiments, the multimodal fusion sub-model guided by the light intensity expert is specifically used to: According to the light intensity of the visible light image, a weight value in a target visible light feature and target infrared feature fusion process is determined.

[0053] Specifically, as shown in Figure 4 , in multi-modal target detection, the correlation between different modalities can vary. For example, under low light conditions, infrared images are often more useful than visible light images. However, in the presence of occlusion or in complex scenes, visible light images can provide more valuable information. Therefore, in the multi-modal feature fusion process, it is crucial to adjust the weight of each modality. One of the key factors in determining the weight of the two modalities is the light intensity of the visible light image in the image pair. The present application designs a light intensity expert guided multi-modal fusion module, which uses a multi-modal light intensity expert evaluation mechanism to determine the optimal contribution of each modality. Optionally, the light intensity representation of the image is determined by selecting the gray value with the highest frequency of occurrence in a region centered on the image center of gravity. The image light intensity value ) is input into a gating network together with enhanced visible light features and enhanced infrared features to aggregate features from different modalities and predict dynamic weights. Optionally, the features are concatenated along the channel dimension in the gating network, and then the concatenated features are applied with an average pooling operation as follows.

[0054] wherein is the position of the feature map . The dynamic adaptive weight is generated based on the modality using a gating function . Finally, a weighted sum is performed to obtain the final multi-modal dynamic fusion feature . That is, in the present application, the light intensity expert guided multi-modal fusion module can closely associate the weight of the visible light with the light intensity, and use learnable parameters to realize dynamic learning and adjustment of the importance of the light intensity, thereby effectively fusing visible light and infrared light and improving the accuracy of target detection.

[0055] In an embodiment, if the light intensity of the visible light image is greater than a first threshold value, the weight value of the target visible light feature in the target visible light feature and target infrared feature fusion process is greater than the weight value of the target infrared feature. If the light intensity of the visible light image is less than the first threshold value, the weight value of the target visible light feature in the target visible light feature and target infrared feature fusion process is less than the weight value of the target infrared feature.

[0056] Specifically, in multi-modal object detection, the correlation between different modalities can vary. For example, in low-light conditions, infrared images are often more useful than visible light images. However, in the presence of occlusions or in complex scenes, visible light images can provide more valuable information. Therefore, in the multi-modal feature fusion process, the weight of visible light is closely related to the intensity of light in the embodiments of the present application, that is, in the process of fusing target visible light features and target infrared features, if the light intensity of the visible light image is greater than the first threshold value, the weight value of the target visible light feature is greater than the weight value of the target infrared feature; if the light intensity of the visible light image is less than the first threshold value, the weight value of the target visible light feature is less than the weight value of the target infrared feature, thereby realizing effective fusion of visible light and infrared light and improving the accuracy of target detection.

[0057] In the method of the above embodiment, in the process of fusing target visible light features and target infrared features, if the light intensity of the visible light image is greater than the first threshold value, the weight value of the target visible light feature is greater than the weight value of the target infrared feature; if the light intensity of the visible light image is less than the first threshold value, the weight value of the target visible light feature is less than the weight value of the target infrared feature, thereby closely relating the weight of visible light to the intensity of light in the multi-modal feature fusion process, realizing effective fusion of visible light and infrared light, and improving the accuracy of target detection.

[0058] For example, to verify the effect of the technical scheme of the present application, experiments are conducted, and the specific results are as follows: (1) Data set description The present application trains and tests the present method on the Dronevehicle and LLVIP (Low-Light Vision visible-Infrared Paired) data sets. Dronevehicle is a large-scale RGB-infrared vehicle detection data set based on drones, containing 28439 pairs of images, covering urban roads, residential areas, parking lots and other scenes from day to night. The LLVIP data set is a visible-infrared paired data set for low-light vision, consisting of 15488 pairs of images collected from 26 locations between 6 pm and 10 pm, covering 24 night scenes and 2 day scenes. During training, on the Dronevehicle data set, we scale the input image to 840 x 712 for all experiments, and on the LLVIP data set, we scale the input image to 1280 x 1024 for all experiments, and the performance is reported on the validation data set.

[0059] (2) Experimental setup The present application uses NVIDIA RTX 4090 as a server device. The model is based on the PyTorch framework. During the training process, an automatically selected optimizer (auto) is used, with an initial learning rate of 1e-2, and the learning rate decays to 1e-4 as the training progresses. The total training epoch is 200, and the batch size is 16. We use the mean of Average Precision (mAP) as the evaluation indicator of the model precision.

[0060] (3) Ablation analysis Table 1 shows the ablation experiment results of the method of the present application on the DroneVehicle dataset. The original YOLOv8s model is used as the baseline, without modification, with a precision of 78.0%. After adding the asymmetric multi-modal sparse attention mechanism, the precision is improved to 80.2%, an increase of 2.2% compared to the baseline. After adding the light intensity guided multi-modal hybrid expert, the precision is improved to 80.0%, an increase of 2.0% compared to the baseline. When both modules are added, that is, the proposed method, the precision reaches 80.4%, an increase of 2.4%. Through the ablation experiment, it is proved that the two modules proposed can effectively improve the performance of multi-modal target detection, and the combination of the two can play a more significant synergistic effect, further optimizing the detection results.

[0061] Table 1

[0062] (4) Comparison with other methods Table 2 shows the average precision values of different vehicle types and processing speeds of eight network architectures on the DroneVehicle dataset. The experimental results show that the performance of the proposed method is at least 16.4% higher than that of UA-CMDet and at least 6.2% higher than that of C²Former-S²ANet. In terms of processing speed, the proposed method has an advantage with a processing speed of 69.7 FPS. These differences are due to the proposed method enhancing single-modal features through the asymmetric sparse attention mechanism and dynamically adjusting the weights of different modalities by the light intensity expert during the multi-modal fusion process.

[0063] Table 2

[0064] Table 3 presents the average precision values of six different methods on the LLVIP dataset. The proposed method achieves a map of 97.5% on this dataset, far surpassing other methods. The proposed method outperforms DM-Fusion by 9.4%, reaching 88.1%, which is the largest improvement of the proposed method over other methods on this dataset. Compared with YOLOAdaptor, the proposed method improves by 1%, reaching 96.5%. This indicates that the proposed method not only performs well in image detection but also performs excellently on natural scene datasets.

[0065] Table 3

[0066] (5) Visualization results Figure 5 Qualitative results on two multi-modal datasets are shown. It can be seen that the proposed method effectively enhances target features, suppresses background interference, and significantly improves the distinguishability of detected targets and backgrounds, providing more reliable basis for subsequent analysis and decision-making.

[0067] The present application proposes a sparse enhancement dynamic fusion expert network architecture to address the problems of multi-modal feature enhancement information redundancy and difficulty in adapting to light changes in existing multi-modal target detection fusion methods. The network proposes an asymmetric multi-modal sparse attention mechanism and a light intensity expert guided multi-modal fusion module, which helps to improve the performance of multi-modal target detection under different lighting conditions. Among them, the asymmetric multi-modal sparse attention mechanism enhances the features of visible light and infrared light respectively. For visible light, a learnable sparse-k operator is designed to adaptively combine attention scores of different scales, effectively filtering out important information in the image. For infrared light, the combination of sparse-level features enhances the global and local context information of infrared data. The light intensity expert guided multi-modal fusion module uses a multi-modal light intensity expert evaluation mechanism to determine the optimal contribution of each modality based on the light intensity of the image, achieving dynamic multi-modal feature fusion. The proposed method outperforms existing advanced methods in detection capability, promoting the application of multi-modal target detection technology in different scenarios.

[0068] The multi-modal data based sparse enhancement dynamic fusion expert network target detection device provided by the present application is described below. The multi-modal data based sparse enhancement dynamic fusion expert network target detection device described below can be correspondingly referred to the multi-modal data based sparse enhancement dynamic fusion expert network target detection method described above. The multi-modal data based sparse enhancement dynamic fusion expert network target detection device of the present application embodiment is as shown in Figure 6 includes: The acquisition module 610 is configured to acquire a visible light image of a target region collected by a visible light sensor and an infrared image of the target region collected by an infrared sensor. The fusion module 620 is configured to input the visible light image and the infrared image into a sparse enhancement dynamic fusion model to obtain fused multi-modal features; and the sparse enhancement dynamic fusion model is configured to enhance and fuse sensor data of different modalities. The detection module 630 is configured to perform target detection according to the fused multi-modal features.

[0069] Figure 7 An example of an entity structure diagram of an electronic device is shown, which can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 can communicate with each other through the communications bus 740. The processor 710 can invoke a logical instruction in the memory 730 to execute a sparse enhancement dynamic fusion expert network target detection method based on multi-modal data, which includes: acquiring a visible light image of a target region collected by a visible light sensor and an infrared image of the target region collected by an infrared sensor; inputting the visible light image and the infrared image into a sparse enhancement dynamic fusion model to obtain fused multi-modal features; the sparse enhancement dynamic fusion model is configured to enhance and fuse sensor data of different modalities; and performing target detection according to the fused multi-modal features.

[0070] In addition, the logical instruction in the memory 730 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0071] In another aspect, the present application also provides a computer program product comprising a computer program, the computer program being stored in a non-transitory computer-readable storage medium, and the computer program being executable by a processor to enable a computer to perform the method of sparse enhanced dynamic fusion expert network target detection based on multi-modal data, which comprises: acquiring a visible light image of a target region collected by a visible light sensor and an infrared image of the target region collected by an infrared sensor; inputting the visible light image and the infrared image into a sparse enhanced dynamic fusion model to obtain fused multi-modal features; the sparse enhanced dynamic fusion model is used for enhancing and fusing sensor data of different modalities; and performing target detection according to the fused multi-modal features.

[0072] In another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program being executable by a processor to implement the method of sparse enhanced dynamic fusion expert network target detection based on multi-modal data, which comprises: acquiring a visible light image of a target region collected by a visible light sensor and an infrared image of the target region collected by an infrared sensor; inputting the visible light image and the infrared image into a sparse enhanced dynamic fusion model to obtain fused multi-modal features; the sparse enhanced dynamic fusion model is used for enhancing and fusing sensor data of different modalities; and performing target detection according to the fused multi-modal features.

[0073] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on a plurality of network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.

[0074] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary universal hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0075] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A sparse enhanced dynamic fusion expert network target detection method based on multimodal data, characterized in that: include: Acquire a visible light image of the target area acquired by the visible light sensor and an infrared image of the target area acquired by the infrared sensor; Inputting the visible light image and the infrared image into a sparse enhanced dynamic fusion model to obtain fused multimodal features; The sparse enhanced dynamic fusion model is used to enhance and fuse sensor data of different modalities; Target detection is performed based on the fused multimodal features.

2. The sparse enhanced dynamic fusion expert network target detection method based on multimodal data according to claim 1 is characterized in that: The sparse enhanced dynamic fusion model includes an asymmetric multimodal sparse attention sub-model and a light intensity expert guided multimodal fusion sub-model; wherein, The asymmetric multimodal sparse attention sub-model is used to enhance the visible light features extracted from the visible light image and the infrared features extracted from the infrared image, filter out important information in the visible light image and enhance the context information of the infrared data, and obtain the target visible light features and target infrared features; The light intensity expert-guided multimodal fusion sub-model is used to determine the weight value in the fusion process of the target visible light feature and the target infrared feature to obtain the fused multimodal feature.

3. The sparse enhanced dynamic fusion expert network target detection method based on multimodal data according to claim 2 is characterized in that: The asymmetric multimodal sparse attention sub-model is specifically used to: Using multiple sparse operators, perform sparse operations on the attention matrix corresponding to the visible light feature to obtain multiple sparse operation results; Different sparse operators among the multiple sparse operators represent different sparsity levels; The multiple sparse operation results are added together to obtain the target visible light feature.

4. The sparse enhanced dynamic fusion expert network target detection method based on multimodal data according to claim 2 is characterized in that: The asymmetric multimodal sparse attention sub-model is specifically used to: Infrared features at different sparse levels are combined to obtain target infrared features; wherein the infrared features at different sparse levels include detail features and semantic features.

5. The sparse enhanced dynamic fusion expert network target detection method based on multimodal data according to claim 2 is characterized in that: The multimodal fusion sub-model guided by the light intensity expert is specifically used for: A weight value in a fusion process of the target visible light feature and the target infrared feature is determined according to the light intensity of the visible light image.

6. The sparse enhanced dynamic fusion expert network target detection method based on multimodal data according to claim 5 is characterized in that: If the light intensity of the visible light image is greater than a first threshold, in the process of fusing the target visible light feature and the target infrared feature, the weight value of the target visible light feature is greater than the weight value of the target infrared feature; If the light intensity of the visible light image is less than a first threshold, during the fusion process of the target visible light feature and the target infrared feature, the weight value of the target visible light feature is less than the weight value of the target infrared feature.

7. A sparse enhanced dynamic fusion expert network target detection device based on multimodal data, characterized in that: include: An acquisition module, configured to acquire a visible light image of a target area acquired by a visible light sensor and an infrared image of a target area acquired by an infrared sensor; a fusion module, configured to input the visible light image and the infrared image into a sparse enhanced dynamic fusion model to obtain fused multimodal features; The sparse enhanced dynamic fusion model is used to enhance and fuse sensor data of different modalities; The detection module is used to perform target detection based on the fused multimodal features.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the sparse enhanced dynamic fusion expert network target detection method based on multimodal data is implemented as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the sparse enhanced dynamic fusion expert network target detection method based on multimodal data is implemented as described in any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the sparse enhanced dynamic fusion expert network target detection method based on multimodal data is implemented as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Humanoid robot multi-modal sensing fusion method based on dynamic sparse activation

    CN121552445A