Multi-modal target detection method based on dynamic illumination perception random mask

By introducing a cross-modal differential feature interaction module and dynamic light perception mask strategy in the YOLOv8n network, the problem of insufficient detection accuracy in the lighting environment in multimodal object detection is solved, more efficient multimodal feature learning and recognition is achieved, and the target detection performance of drones and satellite platforms is improved.

CN120259294AActive Publication Date: 2025-07-04NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510734298.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The existing deep learning object detection algorithm lacks the mining of complementary information between dual modes in multimodal image fusion, resulting in insufficient detection accuracy in complex lighting environments, especially in low-light or dull scenes at night, and insufficient infrared image texture information, making it difficult to achieve multi-condition work throughout the day.

Method used

Using a dual-flow backbone network based on YOLOv8n, a cross-modal differential feature interaction module (CDFIM) and a dynamic light-aware random mask training strategy are designed, and inter-modal feature representation is enhanced through differential feature interaction, and the mask processing is dynamically adjusted using the illumination information of the visible light image to balance inter-modal attention and improve the robustness of the model to different lighting environments.

Benefits of technology

The recognition accuracy and robustness of the multimodal object detection model under complex lighting conditions are improved, the multimodal remote sensing object detection capabilities of drones and satellite platforms are enhanced, and the computing burden is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259294A_ABST
    Figure CN120259294A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal target detection method based on a dynamic illumination perception random mask, relates to the field of multi-modal target detection, and provides a dynamic illumination perception random mask method for performing mask processing on a part of proportional images of a training set and considering the influence of exposure and a strong light environment. Enabling the network to learn multi-modal effective complementary features in an unbiased manner; constructing a bimodal feature level fusion target detection basic network; a cross-modal differential feature interaction module is introduced, and complementary information interaction between double modals is increased by extracting difference features between visible light and infrared modals and dynamically allocating attention weights on a channel-space dimension; and performing model training and testing. According to the multi-modal target detection method provided by the invention, the robustness of the multi-modal target detection model to complex illumination environments such as exposure and strong light is improved, and the multi-modal remote sensing target detection capability of an unmanned aerial vehicle platform / satellite platform and the target recognition precision under the complex illumination condition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-modal object detection, and particularly to a multi-modal object detection method based on dynamic light perception random mask. Background Art

[0002] With the remarkable success of deep learning technology in the field of natural images, deep learning object detection algorithms have been applied in various fields. However, most deep learning algorithms are designed and applied specifically for a single modality. Although visible light images have rich target texture and color information, using visible images alone in low-light or no-light scenes at night cannot effectively monitor targets. Thermal infrared images have the advantage of being less affected by light, but infrared images are often accompanied by low texture information, low resolution, and high noise levels. The above defects of single modality lead to difficult-to-satisfactory accuracy in object detection and inability to achieve all-day multi-condition operation, restricting their application scope.

[0003] Therefore, fusing the complementary information of visible light images and infrared images for object detection has begun to attract the attention of scholars. However, existing methods lack attention to the complementary information between the two modalities, insufficiently mine cross-modal complementary information, and are prone to learning redundant information. In addition, current research focuses on the design of fusion methods and ignores the design of information interaction methods between modalities. The detection accuracy for complex targets such as those in exposure, strong light, dense, and low contrast with the background needs to be improved.

[0004] To solve the above problems, the present invention proposes a multi-modal object detection method based on light perception random mask. Based on the YOLOv8n dual-stream backbone network, the present invention designs a cross-modal differential feature interaction module (CDFIM) to deeply mine and learn the complementary features of visible light and infrared images, enhance the interaction and feature representation between modalities, and improve the learning ability of dual-modal complementary features. In addition, the present invention designs a light perception random mask training strategy to add dynamic complementary masks to the dual-modal images according to the local light information of visible light, forcing the model to balance the learning of cross-modal complementary features during training and improving the robustness to different light environments. Summary of the Invention

[0005] To solve the above problems, the present application proposes a multi-modal object detection method based on dynamic light perception random mask, including the following steps: S1. Obtain the image to be detected, and preprocess the image to be detected to obtain an initial image; S2. Based on YOLOv8n and an improved cross-modal differential feature interaction module, construct an improved multi-modal detection model, obtain a multi-modal object detection dataset, and train the improved multi-modal detection model using a dual-modal model training strategy that combines the multi-modal object detection dataset with a dynamic light perception random mask. The improved multi-modal detection model includes a two-stream backbone network, a cross-modal differential feature interaction enhancement module, and a YOLO detection head. S3. Use the improved multi-modal detection model to process the initial image to obtain the object detection result.

[0006] Preferably, the two-stream backbone network is used to perform multi-scale feature extraction on the initial image for subsequent feature fusion. The initial image includes a visible light image and an infrared image, and the image features include visible light features and infrared features .

[0007] Preferably, the cross-modal differential feature interaction enhancement module is used to utilize the differential complementary information between the initial images, enhance the features between modalities in both the channel and spatial aspects, fuse the multi-scale features using channel concatenation means, send them into a feature pyramid structure, and finally input them to the YOLO detection head. The YOLO detection head detects the information to be detected to obtain the object detection result, and the object detection result includes the object position, object category, and prediction probability.

[0008] Preferably, the differential feature interaction enhancement module includes two parts: channel attention and spatial attention, and the channel attention and spatial attention are placed in series.

[0009] Preferably, the enhancement part of the channel attention is processed as follows: S2101. Differential feature enhancement. The visible light features and infrared features are used as inputs together. After subtracting the two, the differential features are obtained. S2102. After passing through a 1×1 convolution and a GLUE activation function, an enhanced differential feature map is obtained ; S2103. The spatial information of the differential feature map is obtained through average pooling and max pooling operations to obtain a global receptive field for spatial aggregation. S2104. The spatial information is then processed through a shared MLP and a channel attention map is generated , and the channel attention enhanced feature is obtained through element-wise multiplication of the channel attention map and the input differential feature . S2105. Introduce the idea of residual addition, and add the channel attention enhanced feature to the visible light features respectively and infrared features Add them together to obtain and ; S2106. Add and together to obtain the output result of the channel attention part .

[0010] Preferably, the enhanced expression of the channel attention is: ; wherein, ; ; ; ; wherein, represents element-wise multiplication; represents the sigmoid function; represents a 1×1 convolution operation; and represent global average pooling and global max pooling; MLP represents a shared multi-layer perceptron.

[0011] Preferably, after calculating the channel attention map, use a depthwise separable convolution module to extract important spatial regions for each channel, and finally obtain a dynamically allocated spatial attention map; The enhanced part of the channel attention adopts the Inception style.

[0012] Preferably, the processing steps of the enhanced part of the spatial attention are as follows: S2107. Use a 5×5 small kernel convolution to obtain the local feature information of the output result of the channel attention part. The local feature is represented by ; S2108. Secondly, use multiple groups of parallel depthwise strip convolutions to capture local context information across multiple scales; S2109. Obtain the final output feature through the element-wise multiplication output of the spatial attention mechanism channel mixing result and the channel prior feature; Among them, using 1×1 convolution to achieve channel mixing helps to obtain a more refined spatial attention map.

[0013] Preferably, the enhanced expression of the spatial attention is: ; ; wherein, DwConv represents depthwise separable convolution; represents the i-th branch, is the original input, and the rest are all composed of strip convolutions; represents a 1×1 convolution operation.

[0014] The multi-modal object detection dataset includes visible-infrared images. After dividing the multi-modal object detection dataset into a training set, a validation set, and a test set; The specific steps of the random mask training strategy with dynamic light perception are as follows: Step 1: Evenly divide a visible light image into N×N regions, calculate the average illumination value of each region, and arrange these regions in ascending order of illumination value. The K regions with the lowest brightness, that is, the regions with the lowest ranking, are dark regions (referred to as Top-K regions), and the K regions with the highest ranking are bright regions (referred to as Bottom-K regions); Step 2: Also divide the infrared image matching the visible light image into N×N regions; Step 3: For the dark regions, add a mask at the positions of all dark regions in the visible light image, that is, set the pixel value to 0; Step 4: For the bright regions, randomly select the positions of K / 2 bright regions to add a mask in the infrared modality, that is, set the pixel value to 0, and add a mask at the positions of the remaining K / 2 in the visible light modality, that is, set the pixel value to 0; Step 5: Select a target proportion of visible-infrared images in the training set of the multi-modal dataset, and repeat the above steps 1-4, that is, perform mask processing; Step 6: The mask processing is only performed on the training set, and the validation set and the test set remain unchanged. Use the training set after mask processing and the original validation set and test set to train and test the multi-modal object detection model.

[0015] In summary, a multi-modal object detection method based on dynamic light perception random mask according to the present invention, compared with the traditional technology, the present invention proposes a dynamic light perception random mask method, performs mask processing on a partial proportion of images in the training set, and considers the influence of exposure and strong light environments, enabling the network to learn multi-modal effective complementary features without bias; constructs a dual-modal feature-level fusion object detection basic network; introduces a cross-modal differential feature interaction module, improves the multi-modal remote sensing object detection ability of the drone platform / satellite platform, and improves the object recognition accuracy under complex lighting conditions.

[0016] Next, through the drawings and embodiments, the technical method of the present invention will be further described in detail. Description of the Drawings

[0017] Figure 1 is a schematic diagram of the principle process of the present invention; Figure 2The improved multi-modal detection model diagram of the present invention; Figure 3 The schematic diagram of the cross-modal differential feature interaction module of the present invention; Figure 4 The comparison diagram of the object detection results of different models under the VEDAI dataset of the present invention. Detailed implementation manners

[0018] The technical method of the present invention will be further described below through the drawings and embodiments. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions and values of the components and steps described in these embodiments do not limit the scope of the present application.

[0019] The following description of at least one exemplary embodiment is actually only illustrative and in no way restricts the present application and its application or use.

[0020] The technologies, systems and devices known to those of ordinary skill in the relevant fields may not be discussed in detail, but where appropriate, the technologies, systems and devices should be regarded as part of the specification.

[0021] In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0022] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the field to which the present invention belongs.

[0023] Based on the YOLOv8n model, the present invention designs a two-stream feature extraction backbone network model to extract depth features from visible light and infrared images respectively. Secondly, the present invention designs a CDFIM differential feature interaction enhancement module, which uses the differential complementary information between multi-modal images to enhance the features between modalities in both the channel and spatial aspects. And the multi-scale features are fused by means of channel concatenation, sent into the feature pyramid structure, and finally input to the YOLO detection head to obtain the fusion recognition result. For the training of the model, the present invention proposes a random masking strategy based on light perception. Based on the light prior information of visible light images, while balancing the attention of the model to the two modalities, it improves the learning of effective modal information, and takes into account the influence of exposure and strong light, so that the network can unbiasedly focus on the complementary and effective information of different modalities.

[0024] A multi-modal object detection method based on dynamic light perception random masking provided by the present invention, as Figure 1 and Figure 2 shown, S1, obtain the image to be detected, and preprocess the image to be detected to obtain the initial image.

[0025] S2. Based on YOLOv8n and an improved cross-modal differential feature interaction module, an improved multi-modal detection model is constructed, a multi-modal object detection dataset is obtained, and the improved multi-modal detection model is trained using a dual-modal model training strategy that combines the multi-modal object detection dataset with a dynamic light perception random mask.

[0026] The improved multi-modal detection model includes a two-stream backbone network, a cross-modal differential feature interaction enhancement module, and a YOLO detection head.

[0027] S3. The initial image is processed using the improved multi-modal detection model to obtain the object detection result.

[0028] Furthermore, the two-stream backbone network is used to perform multi-scale feature extraction on the initial image for subsequent feature fusion. The initial image includes visible light images and infrared images, and the image features include visible light features and infrared features .

[0029] Furthermore, the cross-modal differential feature interaction enhancement module is used to utilize the differential complementary information between the initial images, enhance the features between modalities in both the channel and spatial dimensions, fuse the multi-scale features using channel concatenation, send them into the feature pyramid structure, and finally input them to the YOLO detection head.

[0030] The YOLO detection head detects the information to be detected to obtain the object detection result, and the object detection result includes the object position, object category, and prediction probability.

[0031] The present invention designs the CDFIM module as a bridge between the visible light and infrared branches. As shown in the figure. For the dual-modal information input, traditional processing methods usually use direct addition or channel concatenation, which will increase both differential information and redundant information. Therefore, the CDFIM designed by the present invention focuses on differential features to achieve the enhancement of cross-modal complementary features. The CDFIM module contains two parts: channel attention and spatial attention, which are placed in series. A design of dynamically distributing attention weights in both the channel and spatial dimensions is adopted to more accurately focus on important target features.

[0032] Furthermore, as Figure 3 shown, the differential feature interaction enhancement module contains two parts: channel attention and spatial attention, and the channel attention and spatial attention are placed in series.

[0033] Furthermore, the enhancement part of the channel attention is processed as follows: S2101. Differential feature enhancement, visible light features and infrared features Take them as inputs together, subtract them from each other, and obtain differential features.

[0034] S2102. After 1×1 convolution and GLUE activation function, obtain an enhanced differential feature map .

[0035] S2103. Obtain the global receptive field through average pooling and max pooling operations on the spatial information of the differential feature map, and perform spatial aggregation.

[0036] S2104. Subsequently, process the spatial information through a shared MLP and generate a channel attention map , and obtain the channel attention enhanced feature through element-wise multiplication of the channel attention map and the input differential feature .

[0037] S2105. Introduce the idea of residual addition, add the channel attention enhanced feature to the visible light feature and the infrared feature respectively, and obtain and .

[0038] S2106. Add and to obtain the output result of the channel attention part .

[0039] Furthermore, the enhanced expression of the channel attention is: .

[0040] Among them, . . . .

[0041] Among them, represents element-wise multiplication. represents the sigmoid function. represents the 1×1 convolution operation. and represent global average pooling and global max pooling. MLP represents the shared multi-layer perceptron.

[0042] Furthermore, after calculating the channel attention map, use the depthwise separable convolution module to extract important spatial regions for each channel, and finally obtain the dynamically allocated spatial attention map.

[0043] The enhanced part of the channel attention adopts the Inception style.

[0044] Furthermore, the processing steps of the enhanced part of the spatial attention are as follows: S2107. Use a 5×5 small kernel convolution to obtain the output result of the channel attention part of the local feature information, and the local feature is represented by .

[0045] S2108. Secondly, use multiple groups of parallel depth strip convolutions to capture local context information across multiple scales.

[0046] One group of strip convolutions can be equivalent to the receptive field of a standard convolution, and the computational cost of depth strip convolution is smaller compared to 2D convolution. In addition, strip convolution is convenient for extracting the features of slender-shaped objects such as vehicles, ships, airplanes, etc., and is more suitable for remote sensing targets.

[0047] S2109. Obtain the final output feature through the element-wise multiplication output of the channel mixing result and the channel prior feature by the spatial attention mechanism .

[0048] Among them, using 1×1 convolution to achieve channel mixing helps to obtain a more refined spatial attention map.

[0049] Furthermore, the expression of the enhanced part of the spatial attention is: .

[0050] .

[0051] Among them, DwConv represents depthwise separable convolution. represents the i-th branch, is the original input, and the rest are all composed of strip convolution pairs. represents a 1×1 convolution operation.

[0052] When using an intelligent model to learn bimodal features, due to the rich texture and color information in the visible light modality, the backbone network usually tends to extract visible light modality features and ignores infrared modality features. The idea of an asymmetric mask can force the network to learn more effective features from complementary modes and avoid model bias learning. However, the typical random mask strategy does not consider the importance of different modalities in different regions of the image. For this reason, the present invention proposes a dynamic perception random mask training strategy. Different from the conventional random mask method, the strategy of the present invention adds a dynamic mask to each group of bimodal images based on the local illumination perception information of the visible light image. In addition, the present invention considers the modality feature learning method in the case of strong light and exposure, and improves the learning ability of the model for complex environments.

[0053] The multimodal object detection dataset includes visible light-infrared images. After dividing the multimodal object detection dataset into a training set, a validation set, and a test set.

[0054] The specific steps of the random mask training strategy using dynamic light perception are as follows: Step 1: Evenly divide a visible light image into N×N regions, calculate the average illumination value of each region, and arrange these regions in ascending order of illumination value. The K regions with the lowest brightness, that is, the regions with the lowest illumination value, are the dark light regions (referred to as the Top-K regions), and the K regions with the highest brightness are the bright light regions (referred to as the Bottom-K regions).

[0055] Step 2: Also divide the infrared image matching the visible light image into N×N regions.

[0056] Step 3: For the dark light regions, add a mask at the positions of all dark light regions in the visible light image, that is, set the pixel value to 0.

[0057] Step 4: For the bright light regions, randomly select the positions of K / 2 bright light regions to add a mask in the infrared modality, that is, set the pixel value to 0, and add a mask at the positions of the remaining K / 2 in the visible light modality, that is, set the pixel value to 0.

[0058] Step 5: Select a target proportion of visible light-infrared images in the training set of the multi-modal dataset, and repeat the above steps 1-4, that is, perform mask processing.

[0059] Finally: 。

[0060] 。

[0061] Step 6: The mask processing is only performed on the training set, and the validation set and test set remain unchanged. Use the training set after mask processing and the original validation set and test set to train and test the multi-modal object detection model.

[0062] To verify the effectiveness of the multi-modal detection model and the dynamic light perception random mask training strategy designed in the present invention, the present invention conducts object detection tests under the VEDAI dataset. VEDAI is a public dataset for small object detection in aerial images, which contains challenges such as illumination / shadow changes, strong light reflection, and occlusion, covering various scenarios such as rural, urban areas, and mountainous areas. The dataset contains 1200 pairs of RGB-IR images, more than 3700 annotated objects, with an average of 5.5 objects per image, which account for about 0.7% of the total image pixels. Nine types of ground targets are provided. In this article, the 1024×1024 pixel version is used for experiments in the present invention, and 1089 pairs of image pairs are selected as the training set, and 121 pairs are used as the validation set and test set.

[0063] During the experiment of the present invention, the Windows 11 operating system was adopted, and the PyCharm integrated development environment (IDE) was used for programming. This development environment was configured with Python 3.8 version, PyTorch 1.8 deep learning framework, and CUDA 11.1 version. The hardware resources relied on for the experiment included a CPU of the Intel Core i7-13700KF model and a GPU of the NVIDIA RTX 4090 model. During the model training stage, a total of 300 epochs were executed, and the batch size for each epoch was 16 samples.

[0064] The ablation experiments of each module on the VEDAI dataset are shown in Table 1. The basic model of the present invention adopts the method of Concat fusion in the middle layer of the bimodal YOLOv8n. It can be found that on the VEDAI dataset, when the CDFIM module is added to the basic model, the mAP50 of the model is increased by 1.5% respectively, indicating that the complementary learning ability and feature interaction ability of the bimodal features can be effectively improved by adding the CDFIM module of the present invention. In addition, the number of parameters only increases by 0.26M, and the computational volume only increases by 0.7 GFLOPs. This is attributed to the cross-modal differential features extracted in the CDFIM module of the present invention, which avoids the learning of redundant information and introduces depthwise strip convolution to achieve the purpose of feature enhancement with extremely low computational volume. Combined with the dynamic light perception random mask training strategy, the recognition accuracy on this dataset is improved to 79.5%, and no additional computational burden is added to the model. The above results prove the effectiveness of the cross-modal differential feature extraction module and the dynamic perception random mask training strategy proposed by the present invention.

[0065] Table 1 Ablation experiments of each module on the VEDAI dataset ;

[0066] As shown in Table 2, the model of the present invention is 7.4% higher than the multi-modal baseline model of the present invention in terms of mAP50, which indicates the advancement and effectiveness of the model proposed by the present invention. Compared with the YOLOv8n visible light single-mode model, the recognition accuracy mAP50 of the present invention is increased by 8.5%, and compared with the YOLOv8n infrared single-mode model, the recognition accuracy mAP50 of the present invention is increased by 8.1%. Compared with the MMYFnet model with the highest recognition accuracy among multi-modal models. Although the model of the present invention is slightly lower by 0.5% in terms of mAP, the number of parameters of the present invention is reduced by 72.8%, and the computational volume is reduced by 71.6%. The lightweight model will be more conducive to the actual deployment and rapid inference on the spaceborne and airborne platforms. Therefore, the multi-modal object detection method proposed by the present invention shows a good accuracy-computation trade-off.

[0067] Table 2 Comparison results of the recognition performance of different object detection models under the VEDAI dataset ;

[0068] To more intuitively display the detection performance of this method, the detection results of the model of the present invention are compared with those of two other models. As Figure 4 shown, for the truck target in the middle area, due to the small size of the target and its very close color to the background, redundant positioning occurs in the YOLOv8n visible light single-mode model. Due to the blurred infrared target contour, classification errors and missed detections occur in the YOLOv8n infrared single mode. The benchmark model also has classification errors. The multi-modal detection model of the present invention correctly detects and locates all truck targets. This proves the detection robustness of the model of the present invention for targets in complex scenarios.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical method of the present invention, and these modifications or equivalent replacements cannot make the modified technical method deviate from the spirit and scope of the technical method of the present invention.

Claims

1. A multi-modal object detection method based on dynamic light perception random mask, characterized in that, It includes the following steps: S1. Obtain the image to be detected, and preprocess the image to be detected to obtain an initial image; S2. Based on YOLOv8n and an improved cross-modal differential feature interaction module, construct an improved multi-modal detection model, obtain a multi-modal object detection dataset, and train the improved multi-modal detection model using a dual-modal model training strategy that combines the multi-modal object detection dataset with a dynamic illumination-aware random mask; The improved multi-modal detection model includes a two-stream backbone network, a cross-modal differential feature interaction enhancement module, and a YOLO detection head; S3. Use the improved multi-modal detection model to process the initial image to obtain an object detection result.

2. The multimodal object detection method based on dynamic light perception random mask according to claim 1, wherein, The two-stream backbone network is used to perform multi-scale feature extraction on the initial image for subsequent feature fusion; The initial image includes a visible light image and an infrared image, and the image features include visible light features and infrared features .

3. The multimodal object detection method based on dynamic light perception random mask according to claim 2, wherein, The cross-modal differential feature interaction module is used to utilize the differential complementary information between the initial images, enhance the features between modalities in terms of both channels and space, fuse the multi-scale features using channel concatenation means, send them into a feature pyramid structure, and finally input them to the YOLO detection head; The YOLO detection head detects the information to be detected to obtain an object detection result, and the object detection result includes the object position, object category, and prediction probability.

4. The multimodal object detection method based on dynamic light perception random mask according to claim 3, wherein, The differential feature interaction enhancement module includes two parts: channel attention and spatial attention, and the channel attention and spatial attention are placed in series.

5. A multimodal object detection method based on dynamic light perception random masking according to claim 4, characterized in that, The enhancement processing steps of the channel attention are as follows: S2101. Differential feature enhancement, visible light feature and infrared feature are jointly used as inputs, and the differential feature is obtained by subtracting the two. S2102. Through 1×1 convolution and the GLUE activation function, an enhanced differential feature map is obtained. ; S2103. The spatial information of the differential feature map obtains a global receptive field through average pooling and max pooling operations for spatial aggregation; S2104. The spatial information is then processed by the shared MLP to generate a channel attention map. , and the channel attention enhanced features are obtained by element-wise multiplication of the channel attention map and the input differential features. ​ S2105. Introduce the idea of residual addition, and add the channel attention enhanced features to the visible light features and the infrared features respectively, to obtain and ; S2106. Add and to obtain the output result of the channel attention part .

6. The multimodal object detection method based on dynamic light perception random mask according to claim 5, characterized in that, The enhancement expression of the channel attention is: ; Among them, ; ; ; ; Among them, represents element-wise multiplication; represents the sigmoid function; represents a 1×1 convolution operation; and represent global average pooling and global max pooling; MLP represents a shared multi-layer perceptron.

7. A multi-modal object detection method based on dynamic light perception random mask according to claim 6, characterized in that, After calculating the channel attention map, use a depthwise separable convolution module to extract important spatial regions for each channel, and finally obtain a dynamically allocated spatial attention map; The enhancement part of the channel attention adopts the Inception style.

8. A multimodal object detection method based on dynamic light perception random masking according to claim 4, characterized in that The enhancement processing steps of the spatial attention are as follows: S2107. Use a 5×5 small kernel convolution to obtain the output result of the channel attention part of the local feature information, and the local feature is represented by ; S2108. Secondly, use multiple groups of parallel depthwise strip convolutions to capture local context information across multiple scales; S2109. Obtain the final output feature by performing element-wise multiplication on the channel mixing result and the element-level multiplication output of the channel prior feature through the spatial attention mechanism ; Among them, using 1×1 convolution to achieve channel mixing helps to obtain a more refined spatial attention map.

9. A multimodal object detection method based on dynamic light perception random mask according to claim 8, characterized in that, The enhancement expression of the spatial attention is: ; ; Among them, DwConv represents depthwise separable convolution; represents the i-th branch, is the original input, and the rest are all composed of strip convolution pairs; represents a 1×1 convolution operation.

10. A multi-modal object detection method based on dynamic light perception random mask according to claim 1, characterized in that The multi-modal object detection dataset includes visible light-infrared images. After dividing the multi-modal object detection dataset into a training set, a validation set, and a test set; The specific steps of the dynamic illumination-aware random mask training strategy are as follows: Step 1: Evenly divide a visible light image into N×N regions, calculate the average illumination value of each region, and arrange these regions in ascending order of illumination value. The K regions with the lowest brightness, that is, the regions with the highest ranking, are called dark regions, namely Top-K regions, and the K regions with the highest brightness, that is, the regions with the lowest ranking, are called bright regions, namely Bottom-K regions; Step 2: Also divide the infrared image matching the visible light image into N×N regions; Step 3: For the dark regions, add a mask at the positions of all dark regions in the visible light image, that is, set the pixel values to 0; Step 4: For the bright regions, randomly select the positions of K / 2 bright regions to add masks in the infrared modality, that is, set the pixel values to 0, and add masks to the remaining K / 2 positions in the visible light modality, that is, set the pixel values to 0; Step 5: Select the target proportion of visible light-infrared images in the training set of the multi-modal dataset, and repeat the above Steps 1-4, that is, perform mask processing; Step 6: The mask processing is only performed on the training set, and the validation set and the test set remain unchanged. Use the training set after mask processing and the original validation set and test set to train and test the multi-modal object detection model.

Citation Information

Patent Citations

  • Organic dynamic random access memory based on PEDOT (Poly 3, 4-ethylenedioxythiophene):PSS (poly styrenesulfonate) and manufacturing method thereof

    CN102214791A

  • Semantic segmentation method, system and equipment for aerial image and medium

    CN119027834A

  • Marine work equipment identification method based on vision

    CN119863700A

Cited By

  • Image fusion method based on cross-modal difference and biaxial attention network

    CN120598800A

  • Channel buoy detection method based on fusion of multi-mode pulse neural network and visual Transform

    CN121616952A