A small target detection method and system based on visible light and infrared multimodal fusion
By adopting a pixel-level processing method of multimodal fusion of visible light and infrared in remote sensing small object detection, and using the improved YOLOv5 neural network model, the problems of low detection rate, high false alarm rate and low real-time in the prior art are solved, and more efficient remote sensing small object recognition is achieved.
Patent Information
- Application Number
- CN202410601737.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-05-15
AI Technical Summary
The prior art has problems such as low detection rate, high false alarm rate, high false alarm rate and low real-time in remote sensing small object detection, mainly due to insufficient network feature information extraction and feature fusion capabilities, and low recognition accuracy in extreme environments.
A small object detection method based on multimodal fusion of visible light and infrared is proposed. The visible light and infrared images collected by satellites or drones are processed through pixel-level fusion, and the improved YOLOv5 neural network model, including improved backbone network, feature expression network and head network structure, improve the feature extraction and fusion capabilities of the detection model.
It effectively improves the recognition accuracy of remote sensing small targets, reduces false alarms and false alarm rates, enhances the real-time nature of the detection system, and is suitable for environments where multimodal data is complementary information and compute resources are limited.
Smart Images

Figure CN118470557B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual detection, and specifically relates to a small target detection method and system based on visible light and infrared multimodal fusion. Background Art
[0002] In recent years, target detection technology applied to satellites and drones has had an important impact on the civil and military fields. Since remote sensing images have the characteristics of small and dense targets, complex backgrounds and more interference, the target detection and recognition accuracy is low, and the false detection rate and false alarm rate are not ideal. Faced with many technical challenges, traditional detection methods are not enough to effectively detect small remote sensing targets. The main reason is that the network's feature information extraction and feature fusion capabilities are insufficient and the feature expression capabilities are limited. Improving the accuracy of remote sensing small target recognition has become an important research field.
[0003] Since the visible light imaging system relies on external light illumination for imaging, the visible light image data usually contains rich color and texture information of the target. It is mainly suitable for scenes with good lighting conditions, but the recognition accuracy is low in complex environments such as local strong light or backlight environments, rain and snow extreme weather, and smoke. The infrared imaging system uses the thermal radiation characteristics of the target for passive detection and is not easily interfered with. It has good detection effects at night and in low light environments, but the shape and texture information of infrared images is very limited. It is difficult to achieve ideal detection effects for small target recognition through a single modality. With the rapid development of imaging technology, it has become a feasible path to collect remote sensing images using multispectral cameras. The information complementarity between multiple modal data can improve the detection accuracy. Due to the diverse characteristics of multimodal data, it is difficult to achieve ideal results by directly applying the general single-modal detection algorithm to the infrared and visible light multimodal fusion small target detection scene.
[0004] The main difficulties of multimodal fusion detection are scale diversity, differential feature fusion, alignment and registration of different modal data, feature information loss, noise interference, computing resource limitations, and scarce training data. The cause of scale diversity is that small targets may show different scale characteristics in different modal data. Differential feature fusion means that different modal data may have differences in the representation of small target color, texture and other features. Effective fusion and extraction of useful information is the key. In the process of multimodal data acquisition, due to the differences in acquisition perspective and time, the spatial alignment and registration of small targets may be inaccurate, resulting in unsatisfactory data fusion capabilities. The main cause of feature information loss and noise interference is that small targets occupy fewer pixels and are easily lost during feature extraction. Multimodal algorithms need to process and fuse data from different sensors, which increases the complexity and computational cost of the algorithm. Since remote sensing small target detection models are mainly deployed on satellites and drones for real-time processing, they are also limited by computing resources. In addition, small target data annotation in multimodal scenarios requires a lot of money and time costs, and currently available multimodal small target data sets are relatively scarce.
[0005] According to the level of feature fusion, multimodal algorithms can be divided into pixel-level fusion, feature-level fusion and decision-level fusion. Among them, pixel-level fusion refers to fusion on raw data. By combining data information from different sensors in the data preprocessing stage. The specific process of pixel-level fusion is as follows: Figure 1 As shown in the figure, pixel-level fusion refers to fusion on raw data. By combining data information from different sensors in the data preprocessing stage, it is more suitable for the fusion of satellite images of different bands in the remote sensing field to obtain richer feature information. The specific process of feature-level fusion is as follows: Figure 2 As shown in , feature-level fusion refers to the fusion at the feature representation level where data is converted into feature descriptors. The specific process of decision-level fusion is as follows: Figure 3 As shown in the figure, decision-level fusion refers to the fusion of the results of independent decisions of different modalities, that is, each sensor independently identifies and classifies the target. Compared with pixel-level fusion, feature-level fusion is the fusion of high-level abstract features, which may cause feature loss and thus missed detection. Decision-level fusion has the problem of repeated calculation of different multimodal branches, which consumes a lot of computing resources. Therefore, pixel-level fusion is closer to the data source and more suitable for satellite or airborne applications.
[0006] It is a meaningful task to study small target detection based on multimodal fusion. Summary of the invention
[0007] In view of the problems of low detection rate, high false alarm rate, high false alarm rate and low real-time performance existing in the above-mentioned technology, the purpose of the present invention is to overcome the above-mentioned defects of the prior art and propose a small target detection method and system based on visible light and infrared multimodal fusion.
[0008] To achieve the above object, the present invention proposes a small target detection method based on visible light and infrared multimodal fusion, including:
[0009] Performing pixel-level fusion processing on visible light images and infrared images collected by satellites or drones;
[0010] Inputting the fused image into a trained detection model to output a detection result;
[0011] The detection model is an improved YOLOv5 neural network, including: an improved backbone network structure, an improved feature expression network, and an improved Head network structure; wherein,
[0012] The improved backbone network structure is used to capture feature maps of different scales by introducing the BottleneckCSP structure and adopting an improved SPP structure, and input them into the improved feature expression network;
[0013] The improved feature expression network; is used to obtain a feature layer by introducing the Leaky_Relu activation function into the convolutional layer part and adjusting the downsampling ratio of the average pooling layer, and input it into the improved Head network structure;
[0014] The improved Head network structure is used to perform target classification and regression by introducing the BottleneckCSP2 structure and output a detection result.
[0015] Preferably, the improved backbone network structure includes: an MF module, an improved convolutional module, a BottleneckCSP structure, 3 improved convolutional modules, and an improved SPP structure connected in sequence;
[0016] The MF module is used to extract shared and special information from different modalities, and can bidirectionally combine multimodal internal information in a symmetric and compact manner;
[0017] The improved BottleneckCSP structure is used to further improve the detection performance of the detection model;
[0018] The improved convolutional module is used to capture subtle feature changes;
[0019] The improved SPP structure is used to capture feature maps of different scales.
[0020] Preferably, the improvement of the convolutional module includes:
[0021] Replacing the SiLu activation function of the convolutional module with the Leaky_Relu function, and the improved convolutional module includes Conv2d, BatchNorm2d, and Leaky_Relu.
[0022] Preferably, the improvement of the SPP structure includes:
[0023] Introduce dilated convolution and max pooling of different scales to capture multi-scale features, and use grouped convolution in dilated convolution with the number of groups equal to the number of input channels to reduce the computational amount;
[0024] Use nearest neighbor interpolation to upsample the pooled feature map so that feature maps of different scales have the same spatial size;
[0025] After each convolution operation, use the Leaky_ReLU activation function to introduce non-linearity.
[0026] Preferably, the processing of the improved SPP structure includes:
[0027] The input feature map x passes through the convolutional layer Conv1, the batch normalization layer BN1, and the Leaky_ReLU activation function, and the results output by Conv1 and BN1 then pass through the dilated convolutional layer atrous_conv, the batch normalization layer BN_atrous, and the Leaky_ReLU activation function;
[0028] The max pooling layer Maxpool pools the input x to capture features of different scales and forms Pooled_features;
[0029] Upsample each pooled feature map in Pooled_features by nearest neighbor interpolation to match the spatial size of the original input feature map and form Pooled_features_upsampled;
[0030] Concatenate the feature map x, the output x_atrous of the dilated convolution, and the upsampled pooled feature map Pooled_features_upsampled in the channel dimension to form features;
[0031] Pass the concatenated feature map features through the one-dimensional convolutional layer Conv2 and the batch normalization layer BN2, and apply the Leaky_ReLU activation function to obtain the output x_concatenated.
[0032] Preferably, in the improved feature expression network, adjusting the downsampling ratio of the average pooling layer is achieved by adjusting subscale.
[0033] Preferably, the improved Head network structure includes 4 branches with the same structure and 1 classification and regression network. Each branch generates a set of detection results, and the detection results include object categories and positions. Each branch includes: an improved convolutional module, upsampling, splicing, an improved convolutional module, upsampling, splicing, and a BottleneckCSP2 structure connected in sequence;
[0034] The improved convolutional module is used to further process the feature map obtained by processing the improved backbone network structure; it includes Conv2d, BatchNorm2d, and Leaky_Relu;
[0035] The BottleneckCSP2 structure is used to fuse feature maps at different levels.
[0036] Preferably, the classification and regression network is used to fuse the detection results of the 4 branches to obtain the final detection results.
[0037] Preferably, it includes the training steps of the detection model, including:
[0038] Input the training set into the improved YOLOv5 neural network, adjust the parameters until the training requirements are met, and obtain the trained detection model.
[0039] On the other hand, the present invention provides a small target detection system based on visible light and infrared multimodal fusion, including:
[0040] A fusion module for performing pixel-level fusion processing on visible light images and infrared images collected by satellites or drones; and
[0041] A detection module for inputting the fused image into the trained detection model and outputting the detection results;
[0042] The detection model is an improved YOLOv5 neural network, including: an improved backbone network structure, an improved feature expression network, and an improved Head network structure; wherein,
[0043] The improved backbone network structure is used to capture feature maps of different scales by introducing the BottleneckCSP structure and adopting the improved SPP structure, and input them into the improved feature expression network;
[0044] The improved feature expression network is used to obtain the feature layer by introducing the Leaky_Relu activation function into the convolutional layer part and adjusting the downsampling ratio of the average pooling layer, and input it into the improved Head network structure;
[0045] The improved Head network structure is used to perform object classification and regression and output detection results by introducing the BottleneckCSP2 structure.
[0046] Compared with the prior art, the advantages of the present invention are as follows:
[0047] 1. A pixel-level fusion algorithm is introduced to address the problem of limited computing resources in multi-modal fusion small object detection.
[0048] 2. Aiming at the problem that the proportion of remote sensing image objects in the image is relatively small, resulting in relatively less feature information and prone to feature information loss. The BottleneckCSP structure is introduced into the Backbone part, and the BottleneckCSP2 structure is introduced into the Head module to enhance the feature extraction ability of the network.
[0049] 3. In view of the characteristics of scale diversification and differential feature fusion in multi-modal fusion small object detection, the SPP feature fusion structure is improved to effectively fuse feature information.
[0050] 4. To address the problem that small objects have a low signal-to-noise ratio and are vulnerable to complex backgrounds and noise interference, the Leaky_Relu activation function is introduced into the convolutional layer part to improve the model performance.
[0051] 5. According to the characteristics of the dataset, the downsampling ratio of the appropriate average pooling layer is adjusted to reduce the computational amount while retaining the image details. Description of the Drawings
[0052] Figure 1 It is a pixel-level fusion framework diagram;
[0053] Figure 2 It is a feature-level fusion framework diagram;
[0054] Figure 3 It is a decision-level fusion framework diagram;
[0055] Figure 4 It is an unimproved Backbone block diagram;
[0056] Figure 5 The block diagram of the small object real-time recognition model based on visible light and infrared multi-modal fusion of the present invention;
[0057] Figure 6 It is a comparison diagram of the recognition results of a small object real-time recognition method based on visible light and infrared multi-modal fusion of the present invention and other methods. Detailed Embodiments
[0058] The present invention proposes an improved YOLOv5 neural network model. The improvements include: improving the backbone network structure, improving the feature expression network, and improving the Head network structure; among which,
[0059] The improved backbone network is used to distinguish small targets from complex backgrounds, and the feature extraction ability of the network directly affects the accuracy of the recognition system. It includes: an MF module, an improved convolutional module, a BottleneckCSP structure, 3 improved convolutional modules, and an improved SPP structure connected in sequence. The collected original visible light images and infrared images are subjected to image preprocessing and pixel-level fusion technology processing and then sent into the improved backbone feature extraction network for processing. The improvements to the backbone network structure include: ① introducing an improved SPP feature fusion module, ② introducing an improved convolutional module, ③ introducing an MF module.
[0060] The main function of the improved SPP spatial pyramid pooling module is to capture features of different scales to improve the model's recognition ability for objects of different sizes. Based on the original SPP structure, our improvements are mainly: ① introducing dilated convolution, ② changing to the Leaky_ReLU activation function, ③ introducing upsampling. Among them, the role of introducing dilated convolution and max pooling of different scales is to capture multi-scale features. The role of introducing upsampling and channel concatenation is to fuse features of different scales together to increase the network's recognition ability for objects of different scales. To facilitate subsequent concatenation operations, nearest neighbor interpolation is used to upsample the pooled feature map to ensure that feature maps of different scales have the same spatial size. To better express features, the Leaky_ReLU activation function is used after each convolution operation to introduce non-linearity. Introducing the above modules will increase the computational complexity, so grouped convolution is used in dilated convolution, where the number of groups is equal to the number of input channels c_, which helps to reduce the computational complexity.
[0061] The specific implementation of the improved SPP spatial pyramid pooling module is as follows:
[0062] The input feature map x passes through the convolutional layer Conv1, the batch normalization layer BN1, and the Leaky_ReLU activation function, and the results output by Conv1 and BN1 pass through the dilated convolutional layer atrous_conv, the batch normalization layer BN_atrous, and the Leaky_ReLU activation function;
[0063] The max pooling layer Maxpool pools the input x to capture features of different scales and forms Pooled_features;
[0064] Upsample each pooled feature map in Pooled_features through nearest neighbor interpolation to match the spatial dimensions of the original input feature map, forming Pooled_features_upsampled;
[0065] Concatenate the feature map x, the output x_atrous of the dilated convolution, and the upsampled pooled feature map Pooled_features_upsampled along the channel dimension to form features;
[0066] Pass the concatenated feature map features through a one-dimensional convolutional layer Conv2 and a batch normalization layer BN2, and apply the Leaky_ReLU activation function to obtain the output x_concatenated.
[0067] The introduced improved convolutional module replaces the SiLu activation function of the convolutional module with the Leaky_Relu function. The improved convolutional module consists of Conv2d, BatchNorm2d, and Leaky_Relu. In small object detection, the object size is small and the feature information is scarce. Introducing Leaky_Relu can effectively improve the model performance to capture subtle feature changes more sensitively.
[0068] The MF module is introduced in the Backbone part to fuse the features of the visible light modality and the infrared modality, and at the same time alleviate the problem of limited computing resources in multi-modal fusion small object detection. The pixel-level MF module can extract shared and special information from different modalities and can bidirectionally combine the internal information of multiple modalities in a symmetric and compact manner.
[0069] As an improvement of the above method, the improved backbone network enhances the feature extraction ability of the network by introducing the BottleneckCSP structure. The SPP structure is improved to enhance the ability of network feature fusion and transmission. The improved convolutional module is introduced to capture features more sensitively. The Backbone network structure before improvement is as Figure 4 shown. The improved network structure is as Figure 5 shown.
[0070] Improve the Head network structure to fuse feature maps at different levels, so as to retain more feature information with different receptive fields. The feature maps obtained by the backbone network are processed to obtain feature layers and sent to the Head network (including the neck network). The improvements of the Head network structure include: ① introducing the BottleneckCSP2 module, ② introducing the improved convolutional module; where
[0071] The BottleneckCSP2 module is used to fuse feature maps at different levels, thereby retaining more feature information with different receptive fields. The bottleneck structure in the BottleneckCSP2 module adopts residual connections and dimensionality reduction operations, which helps reduce the computational complexity of the model and improve the operation efficiency. The CSP layer can enhance the adaptability and non-linear representation ability of the network, enabling the model to better adapt to small object detection tasks under different scenarios and lighting conditions and improving the robustness of detection. By fusing feature maps at different levels, the BottleneckCSP2 module enables the model to learn richer context information and object structure information, thus more accurately identifying small objects.
[0072] The improved convolutional module is used to further process the feature maps obtained by the Bankbone network. The improved convolutional module consists of Conv2d, BatchNorm2d, and Leaky_Relu.
[0073] The improved feature expression network is used to enhance the expression ability of the model through non-linear transformation. Compared with linear transformation, non-linear transformation helps the model capture complex features of the input data. By introducing the Leaky_Relu activation function into the convolutional layer part, the model performance is improved. According to the characteristics of the dataset, the appropriate downsampling ratio of the average pooling layer is adjusted to reduce the computational complexity while retaining image details.
[0074] The Leaky_Relu activation function is an improved ReLU (Rectified Linear Unit) activation function. It allows negative inputs to pass through a non-zero gradient, thus solving the problem of gradient vanishing in the negative input region of traditional ReLU. It is used to capture features from rough to detailed at different levels, thus effectively handling objects of various sizes. This is particularly important for small object detection because subtle feature changes may have a significant impact on the detection results.
[0075] The improvement of the downsampling ratio of the average pooling layer is mainly achieved by adjusting subscale. Adjusting subscale can change FALoss, thereby affecting the detection ability of small targets. According to the characteristics of the dataset, subscale in the original model is adjusted to 0.0700, so as to minimize the loss of small target features while reducing the computational load of the model. The PicoheadV2 four-detection-head network is used to receive and process the fusion results of feature maps of different scales, and then send the results to the classification and regression network to achieve accurate detection and positioning of objects. It includes 4 branches with the same structure and 1 classification and regression network. Each branch generates a set of detection results, and the detection results include the object category and location, including: an improved convolutional module, upsampling, splicing, an improved convolutional module, upsampling, splicing, and BottleneckCSP2 structure connected in sequence; the classification and regression network is used to fuse the detection results of the 4 branches to obtain the final detection result.
[0076] The fusion results corresponding to feature maps of different scales are respectively used as the inputs of 4 head networks (Head1, Head2, Head3, and Head4). Each detection head independently predicts the received feature information to generate a set of detection results, including object category and location information. By processing the features of these different layers, multi-scale information can be fully learned, thereby enhancing the detection ability for objects of different sizes.
[0077] FALoss, a loss function for calculating the similarity between feature maps. It is characterized in that FALoss calculates the loss function by receiving two feature map parameters, feature1 and feature2, to be compared. The value of feature is determined by the AvgPool2d average pooling operation, and the pooling size is determined by subscale. And the size of the average pooling is calculated according to the formula subscale = int(1 / subscale). Among them, subscale is the downsampling ratio of the average pooling layer. Thus, the model performance can be optimized by adjusting subscale.
[0078] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0079] Example 1
[0080] Example 1 of the present invention proposes a small target detection method based on visible light and infrared multi-modal fusion, including:
[0081] Performing pixel-level fusion processing on the visible light image and the infrared image collected by the satellite or the drone;
[0082] Inputting the fused image into the trained detection model to output the detection result;
[0083] The detection model is an improved YOLOv5 neural network, including: an improved backbone network structure, an improved feature expression network, and an improved Head network structure; among them,
[0084] The improved backbone network structure is used to capture feature maps of different scales by introducing the BottleneckCSP structure and adopting an improved SPP structure, and input them into the improved feature expression network;
[0085] The improved feature expression network is used to obtain a feature layer by introducing the Leaky_Relu activation function into the convolutional layer part and adjusting the downsampling ratio of the average pooling layer, and input it into the improved Head network structure;
[0086] The improved Head network structure is used to perform object classification and regression by introducing the BottleneckCSP2 structure and output the detection result.
[0087] Input the training set into the improved YOLOv5 neural network, adjust the parameters until the training requirements are met, and obtain the trained detection model.
[0088] Figure 4 It is the block diagram of the unimproved backbone Backbone. The SiLU activation function and the unimproved SPP network are adopted. Figure 5 It is the block diagram of a small target real-time recognition model based on visible light and infrared multimodal fusion of the present invention. It mainly consists of two main parts: Backbone and Head (including the neck). The main function of the Backbone is to extract low-level texture and high-level semantic features. The extracted features will be sent into the Head network. Feature fusion is performed through the feature pyramid network structure to generate multi-scale features, and then rectangular boxes for predicting targets are generated according to the set anchor box sizes. After that, operations such as classification and localization are performed to complete object detection and output the recognition result. The main improvements in the Backbone part are: ① introducing an improved SPP feature fusion module. ② introducing an improved convolutional module. ③ introducing the MF module. The main improvements in the Head part are: ① introducing the BottleneckCSP2 module. ② introducing an improved convolutional module. The main improvement of the feature expression network is to introduce the Leaky_Relu activation function into the convolutional layer part to better capture and express small target features. In addition, the downsampling ratio of the average pooling layer is adjusted according to the characteristics of the dataset to minimize the loss of small target features while reducing the computational amount of the model.
[0089] Example 2
[0090] Embodiment 2 of the present invention provides a small target detection system based on visible light and infrared multi-modal fusion, which is implemented based on the method of Embodiment 1. The system includes:
[0091] A fusion module for performing pixel-level fusion processing on visible light images and infrared images collected by satellites or drones;
[0092] A detection module for inputting the fused image into a trained detection model and outputting a detection result;
[0093] The detection model is an improved YOLOv5 neural network, including: an improved backbone network structure, an improved feature expression network, and an improved Head network structure; wherein,
[0094] The improved backbone network structure is used to capture feature maps of different scales by introducing the BottleneckCSP structure and adopting an improved SPP structure, and input them into the improved feature expression network;
[0095] The improved feature expression network is used to obtain a feature layer by introducing the Leaky_Relu activation function into the convolutional layer part and adjusting the downsampling ratio of the average pooling layer, and input it into the improved Head network structure;
[0096] The improved Head network structure is used to perform object classification and regression by introducing the BottleneckCSP2 structure and output a detection result.
[0097] Figure 6 It is a comparison chart of the recognition results of the real-time small target recognition method based on visible light and infrared multi-modal fusion of the present invention and other methods. The specific steps of a real-time small target recognition method based on visible light and infrared multi-modal fusion are as follows: First, preprocess the collected original visible light and infrared images; then use pixel-level fusion technology to fuse the images of the two modalities; then input the fused image into the improved backbone network for small target feature extraction. The improved backbone network introduces the BottleneckCSP structure to enhance the feature extraction ability of the network; and, improve the SPP structure to enhance the ability of network feature fusion and transmission. Then, process the obtained feature map to obtain a feature layer and send it to the Head network (including the neck network). Replace the C3 structure corresponding to the P3 feature map for detecting small targets in the Head part with the BottleneckCSP2 structure, so as to further improve the model detection performance. Finally, send the feature layer into the object detection network for object classification and regression and output the detection result. Thus, the detection function of real-time recognition of small targets in visible light and infrared multi-modal is realized.
[0098] Table 1 shows the comparison of the detection results of the unimproved method and the improved method under the same conditions. The improved Multi-YOLO has increased by 19.77% in terms of mAP50, with a reduction of 7.0696M in Params and a reduction of 4.71 in GFLOPs.
[0099] Table 2 shows the detection results of single-modal visible small target recognition, single-modal infrared small target recognition, and multi-modal fusion small target recognition under the same conditions. The mAP of the single-modal visible small target recognition model is 75.06%. The mAP of the single-modal infrared small target recognition model is 70.83%. In contrast, the multi-modal Multi-YOLO model can effectively improve the detection accuracy of small targets.
[0100] Comparison between the unimproved and improved versions in Table 1
[0101]
[0102] Comparison between single-modal and multi-modal in Table 2
[0103]
[0104] Summary:
[0105] In view of the characteristics of remote sensing small and weak targets, the present invention improves the backbone network structure, the feature expression network structure, and the Head network structure on the basis of YOLOv5, thereby effectively improving the recognition accuracy.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A small target detection method based on visible light and infrared multimodal fusion, comprising: Perform pixel-level fusion processing on visible light images and infrared images collected by satellites or drones; Input the fused image into the trained detection model and output the detection result; The detection model is an improved YOLOv5 neural network, including: an improved backbone network structure, an improved feature expression network and an improved Head network structure; wherein, The improved backbone network structure is used to capture features of different scales to obtain feature maps by introducing a Bottleneck CSP structure and adopting an improved SPP structure, and input the feature maps into an improved feature expression network; The improved feature expression network is used to obtain a feature layer and input it into an improved Head network structure by introducing a Leaky_Relu activation function into the convolution layer part and adjusting the average pooling layer downsampling ratio; The improved Head network structure is used to perform target classification and regression and output detection results by introducing the Bottleneck CSP2 structure; The improved backbone network structure includes: an MF module, an improved convolution module, a BottleneckCSP structure, three improved convolution modules and an improved SPP structure connected in sequence; The MF module is used to extract shared and special information from different modalities and can bidirectionally combine multimodal internal information in a symmetrical and compact way; The improved BottleneckCSP structure is used to further improve the detection performance of the detection model; The improved convolution module is used to capture subtle feature changes; The improved SPP structure is used to capture features of different scales; The improved Head network structure includes 4 branches with the same structure and 1 classification and regression network, each branch generates a set of detection results, the detection results include object categories and positions, and each branch includes: an improved convolution module, upsampling, splicing, an improved convolution module, upsampling, splicing and BottleneckCSP2 structure connected in sequence; The improved convolution module is used to further process the feature map obtained by processing the improved backbone network structure; including Conv2d, BatchNorm2d and Leaky_Relu; The BottleneckCSP2 structure is used to fuse feature maps at different levels.
2. The small target detection method based on visible light and infrared multimodal fusion according to claim 1 is characterized in that: Improvements to the convolutional module include: The SiLu activation function of the convolution module is replaced with the Leaky_Relu function. The improved convolution module includes Conv2d, BatchNorm2d and Leaky_Relu.
3. The small target detection method based on visible light and infrared multimodal fusion according to claim 2 is characterized in that: The improvements of the SPP structure include: Introduce dilated convolution and maximum pooling of different scales to capture multi-scale features, and use grouped convolution in dilated convolution, with the number of groups equal to the number of input channels, to reduce the amount of computation; Use nearest neighbor interpolation to upsample the pooled feature map so that feature maps of different scales have the same spatial size; After each convolution operation, the Leaky_ReLU activation function is used to introduce non-linearity.
4. The small target detection method based on visible light and infrared multimodal fusion according to claim 3 is characterized in that: The processing of the improved SPP structure includes: The input feature map x passes through the convolution layer Conv1 and the batch normalization layer BN1, as well as the Leaky_ReLU activation function. The output of Conv1 and BN1 passes through the atrous convolution layer atrous_conv and the batch normalization layer BN_atrous and the Leaky_ReLU activation function. The maximum pooling layer Maxpool pools the input x to capture features of different scales to form Pooled_features; Each pooled feature map in Pooled_features is upsampled by nearest neighbor interpolation to match the spatial size of the original input feature map to form Pooled_features_upsampled; Concatenate the feature map x, the output x_atrous of the atrous convolution, and the upsampled pooled feature map Pooled_features_upsampled in the channel dimension to form features; The concatenated feature map features passes through the one-dimensional convolution layer Conv2 and the batch normalization layer BN2, and applies the Leaky_ReLU activation function to obtain the output x_concatenated.
5. The small target detection method based on visible light and infrared multimodal fusion according to claim 1 is characterized in that: In the improved feature expression network, adjusting the downsampling ratio of the average pooling layer is achieved by adjusting subscale.
6. The small target detection method based on visible light and infrared multimodal fusion according to claim 1 is characterized in that: The classification and regression network is used to fuse the detection results of the four branches to obtain the final detection result.
7. The small target detection method based on visible light and infrared multimodal fusion according to claim 1 is characterized in that: Includes the training steps for the detection model, including: The training set is input into the improved YOLOv5 neural network, and the parameters are adjusted until the training requirements are met to obtain a trained detection model.
8. A system based on the small target detection method based on visible light and infrared multimodal fusion according to claim 1, characterized in that: include: Fusion module, used to perform pixel-level fusion processing on visible light images and infrared images collected by satellites or drones; and The detection module is used to input the fused image into the trained detection model and output the detection result; The detection model is an improved YOLOv5 neural network, including: an improved backbone network structure, an improved feature expression network and an improved Head network structure; wherein, The improved backbone network structure is used to capture features of different scales to obtain feature maps by introducing a Bottleneck CSP structure and adopting an improved SPP structure, and input the feature maps into an improved feature expression network; The improved feature expression network is used to obtain the feature layer and input it into the improved Head network structure by introducing the Leaky_Relu activation function into the convolution layer part and adjusting the average pooling layer downsampling ratio; The improved Head network structure is used to perform target classification and regression and output detection results by introducing the Bottleneck CSP2 structure.
Citation Information
Patent Citations
Apple flower growth state detection method based on improved YOLOv5
CN115346212A
Substation equipment detection method and system based on bimodal data fusion
CN117351311A