Unmanned aerial vehicle visual angle infrared light and visible light fusion target detection method
By using data registration, residual network, self-attention module and cross-modal attention fusion methods from the perspective of the drone, the modal misalignment and resolution differences in infrared and visible light fusion target detection are solved, and the accuracy of small-object detection and robustness in complex environments are improved.
Patent Information
- Application Number
- CN202510331300.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The existing infrared and visible light fusion target detection technology at the perspective of drones has problems such as sensitivity between modes, poor adaptability at the same time, and low response rate to small targets.
The drone viewing infrared light and visible light fusion object detection method based on attention mechanism is adopted, and the multi-scale features are extracted through data registration, residual network, self-attention module enhances channel correlation, cross-modal attention fusion and dense connection fusion module are realized, combined with the YOLOv11 detection framework, dual-light fusion and object detection are realized.
It improves modal alignment capability, multi-scale feature fusion efficiency and small object detection accuracy, enhances detection robustness in complex environments, and significantly improves the accuracy and real-timeness of small object detection.
Smart Images

Figure CN120259822A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to an infrared light and visible light fusion target detection method from the perspective of an unmanned aerial vehicle (UAV). Background Art
[0002] UAV technology has been widely applied in multiple fields. Functions such as target detection, autonomous flight, and obstacle avoidance from the perspective of UAVs are becoming increasingly complete. In terms of security patrol, UAVs can quickly reach complex terrains or high-risk areas, transmit high-definition images in real time, and expand the monitoring range. In disaster relief, they can penetrate through ruins, mountains and forests, etc., and accurately locate trapped people. In the field of traffic monitoring, UAVs can overlook roads from high altitudes and comprehensively monitor road conditions. It has become a key technology in fields such as security, rescue, and traffic, greatly expanding the limitations of traditional monitoring means.
[0003] However, UAVs also need to cope with complex all-weather environments, and a single visible light modality is difficult to meet the requirements. Visible light images are rich in color and texture but are greatly affected by the environment, while infrared images are resistant to interference but details are lost. Currently, the target detection technology of infrared light and visible light fusion has gradually matured. The dual-light image fusion technology can combine the advantages of visible light and infrared light, enhance the adaptability of target detection, and perform excellently especially in complex environments.
[0004] However, the existing dual-light fusion target detection faces problems such as modality misalignment, resolution mismatch, and low response to small targets. When modality misalignment occurs, the dual-light fusion technology will have poor fusion effects, affecting the accuracy of target detection. When the resolutions do not match, that is, there is a large difference in the resolutions of visible light and infrared light images, the adaptability of the existing dual-light fusion technology is poor and it is difficult to effectively fuse. When the response to small targets is low, the targets are often small and easily blocked, making it difficult to meet the actual application requirements.
[0005] In summary, for the problems of UAV dual-light fusion target detection, such as being sensitive to misalignment between modalities, having poor adaptability when modal resolutions are different, having a low response rate to small targets, being unable to distinguish target rotation, and being affected by complex environmental factors, there is no method that can comprehensively solve them.
[0006] This paper proposes an infrared light and visible light fusion target detection method from the perspective of UAVs based on the attention mechanism, which solves the problems in the field of dual-light fusion target detection, such as being sensitive to misalignment between modalities, having poor adaptability when modal resolutions are different, and having a low response rate to small targets, and realizes effective dual-light fusion target detection from the perspective of UAVs. Summary of the Invention
[0007] The object of the present invention is to solve the drawbacks existing in the prior art, and a method for fusing infrared light and visible light for target detection from the perspective of an unmanned aerial vehicle (UAV) is proposed. The problems in the prior art in the field of dual-light fusion target detection, such as being sensitive to misalignment between modalities, poor adaptability when modal resolutions are different, and low response rate to small targets, are solved, and effective dual-light fusion target detection from the perspective of a UAV is achieved.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions: A method for fusing infrared light and visible light for target detection from the perspective of a UAV, comprising the following steps: Step 1: Collect infrared light and visible light image datasets of different scenes and different categories, divide them into a training set, a validation set and a test set, and perform registration of the datasets; Step 2: The backbone network is composed of residual modules. Each level of residual module performs convolution and pooling operations, and different-scale feature maps are obtained through downsampling; Step 3: An self-attention module is attached to the output of each level of residual module in the backbone network to extract the correlation of features between channels, and the obtained feature information is added to the backbone network; Step 4: Feature maps of the same scale in the infrared light and visible light branches are input into a cross-modal attention module to calculate the correlation information between modalities, and the cross-modal information of different scales is input into a fusion module; Step 5: The fusion module is designed based on a densely connected convolutional neural network. The output of each layer of upsampling is output to the previous layer, and the information of the entire network is aggregated in the last module to generate the final fusion result; Step 6: After the final fusion result of the fusion module is generated, it is input into the detection framework of YOLOv11 to obtain the detection result with annotations.
[0009] Through the above technical solutions, through a six-step process, data registration, a residual network to extract multi-scale features, an self-attention module to enhance channel correlation, cross-modal attention fusion, a densely connected fusion module, and a YOLOv11 detection framework, dual-light fusion and target detection are realized. Combining a residual network, an attention mechanism and a fusion module, the modal alignment ability, the multi-scale feature fusion efficiency and the small target detection accuracy are improved, the detection robustness in complex environments is enhanced, and the problems of the prior dual-light fusion target detection technology being sensitive to misalignment between modalities, poor adaptability to modal resolution differences, and low small target response rate are solved.
[0010] Preferably, in Step 1: Collect infrared light and visible light image datasets of different scenes and different categories, divide them into a training set, a validation set and a test set, and perform registration of the datasets: Data registration is achieved through feature point matching and affine transformation. Let the visible light image be , and the infrared light image be , the registered image is , and its formula is , where T is the affine transformation matrix, the transformation parameters, which are optimized by minimizing the feature point matching error , and are the coordinates of the matching feature points in the two-modal images, is to find the parameter that minimizes the norm.
[0011] Through the above technical solutions, feature point matching and affine transformation are used for data registration, and the transformation parameters are optimized to minimize the feature point matching error, realizing the spatial alignment of the two-modal images, providing geometrically consistent input for subsequent fusion, reducing misalignment interference, improving the fusion accuracy, and solving the modal misalignment problem caused by differences in viewing angles or sensors between infrared and visible light images.
[0012] Preferably, in step 2: the backbone network is composed of residual modules, and each level of residual module performs convolution and pooling operations, and different-scale feature maps are obtained through downsampling, including the following steps: The input feature map of the k-th layer residual module is , and its output feature map can be calculated by the following formula:
[0013] where, represents the combination of convolution, activation function and pooling operations, and the skip connection ensures the effective transmission of features. The specific convolution operation can be expressed as:
[0014] where, is the convolution operation, is the activation function, and Pool is the pooling operation; where, the output feature map of the k-th layer residual module, where, is the combined operation of convolution, activation and pooling, and the skip connection retains the original features through the identity mapping, avoiding the vanishing gradient.
[0015] Through the above technical solutions, the gradient propagation is stabilized through the skip connection, and the pooling operation realizes multi-scale downsampling, supports the detection of targets with different resolutions, enhances the adaptability of the model to complex scenes, solves the problem of vanishing gradient in the training of deep networks, and the lack of multi-scale feature expression ability.
[0016] Preferably, in step 3: an self-attention module is attached to the output of each residual module in the backbone network to extract the correlation between channel features, and the obtained feature information is added to the backbone network, including the following steps: Attach an self-attention module after the output of each residual module. Let the output feature map of the k-th residual module be , and the output of the self-attention mechanism is , which is calculated by the following formula:
[0017] where represents the query matrix, represents the key matrix, represents the value matrix, is the scaling factor, usually taking the square root of the dimension of the key. The calculated attention feature is added to the output feature map of the backbone network:
[0018] where is the output of the original residual module, is the feature obtained from the self-attention mechanism. The query matrix , key matrix , and value matrix of the self-attention module are obtained by linear transformation from the input feature map , and , while , , are learnable weight matrices, represents matrix multiplication.
[0019] Through the above technical solution, an self-attention module is attached after the residual module to calculate the inter-channel attention weights, enhance the key channel features, capture the long-range dependence relationship, highlight the important features, suppress the noise, improve the sensitivity and localization accuracy of the model to the target details, and solve the problem of weak inter-channel feature dependence and easy masking of key information by noise.
[0020] Preferably, in step 4: the feature maps of the same scale of the infrared light and visible light branches are input into the cross-modal attention module to calculate the inter-modal correlation information, and the cross-modal information of different scales is input into the fusion module, including the following steps: The feature maps of the same scale of the infrared light and visible light images are respectively input into the cross-modal attention module to calculate the inter-modal correlation information. Let the feature maps of the infrared light and visible light be and respectively. The output of the cross-modal attention module is M, and the calculation formula is as follows:
[0021] Among them, and are the query matrices for infrared light and visible light, and are the key matrices for infrared light and visible light, and are the value matrices for infrared light and visible light. The purpose of the cross-modal attention module is to calculate the correlation between modalities and generate the cross-modal feature representation M. The scaling factor in the cross-modal attention module is and the square root of the dimension of the key matrices . The inter-modal correlation feature M is calculated through bidirectional attention to enhance modal complementarity.
[0022] Through the above technical solution, the cross-modal attention module calculates the inter-modal correlation feature through bidirectional attention, enhances modal complementarity, combines the infrared global thermal radiation information with the visible light local texture, solves the problems of misalignment and resolution mismatch, improves the spatial consistency of the fused features, and solves the problem of poor fusion effect caused by insufficient complementarity between the infrared and visible light modalities and resolution differences.
[0023] Preferably, in step 5: The fusion module is designed based on a densely connected convolutional neural network. The upsampling output of each layer is fed back to the previous layer, and the information of the entire network is aggregated in the last module to generate the final fusion result, including the following steps: The fusion module is designed based on a densely connected convolutional neural network. At each layer, the features are passed to the previous layer through an upsampling operation. Let the output feature map of the Qth layer be , then:
[0024] Among them, represents the upsampling operation, is the input of the previous layer. Finally, the cross-modal features of all scales are fused through layer-by-layer connection and upsampling operations to generate the final fused feature representation. Assuming the final fused feature is , then:
[0025] Among them, is the final fused feature representation, and n is the number of layers.
[0026] Preferably, in step 5: The fusion module is designed based on a densely connected convolutional neural network. The upsampling output of each layer is fed back to the previous layer, and the information of the entire network is aggregated in the last module to generate the final fusion result, and further includes the following steps: Upsampling operation It is implemented by transposed convolution, and the output feature map of the Qth layer The generation formula is , where is the transposed convolution kernel weight, represents the skip connection, and the final fused feature is obtained through layer-by-layer weighted summation , is the learnable scale weight coefficient, is the feature alignment operation.
[0027] Through the above technical solution, the fusion module adopts dense connection and transposed convolution upsampling, generates the final feature through layer-by-layer weighted summation, retains multi-scale details, balances the contributions of high and low resolution features, enhances the small target detection ability, suppresses noise and improves robustness, and solves the problems of low multi-scale feature fusion efficiency and easy loss of small target information.
[0028] Preferably, after the final fusion result of the fusion module is generated, it is input into the detection framework of YOLOv11 to obtain the detection result with annotations, including the following steps: The final output of the fusion module is input into the YOLOv11 object detection framework. Let the fused feature map be , and the detection framework of YOLOv11 performs object detection on it to obtain the detection result with annotations:
[0029] where is the feature map output from the fusion module, is the predicted object category and bounding box information; The detection output of YOLOv11 contains the bounding box coordinates and the class probability , and its formula is , where is the Sigmoid function, is the detection head convolutional layer, and the output dimension is , S is the grid size, and C is the number of classes.
[0030] Through the above technical solution, the fused feature is input into the YOLOv11 detection framework, and the object bounding box and category are output. Combining the efficient detection ability of YOLOv11, the balance between real-time performance and high precision is achieved, and the detection rate of small targets is significantly improved in complex environments, solving the problems of poor adaptability of traditional detection frameworks to fused features and difficulty in balancing detection speed and accuracy.
[0031] The present invention has the following beneficial effects: The present invention provides an efficient method for fusing infrared light and visible light for target detection, which can effectively address problems such as misalignment between modalities, resolution differences, and small target detection. This method combines a residual network, an attention mechanism, and a fusion module, effectively improving the performance of dual-light fusion target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0033] Figure 1 is a flowchart of the present invention; Figure 2 is a structural diagram of the dual-light fusion and detection model framework of the present invention; Figure 3 is a structural diagram of the ResBlock module of the present invention; Figure 4 is a structural diagram of the SA-Block module of the present invention; Figure 5 is a structural diagram of the CA-Block module of the present invention; Figure 6 is a structural diagram of the FusionBlock module of the present invention; Figure 7 is a comparison diagram of the original visible light image, the original infrared image, and the image after fusion by this algorithm in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0035] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0036] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it is not necessary to further define and explain it in subsequent figures.
[0037] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use, or the orientation or positional relationship commonly understood by those skilled in the art. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.
[0038] In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.
[0039] An infrared light and visible light fusion target detection method from the perspective of an unmanned aerial vehicle, as Figures 1 to 6 shown, includes the following steps: Step 1: Collect infrared light and visible light image datasets of different scenes and different categories, and divide them into a training set, a validation set, and a test set, and perform dataset registration. The data registration is achieved through feature point matching and affine transformation. Let the visible light image be and the infrared light image be . The registered image is , and its formula is , where T is the affine transformation matrix, is the transformation parameter, and is optimized by minimizing the feature point matching error. and are the coordinates of the matching feature points in the two-modal images; Step 2: The backbone network consists of residual modules. Each level of residual module performs convolution and pooling operations and obtains feature maps of different scales through downsampling. The input feature map of the k-th layer of residual module is , and its output feature map can be calculated by the following formula:
[0040] where represents the combination of convolution, activation function, and pooling operations, and the skip connection ensures the effective transmission of features. The specific convolution operation can be expressed as:
[0041] where is a convolution operation, is an activation function, and Pool is a pooling operation; Among them, the output feature map of the k-th residual module , where, is a combined operation of convolution, activation, and pooling, The skip connection preserves the original features through the identity mapping to avoid the vanishing gradient; Step 3: An self-attention module is attached to the output of each level of the backbone network's residual module to extract the correlation between channel features, and the obtained feature information is added to the backbone network. An self-attention module is attached after the output of each residual module. Let the output feature map of the k-th residual module be , and the output of the self-attention mechanism is , which is calculated by the following formula:
[0042] where, represents the query matrix, represents the key matrix, represents the value matrix, is the scaling factor, usually taking the square root of the dimension of the key. The calculated attention feature is added to the output feature map of the backbone network:
[0043] where, is the output of the original residual module, is the feature obtained by the self-attention mechanism. The query matrix , key matrix , value matrix of the self-attention module are obtained by linear transformation from the input feature map , , and , , are learnable weight matrices, represents matrix multiplication; Step 4: The feature maps of the same scale of the infrared light and visible light branches are input into the cross-modal attention module to calculate the inter-modal correlation information. The cross-modal information of different scales is input into the fusion module. The feature maps of the same scale of the infrared light and visible light images are respectively input into the cross-modal attention module to calculate the inter-modal correlation information. Let the feature maps of the infrared light and visible light be and respectively. The output of the cross-modal attention module is M, and the calculation formula is as follows:
[0044] in, and is the query matrix for infrared light and visible light, and is the bond matrix of infrared light and visible light, and is the value matrix of infrared light and visible light. The purpose of the cross-modal attention module is to calculate the correlation between modalities and generate a cross-modal feature representation M. The scaling factor in the cross-modal attention module is is the key matrix and The square root of the dimension , the inter-modality correlation feature M is calculated through bidirectional attention to enhance the modality complementarity; Step 5: The fusion module is designed based on a densely connected convolutional neural network. Each layer is upsampled and output to the previous layer. In the last module, the information of the entire network is collected to produce the final fusion result. The fusion module is designed based on a densely connected convolutional neural network. In each layer, the features are passed to the previous layer through upsampling operations. Let the output feature map of the Qth layer be ,but:
[0045] in, represents the upsampling operation, is the input of the previous layer. Finally, the cross-modal features of all scales are fused through layer-by-layer connection and upsampling operations to generate the final fused feature representation. Assume that the final fused feature is ,but:
[0046] in, is the final fusion feature representation, n is the number of layers; Upsampling Operation Using transposed convolution, the Qth layer outputs the feature map The formula for generating ,in, is the transposed convolution kernel weight, Represents skip connection, and finally fuses features By layer-by-layer weighted summation , is the learnable scale weight coefficient, It is the feature alignment operation; Step 6: After the final fusion result of the fusion module is generated, it is input into the detection framework of yolov11 to obtain the detection result with annotations. The final output of the fusion module is input into the YOLOv11 target detection framework. Let the fused feature map be , the detection framework of YOLOv11 performs object detection on it and obtains the detection results with annotations:
[0047] Among them, is the feature map output from the fusion module, is the predicted object category and bounding box information; The detection output of YOLOv11 contains the bounding box coordinates and the class probabilities , and its formula is , where is the Sigmoid function, is the detection head convolutional layer, and the output dimension is , S is the grid size, and C is the number of classes.
[0048] Combining the above content and Figures 1 to 5 : The ResBlock module is a residual module, corresponding to Figure 2 , and the ResBlock module is the core component of the backbone network, which is used to extract multi-scale features and maintain the stability of feature transmission. Its structure design is as follows: The input and the skip connection input feature map are processed through two branches. The main branch performs convolution, ReLU activation function, and pooling operations in sequence. The skip connection directly transmits the input feature map to the output end to avoid the problem of gradient disappearance in the deep network; the calculation formula , where is the combination operation: , the convolution operation uses a 3×3 convolution kernel, the stride is 1, and the padding is 1 to keep the size of the feature map unchanged. The pooling operation uses max pooling, the kernel size is 2×2, and the stride is 2 to achieve downsampling; The original feature information is retained through the skip connection, the gradient propagation ability is enhanced, and the multi-scale downsampling provides feature maps with different resolutions for subsequent modules to support multi-scale object detection.
[0049] Figure 2 In , the ResBlock is a residual module, the SA-Block is a self-attention module, the CA-Block is a cross-modal attention module, the Fusion Block is a fusion module, and the Detection Block is a detection module.
[0050] The SA-Block module is a self-attention module, corresponding to Figure 3, the SA-Block module is attached to the output of each residual module to capture long-range dependencies between channels. Its structure includes the following steps: Feature transformation: The input feature map generates a query matrix through linear transformation , a key matrix , and a value matrix . The attention weights are calculated by the scaled dot-product attention mechanism: , is the square root of the dimension of the key matrix, used to stabilize the gradient. The function normalizes the attention weights to highlight important channel features; Feature fusion: Add the attention features and the original features to obtain an enhanced feature map . The self-attention mechanism enhances the correlation between channels, improves the model's sensitivity to key features, is suitable for object localization in complex backgrounds, and reduces noise interference.
[0051] Figure 3 In - Figure 3 , ResBlock is the residual module, Conv1×1 is the 1×1 convolution, LReLU is the leaky rectified linear unit (activation function), Conv3×3 is the 3×3 convolution, and Pool2×2 is the 2×2 pooling.
[0052] The CA-Block module is the cross-modal attention module, corresponding to Figure 4 . The CA-Block module is used to fuse the features of two modalities, infrared light (IR) and visible light (VIS), to solve the problems of misalignment and resolution difference between modalities. Its structure is as follows: Input the infrared feature map FIR and the visible light feature map FVIS of the same scale as the input and the feature alignment. Calculate the correlation between modalities through the bidirectional attention mechanism. The bidirectional attention calculates the inter-modal correlation feature M through the following formula: ; . To ensure the stability of the attention weights, the bidirectional attention mechanism considers the correlations of both IR→VIS and VIS→IR, enhances the modality complementarity, solves the problem of geometric misalignment between modalities, improves the spatial consistency of the fused features, combines the global thermal radiation information of the infrared image and the local texture details of the visible light image, and optimizes the object representation ability.
[0053] Figure 4In this, SA-Block is the self-attention module, MatMul is matrix multiplication, Softmax is the Softmax function, MatMul & Scale is matrix multiplication and scaling, LinearQ1 is the linear transformation (query 1), Lineark1 is the linear transformation (key 1), Linearv1 is the linear transformation (value 1), LinearQ2 is the linear transformation (query 2), Lineark2 is the linear transformation (key 2), Linearv2 is the linear transformation (value 2), X^1 is the input 1, and X^2 is the input 2.
[0054] The FusionBlock module is the fusion module, corresponding to Figure 5 , the FusionBlock module is based on a dense connection structure, integrates multi-scale cross-modal features, and generates the final fusion result. Its design is as follows: Upsampling and skip connection The feature map of the l-th layer is upsampled by transposed convolution and skip-connected to the feature of the previous layer, , is the parameter of the transposed convolution kernel, retaining the features of the original resolution; Multi-scale weighted fusion The final fusion feature is generated by layer-by-layer weighted summation , the dense connection retains multi-scale detailed information, enhances the small object detection ability, balances the contributions of high and low resolution features through weighted fusion, suppresses noise and improves the robustness of fusion.
[0055] Figure 5 In this, the FusionBlock is the fusion module, is the bidirectional connection block 0,2, is the bidirectional connection block 0,1, is the bidirectional connection block 0,0, is the bidirectional connection block 1,2, is the bidirectional connection block 1,1, is the bidirectional connection block 2,2.
[0056] Compared with the prior art, the present invention addresses the problems existing in the existing dual - light fusion object detection technology and proposes a novel, more efficient and accurate dual - light fusion object detection method. Specifically, compared with the prior art, the beneficial effects of the present invention are as follows: The present invention addresses the problems existing in the existing dual - light fusion object detection technology and proposes a novel, more efficient and accurate dual - light fusion object detection method. Specifically: It solves the problem of misalignment between modalities: By introducing a self - attention mechanism in the feature extraction module, the present invention effectively captures the relationship between the infrared light and visible light modalities, thus solving the influence of misaligned modality on the detection accuracy in traditional dual - light fusion methods. It improves the multi - modal data fusion ability: Through the cross - modal attention module, the present invention can fully fuse the multi - scale feature information of infrared light and visible light images, thereby optimizing the recognition ability for different scenes and targets, especially the detection of small targets. It improves the accuracy of small target detection: During the object detection process, in combination with the YOLOv11 framework, the present invention effectively improves the detection ability of small targets after multi - modal fusion. Especially in complex backgrounds or low - light environments, the object detection effect is significantly improved. Specifically, see Figure 7 , Figure 7 The first row in it is the original visible light image, the second row is the original infrared image, and the third row is the image after fusion by this algorithm. Combining with the attached Figure 7 It can be seen that the labeled image output after multi - modal fusion by the method of the present invention has higher clarity than the original visible light image, can adapt to the situation where the environmental visible light is darker, so as to improve the object detection effect. Compared with the original infrared image, it can output a colorful image. It is not only convenient for the system to autonomously perform object detection actions, but also can be clearer and have richer details when output to the artificial perspective for review. With the cooperation of artificial and automated image processing, the detection accuracy can be significantly improved.
[0057] And the algorithm framework of the present invention is tested on the DroneVehicle dataset, and the following test results are provided: [1] Ale L, Zhang N, Li L. Road damage detection using RetinaNet[C] / / 2018 IEEE International Conference on Big Data (Big Data). IEEE, 2018: 5197 - 5200. [2] Ren S, He K, Girshick R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE transactions on pattern analysis and machine intelligence, 2016, 39(6): 1137-1149. [3] Han J, Ding J, Li J, et al. Align deep features for oriented object detection[J]. IEEE transactions on geoscience and remote sensing, 2021, 60: 1-11. [4] Redmon J, Divvala S, Girshick R, et al. You only look once: Unified, real-time object detection[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 779-788. [5] Sun Y, Cao B, Zhu P, et al. Drone-based RGB-infrared cross-modality vehicle detection via uncertainty-aware learning[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(10): 6700-6713. [6] Yuan M, Wei X. C 2 former: Calibrated and complementary transformer for rgb-infrared object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024. In the above table, mAP is the mean average precision, a commonly used evaluation metric in object detection, which is used to measure the detection accuracy of the model for different classes of objects. Its calculation method is to first calculate the AP for each class, that is, the area under the precision-recall curve, and then calculate the average value of the APs for all classes. The higher the mAP value, the better the detection performance of the model; Speed is the number of frames processed per second, which is used to measure the inference speed of the model. The higher the fps value, the better the real-time performance of the model; The proposed algorithm framework was tested on the LLVIP dataset, and the following test results are provided: [1] Ma J, Yu W, Liang P, et al. FusionGAN: A generative adversarial network for infrared and visible image fusion[J]. Information fusion, 2019, 48: 11-26. [2] Li H, Wu X J, Durrani T. NestFuse: An infrared and visible image fusion architecture based on nest connection and spatial / channel attention models[J]. IEEE Transactions on Instrumentation and Measurement, 2020, 69(12): 9645-9656. [3] Zhang H, Ma J. SDNet: A versatile squeeze-and-decomposition network for real-time image fusion[J]. International Journal of Computer Vision, 2021, 129(10): 2761-2785. [4] Ma J, Zhang H, Shao Z, et al. GANMcC: A generative adversarial network with multiclassification constraints for infrared and visible image fusion[J]. IEEE Transactions on Instrumentation and Measurement, 2020, 70: 1-14. In the above table, Precision measures the ability of the model to correctly predict positive samples, and is calculated as the ratio of true positives (TP) to the number of samples predicted as positive (TP + FP). The higher the precision, the more accurate the model's prediction of positive samples; Recall measures the ability of the model to find all actual positive samples, and is calculated as the ratio of true positives (TP) to the number of all actual positive samples (TP + FN). The higher the recall, the more comprehensive the model's coverage of positive samples; mAP@0.50 is the mean average precision calculated when the intersection over union (IOU) threshold between the predicted bounding box and the ground truth bounding box is 0.50, and is used to evaluate the detection performance of the model at this threshold; mAP@[0.5:0.95] is the mean average precision calculated when the IOU threshold ranges from 0.5 to 0.95 (step size 0.05), comprehensively evaluating the detection performance of the model under different levels of strictness. The larger this metric value, the better the detection effect of the model under different IOU thresholds.
[0058] In addition, the method of the present invention enhances the fusion of multi-scale information: by introducing a fusion module designed with a densely connected convolutional neural network, the present invention can better combine feature maps of different scales, enhancing the performance of the model in processing multi-scale targets. In summary, the present invention not only has a significant improvement in accuracy and robustness compared to the prior art, but also effectively reduces the computational complexity, has good practicability and broad application prospects, and is particularly suitable for the dual-modal fusion target detection task in complex environments.
[0059] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An infrared light and visible light fusion target detection method from the perspective of an unmanned aerial vehicle, characterized in that, It includes the following steps: Step 1: Collect infrared and visible light image datasets of different scenarios and categories, divide them into training sets, validation sets, and test sets, and perform registration of the datasets; Step 2: The backbone network is composed of residual modules. Each level of residual module performs convolution and pooling operations, and obtains feature maps of different scales through downsampling; Step 3: An self-attention module is attached to the output of each level of residual module in the backbone network to extract the correlation of features between channels, and the obtained feature information is added to the backbone network; Step 4: Input the feature maps of the same scale of the infrared and visible light branches into the cross-modal attention module to calculate the inter-modal correlation information, and input the cross-modal information of different scales into the fusion module; Step 5: The fusion module is designed based on a densely connected convolutional neural network. The upsampling output of each layer is fed to the previous layer, and the information of the entire network is aggregated in the last module to generate the final fusion result; Step 6: After the final fusion result of the fusion module is generated, it is input into the detection framework of yolov11 to obtain the detection result with annotations.
2. The method for detecting a target by fusing infrared light and visible light from an unmanned aerial vehicle perspective according to claim 1, characterized in that, Step 1: Collect infrared and visible light image datasets of different scenarios and categories, divide them into training sets, validation sets, and test sets, and perform registration of the datasets: Data registration is achieved through feature point matching and affine transformation. Let the visible light image be , and the infrared light image be . The registered image is . Its formula is , where T is the affine transformation matrix, is the transformation parameter, which is optimized by minimizing the feature point matching error . and are the coordinates of the matching feature points in the two-modal images, is to find the parameter that minimizes the norm.
3. The infrared light and visible light fusion target detection method from the perspective of an unmanned aerial vehicle according to claim 1, wherein The step 2: The backbone network is composed of residual modules. Each level of residual module performs convolution and pooling operations, and obtains feature maps of different scales through downsampling, including the following steps: The input feature map of the k-th layer residual module is , and its output feature map can be calculated by the following formula: , where represents the combination of convolution, activation function, and pooling operations. The skip connection ensures the effective transmission of features. The specific convolution operation can be expressed as: , where is the convolution operation, is the activation function, and Pool is the pooling operation; among them, the output feature map of the k-th layer residual module , where is the combined operation of convolution, activation, and pooling, The skip connection preserves the original features through the identity mapping, avoiding the vanishing gradient.
4. The infrared light and visible light fusion target detection method from the perspective of an unmanned aerial vehicle according to claim 1, wherein The step 3: An self-attention module is attached to the output of each level of residual module in the backbone network to extract the correlation of features between channels, and the obtained feature information is added to the backbone network, including the following steps, Append a self-attention module after the output of each residual module. Let the output feature map of the k-th layer residual module be and the output of the self-attention mechanism be which is calculated by the following formula: where represents the query matrix, represents the key matrix, represents the value matrix, is the scaling factor, usually taking the square root of the dimension of the key. The calculated attention feature is added to the output feature map of the backbone network: where is the output of the original residual module, is the feature obtained by the self-attention mechanism. The query matrix , key matrix , and value matrix of the self-attention module are obtained by linear transformation from the input feature map , while , , are learnable weight matrices, represents matrix multiplication.
5. The method for fusing infrared light and visible light for target detection from the perspective of an unmanned aerial vehicle according to claim 1, wherein The step 4: Input the feature maps of the same scale of the infrared and visible light branches into the cross-modal attention module to calculate the inter-modal correlation information, and input the cross-modal information of different scales into the fusion module, including the following steps: The feature maps of the same scale of infrared and visible light images are respectively input into the cross-modal attention module to calculate the correlation information between modalities. Suppose the feature maps of infrared and visible light are and , the output of the cross-modal attention module is M, and the calculation formula is as follows: ,in, and is the query matrix for infrared light and visible light, and is the bond matrix of infrared light and visible light, and is the value matrix of infrared light and visible light. The purpose of the cross-modal attention module is to calculate the correlation between modalities and generate a cross-modal feature representation M. The scaling factor in the cross-modal attention module is is the key matrix and The square root of the dimension ,The inter-modality correlation feature M is calculated through bidirectional attention to enhance the modality complementarity.
6. The method for detecting a target by fusing infrared light and visible light from an unmanned aerial vehicle perspective according to claim 1, wherein The step 5: The fusion module is designed based on a densely connected convolutional neural network. The upsampling output of each layer is fed to the previous layer, and the information of the entire network is aggregated in the last module to generate the final fusion result, including the following steps: The fusion module is designed based on a densely connected convolutional neural network. At each layer, features are passed to the previous layer through an upsampling operation. Let the output feature map of the Q-th layer be , then: , where represents the upsampling operation, is the input of the previous layer. Finally, cross-modal features at all scales are fused through layer-by-layer connection and upsampling operations to generate the final fused feature representation. Assuming the final fused feature is , then: , where is the final fused feature representation, and n is the number of layers.
7. A method for detecting a target by fusing infrared light and visible light from a drone perspective according to claim 6, characterized in that, The step 5: The fusion module is designed based on a densely connected convolutional neural network. The upsampling output of each layer is fed to the previous layer, and the information of the entire network is aggregated in the last module to generate the final fusion result, and further includes the following steps: Upsampling operation Implemented using transposed convolution, the output feature map of the Q-th layer The generation formula is , where is the transposed convolution kernel weight,[[]] represents the skip connection, and the final fused feature is obtained by weighted summation layer by layer , is the learnable scale weight coefficient,[[]] is the feature alignment operation.[[]] 8. The method for fusing infrared light and visible light for target detection from the perspective of an unmanned aerial vehicle according to claim 1, wherein The step 6: After the final fusion result of the fusion module is generated, it is input into the detection framework of yolov11 to obtain the detection result with annotations, including the following steps: The final output of the fusion module is input into the YOLOv11 object detection framework. Let the fused feature map be . The detection framework of YOLOv11 performs object detection on it and obtains the detection results with annotations: , where is the feature map output from the fusion module, and is the predicted object category and bounding box information; Detection Output of YOLOv11 Contains bounding box coordinates And class probabilities , and its formula is , where Is the Sigmoid function Is the detection head convolutional layer, and the output dimension is , S is the grid size, and C is the number of classes.
Citation Information
Cited By
Infrared and visible light target detection method, electronic equipment, storage medium and product
CN120495644A
Method and system for detecting nonferrous metal target of scraped car
CN120953758A
A method and system for detecting non-ferrous metal targets in scrapped automobiles
CN120953758B
Target detection method of crossing machine based on heteroocular modal crossing
CN121170645A
Infrared fusion target detection and identification method under complex low-light background condition
CN121190736A