Visible light and infrared image fusion method based on feature matching under view angle of unmanned aerial vehicle
By using a feature matching-based method, the computational efficiency and hardware dependency issues of fusion of visible light and infrared images from the perspective of UAVs were solved, achieving efficient and robust image fusion and generating high-quality fused images.
Patent Information
- Application Number
- CN202510715587.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing methods for fusing visible light and infrared images from the perspective of UAVs suffer from low computational efficiency, high hardware dependence, and unstable fusion quality. In particular, they are difficult to meet real-time requirements and are susceptible to noise interference in UAV inspections.
We employ a feature-matching-based approach, combining image preprocessing, object detection, feature extraction, and feature-level fusion with a lightweight network for image reconstruction, achieving efficient and robust image fusion.
The real-time and robustness of image fusion are improved, the hardware cost is reduced, and information-rich and clear fused images are generated.
Smart Images

Figure CN120823463A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field, and in particular relates to a visible light and infrared image fusion method based on feature matching from the perspective of an unmanned aerial vehicle. Background Art
[0002] Object detection systems from drones have been widely used in engineering inspections. Most existing drone-based object detection solutions often rely on visible light or infrared imagery for prediction. Visible light images offer high resolution and contrast, but they struggle to capture low-light conditions, such as at night and inclement weather. Infrared images, on the other hand, allow for all-day detection and can capture objects that visible light images cannot. However, infrared images typically have lower resolution and suffer from interlaced textures. Therefore, the rational use of visible light and infrared images for complementary fusion can yield richer semantic information, resulting in a robust and information-rich fused image.
[0003] Existing technologies usually use pixel fusion to fuse visible light and infrared light, but they still face many technical bottlenecks and challenges in practical applications, specifically in the following aspects:
[0004] Computational efficiency and real-time performance: In practical applications such as drone inspections, the raw image data is massive. Traditional pixel-level fusion methods (such as multi-scale transformations and deep learning networks) require pixel-by-pixel calculations, making the fusion process extremely time-consuming and difficult to meet real-time requirements. This is especially true for end-to-end fusion networks based on deep learning, whose complex model structure and training process further exacerbate the computational burden.
[0005] Hardware dependence and cost issues: Pixel-level fusion requires precise alignment of visible light and infrared images at the pixel level, which places extremely high demands on the synchronization and stability of imaging equipment and increases hardware costs.
[0006] Fusion quality and robustness: Pixel information is susceptible to contamination and noise, leading to unstable quality of raw pixels, which can affect the fusion effect after superposition. Traditional multi-scale transformation methods struggle to effectively distinguish noise from real features when processing complex scenes. While deep learning networks have strong feature extraction capabilities, they are prone to overfitting or insufficient generalization when data is insufficient or scenes vary widely, affecting the robustness of the fusion results.
[0007] In response to the shortcomings of pixel-level fusion methods in terms of computational efficiency, hardware dependence, and fusion quality, this patent proposes a feature-level visible light and infrared image fusion algorithm. This algorithm abandons the traditional pixel-by-pixel processing approach and instead starts from the feature level. By extracting and aligning the salient features of the same target in the visible light image and the infrared image, it achieves efficient and robust image fusion. Summary of the Invention
[0008] The purpose of the present invention is to solve the problems in the prior art and propose a visible light and infrared image fusion method based on feature matching from the perspective of a drone.
[0009] To address the above issues, the present invention aims to provide a method for fusion of visible light and infrared images based on feature matching from the perspective of a drone, which includes the following parts:
[0010] Image preprocessing module, used to receive visible light image I vis and infrared image I ir As input, preprocessing operations are performed to improve image quality, eliminate modality differences, and provide high-quality input data for subsequent modules;
[0011] The target detection module is used to detect the target from the visible light image I vis and infrared image I ir Detect targets in the two images and match them in two images through cross-modal association, providing target information for subsequent feature extraction and fusion;
[0012] The feature extraction module is used to extract the complementary modal features of targets in visible light and infrared images, including shape, texture, and edge information, providing rich and robust feature representation for subsequent feature fusion;
[0013] The feature-level fusion module aligns and fuses the features extracted from visible light images and infrared images to generate information-rich and complementary fused features, providing high-quality feature representation for subsequent tasks. It includes a feature alignment submodule and a feature fusion submodule.
[0014] Feature reconstruction module, the optimized fusion feature F optimized The reconstruction is performed into high-quality images, preserving the complementary information of visible and infrared images while ensuring the clarity and details of the reconstructed images.
[0015] In the above-mentioned visible light and infrared image fusion method based on feature matching from the perspective of a drone, the image preprocessing module includes the following processing steps:
[0016] S1: Cross-modal registration: Use the affine transformation matrix to perform geometric correction on the image to eliminate the field of view offset. The reference formula is as follows:
[0017]
[0018] Where a, b, c, d are rotation and scaling parameters, t x , t y The translation amount is calculated by optimizing the parameters through SIFT feature point matching, and then the registered visible light image and infrared image are output;
[0019] S2: Noise suppression. On the one hand, Gaussian filtering is performed on the infrared image to smooth the noise. The reference formula is as follows:
[0020]
[0021] Among them, w(x,y) is the Gaussian kernel function used to calculate the weight between pixels, and Z is the normalization factor;
[0022] On the other hand, non-local mean denoising is used for visible light images to retain image details. The reference formula is as follows:
[0023]
[0024] Among them, w(x,y) is the similarity weight, which is calculated based on the similarity between pixel blocks, Z is the normalization factor, and finally the denoised visible light image and infrared image are output.
[0025] S3: Dynamic range compression: Perform histogram equalization on the infrared image to avoid overexposed or underexposed areas. The reference formula is as follows:
[0026]
[0027] Among them, n i is the gray level frequency, N is the total number of pixels;
[0028] Finally, the infrared image with dynamic range compression is output.
[0029] In the above-mentioned visible light and infrared image fusion method based on feature matching from the perspective of a drone, the target detection module includes the following processing steps:
[0030] S1: Bimodal detection network: using a YOLOv5 network with a shared backbone to process visible light images simultaneously vis and infrared image I ir , output target bounding box B ir , B vis And confidence C, classification probability P cls Calculated by Softmax function:
[0031] P cls =Softmax(W·F roi +b)
[0032] Among them, F roi is the feature extracted by ROIAlign, W is the classification weight, and b is the bias term;
[0033] S2: Cross-modal association: Based on the Hungarian algorithm, the intersection over union (IoU) of the target bounding box and feature similarity are combined to match the same target in the visible light and infrared images. The matching score is S match The calculation is as follows:
[0034] S match =α·IoU(B ir ,B vis )+β·cos(F ir ,F vis )
[0035] α and β are balancing weights, controlling the contribution of IoU and feature similarity respectively;
[0036] Output the matching target pair and the matching score S match .
[0037] In the above-mentioned visible light and infrared image fusion method based on feature matching from the perspective of the drone, the feature extraction module extracts:
[0038] Visible light features, which use the ResNet-50 network to extract high-frequency texture features in visible light images, including edges and gradients. The calculation is as follows:
[0039]
[0040] Among them, * represents the convolution operation, W l and b l are the weight and bias terms of the lth layer respectively;
[0041] Then output the visible light feature F vis ;
[0042] Infrared features: The infrared features use an attention mechanism to enhance the saliency of thermal targets in infrared images;
[0043] Among them, the attention weight A ir The calculation is as follows:
[0044] A ir =σ(W a ·[F ir ,F vis ]+b a )
[0045] Among them, σ is the Sigmoid function, W a and b a are the weights and biases of the attention mechanism, [F ir ,F vis ] represents infrared signature F ir and visible light Fvis splicing of features;
[0046] Then use the attention weight A ir Used to weight infrared features and highlight the significance of thermal targets:
[0047]
[0048] Finally output infrared characteristics
[0049] In the above-mentioned visible light and infrared image fusion method based on feature matching from the perspective of a drone, the feature alignment submodule is used to ensure the spatial and semantic consistency of features. The workflow includes:
[0050] Geometric alignment: The geometric alignment is based on the matching pairs (B vis ,B ir ), and use the affine transformation matrix to perform spatial registration on the infrared feature map:
[0051] θ=argmin|F vis -T θ (F ir )|2
[0052] Where T θ is the affine transformation operator; θ is optimized by back propagation to minimize the visible light feature F vis and transformed infrared signature T θ (F ir ), and then output the aligned infrared features
[0053] Semantic alignment: Use cross-modal contrastive learning to construct feature similarity constraints and maximize the similarity of target cross-modal features:
[0054]
[0055] where cos(F vis ,F ir ) represents the visible light feature F vis and infrared signature F ir The cosine similarity between them; τ is the temperature coefficient, which controls the smoothness of the similarity distribution; K is the number of negative samples, and its output is the semantically aligned visible light feature F vis and infrared signature F ir .
[0056] In the above-mentioned visible light and infrared image fusion method based on feature matching from the perspective of a drone, the feature fusion submodule includes:
[0057] Feature fusion:
[0058] F fused =A·F vis +(1-A)·F ir
[0059] Among them, A is the attention weight, which is used to dynamically allocate the contribution of the two modal features, and its output is the preliminary fusion feature F fused .
[0060] Feature optimization: Optimize the fused features to eliminate noise and inconsistency, and improve the robustness and discriminability of the features. The optimization process is achieved through convolutional neural networks:
[0061] F optimized =ReLU(W o *F fused +b o )
[0062] Among them, W o and b o To optimize the weights and biases of the network, the output is the optimized fusion feature F optimized .
[0063] In the above-mentioned visible light and infrared image fusion method based on feature matching from the perspective of a UAV, the feature reconstruction module includes a decoding network part and an image post-processing part;
[0064] The decoding network uses a lightweight convolutional neural network or a generative adversarial network as a decoder to fuse the features F optimized Decoded into an image,
[0065] I reconstructed =Decoder(F optimized ).
[0066] In the above-mentioned visible light and infrared image fusion method based on feature matching from the perspective of a drone, the image post-processing part includes the following steps:
[0067] S1: Denoising: removes noise from the reconstructed image.
[0068] S2: Sharpening: Use the Laplacian operator or adaptive sharpening filter to enhance image edges and details;
[0069] S3: Dynamic range adjustment: Optimize image brightness and contrast using histogram equalization or adaptive contrast stretching:
[0070] I final =PostProcess(I reconstructed )
[0071] Then output the final reconstructed image I final .
[0072] In the above-mentioned visible light and infrared image fusion method based on feature matching under the perspective of the UAV,
[0073] This patent addresses the shortcomings of pixel-level fusion methods in terms of computational efficiency, hardware dependency, and fusion quality. This patent proposes a feature-level method for fusion of visible and infrared images. This method abandons the traditional pixel-by-pixel approach and instead approaches the image fusion process from a feature-level perspective. By extracting and aligning salient features of the same target in the visible and infrared images, this method achieves efficient and robust image fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 This is a flow chart of a visible light and infrared image fusion method based on feature matching from the perspective of a drone.
[0075] Figure 2 This is a decoding network structure diagram of a visible light and infrared image fusion method based on feature matching from the perspective of a drone. DETAILED DESCRIPTION
[0076] Reference Figure 1-2 , a visible light and infrared image fusion method based on feature matching from the perspective of a drone, including the following modules:
[0077] 1) Image preprocessing module
[0078] This module receives the visible light image I vis and infrared image I ir As input, preprocessing operations are performed to improve image quality, eliminate modality differences, and provide high-quality input data for subsequent modules.
[0079] Specific workflow:
[0080] Step 1: Cross-modal registration: Use the affine transformation matrix to perform geometric correction on the infrared / visible light image to eliminate the field of view offset.
[0081]
[0082] Where a, b, c, d are rotation and scaling parameters, t x , t y is the translation amount, and the parameters are optimized through SIFT feature point matching.
[0083] Output: registered visible light image and infrared image.
[0084] Step 2: Noise suppression, including:
[0085] ① Perform Gaussian filtering on the infrared image to smooth the noise.
[0086]
[0087] Where w(x,y) is the Gaussian kernel function used to calculate the weights between pixels, and Z is the normalization factor.
[0088] ② Non-local mean denoising is used for visible light images to preserve image details:
[0089]
[0090] Where w(x,y) is the similarity weight, which is calculated based on the similarity between pixel blocks, and Z is the normalization factor.
[0091] Output: Denoised visible light image and infrared image.
[0092] Step 3: Dynamic range compression: Perform histogram equalization on the infrared image to avoid overexposed or underexposed areas.
[0093]
[0094] Among them, n i is the grayscale frequency, and N is the total number of pixels.
[0095] Output: Infrared image after dynamic range compression.
[0096] 2) Object Detection Module
[0097] This module is responsible for converting the visible light image I vis and infrared image I ir The target is detected in the two images and matched with each other through cross-modal association, providing accurate target information for subsequent feature extraction and fusion.
[0098] Specific workflow:
[0099] Step 1: Dual-modal detection network: Use a YOLOv5 network with a shared backbone to process visible light images simultaneously. vis and infrared image I ir , output target bounding box B iy , B vis And confidence C. Classification probability P cls Calculated by Softmax function:
[0100] P cls =Softmax(W·F roi +b)
[0101] Among them, F roiis the feature extracted by ROIAlign, W is the classification weight, and b is the bias term.
[0102] Step 2: Cross-modal association: Based on the Hungarian algorithm, the intersection over union (IoU) of the target bounding box and feature similarity are combined to match the same target in the visible light and infrared images. Matching score S match The calculation is as follows:
[0103] S match =α·IoU(B ir ,B vis )+β·cos(F ir ,F vis )
[0104] α and β are balancing weights that control the contribution of IoU and feature similarity respectively.
[0105] Output: matched target pair, and matching score S match .
[0106] 3) Feature extraction module
[0107] Extract the modal complementary features of targets in visible light and infrared images, including key information such as shape, texture, and edge, to provide rich and robust feature representation for subsequent feature fusion.
[0108] Among them, visible light features: use the ResNet-50 network to extract high-frequency texture features (such as edges and gradients) in visible light images. Features of the first layer The calculation is as follows:
[0109]
[0110] Among them, * represents the convolution operation, W l and b l are the weight and bias of the lth layer respectively.
[0111] Output: Visible light feature F vis .
[0112] Infrared Features: Enhancing the saliency of thermal targets in infrared images using attention mechanism.
[0113] Attention weight A ir The calculation is as follows:
[0114] A ir =σ(W a ·[F ir ,F vis ]+b a )
[0115] Among them, σ is the Sigmoid function, W aand b a are the weights and biases of the attention mechanism. [F ir ,F vis ] represents infrared signature F ir and visible light F vis Feature splicing.
[0116] Use attention weight A ir Used to weight infrared features and highlight the significance of thermal targets:
[0117]
[0118] Output: Infrared signature
[0119] 4) Feature-level fusion module
[0120] This module is responsible for aligning and fusing the features extracted from visible light images and infrared images to generate information-rich and complementary fused features, providing high-quality feature representation for subsequent tasks.
[0121] (1) Feature alignment submodule
[0122] Due to the different imaging principles of visible light and infrared images, the features of the same target in the two modalities may have spatial offsets or scale differences. Therefore, it is necessary to perform geometric and semantic alignment on the extracted features to ensure spatial and semantic consistency of the features.
[0123] Working principle and process:
[0124] Step 1: Geometric alignment: Based on the matching pairs output by the target detection module (B vis ,B ir ), and use the affine transformation matrix to perform spatial registration on the infrared feature map:
[0125] θ=argmin|F vis -T θ (F ir )|2
[0126] Where T θ is the affine transformation operator; θ is optimized by back propagation to minimize the visible light feature F vis and transformed infrared signature T θ (F ir ) between the two.
[0127] Output: aligned infrared signatures
[0128] Step 2: Semantic alignment: Use cross-modal contrastive learning (CMCL) to construct feature similarity constraints to maximize the similarity of target cross-modal features:
[0129]
[0130] where cos(F vis ,F ir ) represents the visible light feature F vis and infrared signature F ir The cosine similarity between them; τ is the temperature coefficient, which controls the smoothness of the similarity distribution; K is the number of negative samples.
[0131] Output: semantically aligned visible light features F vis and infrared signature F ir .
[0132] (2) Feature fusion submodule
[0133] The aligned visible light features and infrared features are weightedly fused to generate a fused feature representation.
[0134] Working principle and process:
[0135] Step 1: Feature fusion:
[0136] F fused =A·F vis +(1-A)·F ir
[0137] Among them, A is the attention weight, which is used to dynamically allocate the contributions of the two modal features.
[0138] Output: Preliminary fusion feature F fused .
[0139] Step 2: Feature Optimization:
[0140] The fused features are optimized to eliminate noise and inconsistency, and to improve the robustness and discriminative ability of the features. The optimization process is achieved through convolutional neural networks:
[0141] F optimized =ReLU(W o *F fused +b o )
[0142] Among them, W o and b o To optimize the weights and biases of the network
[0143] Output: optimized fusion feature F optimized .
[0144] 5) Feature reconstruction
[0145] The fused feature F optimized The reconstruction is performed into high-quality images, preserving the complementary information of visible and infrared images while ensuring the clarity and details of the reconstructed images.
[0146] Working principle and process:
[0147] Step 1: Decoding the network:
[0148] Use a lightweight convolutional neural network (CNN) or generative adversarial network (GAN) as a decoder to fuse the features F optimized Decode into an image.
[0149] I reconstructed =Decoder(F optimized )
[0150] Step 2: Image post-processing,
[0151] include:
[0152] ① Denoising: Use non-local means or bilateral filtering to remove noise from the reconstructed image.
[0153] ②Sharpening: Use the Laplacian operator or adaptive sharpening filter to enhance image edges and details.
[0154] ③Dynamic range adjustment: Use histogram equalization or adaptive contrast stretching to optimize the brightness and contrast of the image.
[0155] I final =PostProcess(I reconstructed )
[0156] Output: Final reconstructed image I final , which fuses the complementary information of visible and infrared images with high-quality clarity and details.
[0157] It is understood from common technical knowledge that the present invention may be implemented by other embodiments that do not depart from its spirit or essential features. Therefore, the embodiments disclosed above are, in all respects, merely illustrative and not exclusive. All modifications within the scope of the present invention or equivalent to the scope of the present invention are intended to be encompassed by the present invention.
Claims
1. A visible light and infrared image fusion method based on feature matching from the perspective of a drone, characterized by: Includes the following sections: Image preprocessing module, used to receive visible light image I vis and infrared image I ir As input, preprocessing operations are performed to improve image quality, eliminate modality differences, and provide high-quality input data for subsequent modules; The target detection module is used to detect the target from the visible light image I vis and infrared image I ir Detect targets in the two images and match them in two images through cross-modal association, providing target information for subsequent feature extraction and fusion; The feature extraction module is used to extract the complementary modal features of targets in visible light and infrared images, including shape, texture, and edge information, providing rich and robust feature representation for subsequent feature fusion; The feature-level fusion module aligns and fuses the features extracted from visible light images and infrared images to generate information-rich and complementary fusion features, providing high-quality feature representation for subsequent tasks. Includes feature alignment submodule and feature fusion submodule; Feature reconstruction module, the optimized fusion feature F optimized The reconstruction is performed into high-quality images, preserving the complementary information of visible and infrared images while ensuring the clarity and details of the reconstructed images.
2. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 1, characterized in that: The image preprocessing module includes the following processing steps: S1: Cross-modal registration: Use the affine transformation matrix to perform geometric correction on the image to eliminate the field of view offset. The reference formula is as follows: Where a, b, c, d are rotation and scaling parameters, t x , t y The translation amount is calculated by optimizing the parameters through SIFT feature point matching, and then the registered visible light image and infrared image are output; S2: Noise suppression. On the one hand, Gaussian filtering is performed on the infrared image to smooth the noise. The reference formula is as follows: Among them, w(x,y) is the Gaussian kernel function used to calculate the weight between pixels, and Z is the normalization factor; On the other hand, non-local mean denoising is used for visible light images to retain image details. The reference formula is as follows: Among them, w(x,y) is the similarity weight, which is calculated based on the similarity between pixel blocks, Z is the normalization factor, and finally the denoised visible light image and infrared image are output. S3: Dynamic range compression: Perform histogram equalization on the infrared image to avoid overexposed or underexposed areas. The reference formula is as follows: Among them, n i is the gray level frequency, N is the total number of pixels; Finally, the infrared image with dynamic range compression is output.
3. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 1 is characterized in that: The target detection module includes the following processing steps: S1: Bimodal detection network: using a YOLOv5 network with a shared backbone to process visible light images simultaneously vis and infrared image I ir , output target bounding box B ir , B vis And confidence C, classification probability P cls Calculated by Softmax function: P cls =Softmax(W·F roi +b) Among them, F roi is the feature extracted by ROIAlign, W is the classification weight, and b is the bias term; S2: Cross-modal association: Based on the Hungarian algorithm, the intersection over union (IoU) of the target bounding box and feature similarity are combined to match the same target in the visible light and infrared images. The matching score is S match The calculation is as follows: S match =α·IoU(B ir ,B vis )+β·cos(F ir ,F vis ) α and β are balancing weights, controlling the contribution of IoU and feature similarity respectively; Output the matching target pair and the matching score S match .
4. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 1 is characterized in that: The feature extraction module includes: Visible light features, which use the ResNet-50 network to extract high-frequency texture features in visible light images, including edges and gradients. The calculation is as follows: Among them, * represents the convolution operation, W l and b l are the weight and bias terms of the lth layer respectively; Then output the visible light feature F vis ; Infrared features: The infrared features use an attention mechanism to enhance the saliency of thermal targets in infrared images; Among them, the attention weight A ir The calculation is as follows: A ir =σ(W a ·[F ir ,F vis ]+b a ) Among them, σ is the Sigmoid function, W a and b a are the weights and biases of the attention mechanism, [F ir ,F Vis ] represents infrared signature F ir and visible light F vis splicing of features; Then use the attention weight A ir Used to weight infrared features and highlight the significance of thermal targets: Finally output infrared characteristics 5. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 1 is characterized in that: The feature alignment submodule is used to ensure the spatial and semantic consistency of features. The workflow includes: Geometric alignment: The geometric alignment is based on the matching pairs (B vis ,B ir ), and use the affine transformation matrix to perform spatial registration on the infrared feature map: Where T θ is the affine transformation operator; θ is optimized by back propagation to minimize the visible light feature F vis and transformed infrared signature T θ (F ir ), and then output the aligned infrared features Semantic alignment: Use cross-modal contrastive learning to construct feature similarity constraints and maximize the similarity of target cross-modal features: where cos(F vis ,F ir ) represents the visible light feature F vis and infrared signature F ir The cosine similarity between them; τ is the temperature coefficient, which controls the smoothness of the similarity distribution; K is the number of negative samples, and its output is the semantically aligned visible light feature F vis and infrared signature F ir .
6. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 2 is characterized in that: The feature fusion submodule includes: Feature fusion: F fused =A·F vis +(1-A)·F ir Among them, A is the attention weight, which is used to dynamically allocate the contribution of the two modal features, and its output is the preliminary fusion feature F fused . Feature optimization: Optimize the fused features to eliminate noise and inconsistency, and improve the robustness and discriminability of the features. The optimization process is achieved through convolutional neural networks: F optimized =ReLU(W o *F fused +b o ) Among them, W o and b o To optimize the weights and biases of the network, the output is the optimized fusion feature F optimized .
7. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 2, characterized in that: The feature reconstruction module includes a decoding network part and an image post-processing part; The decoding network uses a lightweight convolutional neural network or a generative adversarial network as a decoder to fuse the features F optimized Decoded into an image, I reconstructed =Decoder(F optimized )。 8. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 2, characterized in that: The image post-processing part The following steps are involved: S1: Denoising: removing noise from the reconstructed image; S2: Sharpening: Use the Laplacian operator or adaptive sharpening filter to enhance image edges and details; S3: Dynamic range adjustment: Optimize image brightness and contrast using histogram equalization or adaptive contrast stretching: I final =PostProcess(I reconstructed ) Then output the final reconstructed image I final .
9. The visible light and infrared image fusion method based on feature matching from the perspective of a drone according to claim 8, characterized in that: The denoising method adopts non-local mean or bilateral filtering method.
Citation Information
Patent Citations
Multi-ship tracking method and device based on multi-modal information fusion
CN118379328A
RGB-t multispectral pedestrian detection method based on target aware fusion strategy
US20240331403A1
Cited By
Infrared and visible light image fusion method and system combined with self-supervised feature alignment
CN121213374A
Low, small and slow target detection and tracking method and system based on multi-modal fusion
CN121884015A
A low, small and slow target detection and tracking method and system based on multi-modal fusion
CN121884015B