Infrared and visible light image real-time fusion method and device

Through the lightweight convolutional neural network (CNN) and a specially designed loss function, the problem of information distinction and speed contradiction in the fusion of infrared and visible light images is solved, and real-time and efficient image fusion is achieved, preserving salient targets and background details.

CN120707405APending Publication Date: 2025-09-26INSPUR QILU SOFTWARE IND

Patent Information

Application Number
CN202511046625.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods have a contradiction between fusion quality and processing speed on edge computing platforms, and traditional methods are difficult to effectively distinguish information from different regions, resulting in the weakening of useful information.

Method used

A lightweight convolutional neural network (CNN) is used to highlight the thermal target area in the infrared image by designing a special loss function and a binary mask image. During the training process, the network is guided to focus on the salient area, and pixel-level fusion is performed by combining similarity loss and target information loss.

Benefits of technology

It achieves real-time and efficient image fusion on the edge computing platform, with complete preservation of salient area information and rich background texture details, reducing model complexity and computational complexity, and is suitable for multiple computing platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707405A_ABST
    Figure CN120707405A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and artificial intelligence, and particularly provides an infrared and visible light image real-time fusion method and device, and the method comprises the steps: firstly, the important information of an infrared image is reflected in a target which can emit more heat, and a lightweight CNN (convolutional neural network) is used for learning the automatic detection of the regions from the infrared image; then, required information is accurately extracted from the detected area, and effective fusion and reconstruction are carried out; in a lightweight convolutional neural network CNN, a loss functional expression is designed, a binary mask image is added, an object radiating a large amount of heat in an infrared image is highlighted, and a salient region where a CNN concerned target is located is guided during training. Compared with the prior art, the method can solve the problem that the salient target area in the image cannot be effectively utilized in the prior art, guarantees the real-time performance of fusion calculation, effectively reduces the operation cost of the algorithm, and achieves the precise and efficient fusion of the infrared visible light image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and artificial intelligence, and specifically provides a method and device for real-time fusion of infrared and visible light images. Background Art

[0002] Infrared imaging technology is based on differences in the target's thermal radiation, making it highly resistant to interference. Unaffected by adverse weather conditions such as smoke, fog, and low illumination, it operates around the clock and is often used for target identification. However, infrared images generally have low resolution, and background texture details are unclear. In contrast, visible light imaging technology can capture rich texture details in a scene, resulting in higher resolution and clarity for both targets and backgrounds. However, visible light imaging systems are susceptible to imaging conditions, low illumination, and occlusions, making it difficult for visible light images to capture and identify targets under these adverse conditions. This complementary nature creates the opportunity to fuse the two, achieving the desired result of capturing both thermal targets and rich texture details. Fusion of infrared and visible light images is a fundamental and key technology, widely used in various fields, including optoelectronic detection, military affairs, and remote sensing.

[0003] Image fusion can generally be categorized into three different types: pixel-based, feature-based, and decision-based. Pixel-based fusion directly fuses the source images pixel by pixel in the spatial or frequency domain according to a specific fusion principle. Feature-based fusion first extracts image features and then uses these features to generate a fusion result. Decision-based fusion independently completes the downstream tasks of each sensor and generates the optimal decision.

[0004] Traditional infrared and visible light image fusion methods mostly use pixel-based fusion, and the main processing stages include image transformation, transformation coefficient fusion, and inverse transformation, such as multi-scale transformation methods and sparse representation methods. To achieve higher image fusion quality, traditional methods design sophisticated models and complex fusion rules, which significantly increases the difficulty of algorithm design. Due to the powerful feature extraction and reconstruction capabilities of deep neural networks, learning-based infrared and infrared image recognition technologies have been well applied. However, deep learning networks tend to adopt feature-based fusion. Generally speaking, end-to-end approaches require effective feature representation, which often relies on large-scale network architectures and complex training strategies. Therefore, balancing the contradiction between fusion quality and processing speed becomes a serious challenge, especially on edge computing platforms, where real-time deployment of algorithms is crucial.

[0005] On the other hand, while many deep learning-based image fusion methods exist, some shortcomings remain. For example, they often fail to distinguish between different regions of the source images when constructing loss functions. This introduces a large amount of redundant or even invalid information into the fusion process, inevitably weakening the useful information in the fused image. Consequently, the desired information in the fused image is difficult to define, which in turn constrains the training process and the design of the loss function. Summary of the Invention

[0006] The present invention aims to overcome the above-mentioned deficiencies in the prior art and provides a highly practical method for real-time fusion of infrared and visible light images.

[0007] A further technical task of the present invention is to provide a reasonably designed, safe and applicable real-time fusion device for infrared and visible light images.

[0008] The technical solution adopted by the present invention to solve its technical problem is:

[0009] A real-time fusion method for infrared and visible light images. First, the important information of infrared images is reflected in the presence of targets that emit more heat. A lightweight convolutional neural network (CNN) learns to automatically detect these areas from infrared images.

[0010] Then, the required information is accurately extracted from the detected areas and effectively fused and reconstructed;

[0011] In the lightweight convolutional neural network (CNN), a loss function is designed and a binary mask image is added to highlight objects that radiate a lot of heat in the infrared image, guiding the CNN to focus on the salient area where the target is located during training.

[0012] Furthermore, during the training phase of the model, the loss function determines the type of information retained in the fused image and the proportional relationship between various types of information;

[0013] Define H, W to represent the height and width of the image, I ir ,I vis Represent the infrared image and visible light image to be fused, I f represents the fused image;

[0014] The target mask image is a binary mask image obtained by annotation. During training, it guides the network to detect the salient area where the target is located. The pixel value of the target position in the mask image is 1, and the pixel value of the rest of the area is 0.

[0015] Furthermore, the loss function of the network structure consists of two losses: similarity loss and target information loss;

[0016] The similarity loss uses the structural similarity index (SSIM) to calculate the similarity between the fused image and the infrared and visible light images respectively. Since the SSIM value is between -1 and 1, the following similarity loss function can be defined:

[0017] L ssim =1-ssim(I f ,I ir )+1-ssim(I f ,I vis ) (1)

[0018] The similarity loss constrains the fused image to learn the brightness, contrast, and structural similarity of infrared and visible light in order to learn the overall intensity distribution of the input image.

[0019] Furthermore, the fused image should be close to the infrared image in the target area and close to the visible light image in the background area. The information contained in the fused image should be measured in two dimensions: pixel and gradient. Therefore, the target information loss should be further divided into pixel loss and gradient loss, as shown below:

[0020] L sa =αL sa-pix +(1-α)L sa-grad (1)

[0021] Pixel loss constrains the pixel intensity of the fused image to be consistent with the source image, while gradient loss focuses on the edge and detail information of the image. α is a balance parameter, and the mask divides the entire image into target and background areas. Therefore, pixel loss and gradient loss are performed separately in two areas.

[0022] Furthermore, the pixel loss constrains the similarity between the fused image and the source image in the salient area and background area through the pixel-by-pixel absolute error (L1 norm). The specific formula is as follows:

[0023]

[0024] Among them, λ is the balance parameter, e is the dot product, and the loss function adopts l1-norm. The mask image avoids the interference of the visible light image in the target area through the dot product, retaining only the thermal target presented by infrared, while only the texture information of the visible light image is required in the background area;

[0025] Similarly, the gradient loss also consists of the following two parts:

[0026]

[0027] Where μ is the equilibrium parameter, Represents the gradient operator, using the Sobel operator;

[0028] In summary, the expression of the total loss function is as follows, where β is used to balance the similarity loss and the target information loss;

[0029] L total =L ssim +βL sa (5).

[0030] Furthermore, the convolutional neural network (CNN) performs pixel-level fusion of images in the spatial domain. Given a pair of registered infrared and visible light images, which are input into the corresponding pathways of the network respectively, the network will output two weight maps.

[0031] The double channel stacking operation is used to address the insufficient feature extraction capability of shallow networks, and the spatial resolution of the feature map of each convolution module remains the same as that of the input image;

[0032] After the input image has undergone feature extraction in one layer of convolution, it will be spliced ​​and input into the shared convolution block.

[0033] Furthermore, the convolutional neural network (CNN) contains a total of 8 convolution blocks, with a convolution kernel size of 3×3. The output channel of the last convolution block conv7 and conv8 is 1. In addition, the number of output channels of the remaining convolution blocks is 8. The last convolution layer is followed by a Sigmoid layer to limit the pixel value range of the weight map to [0, 1].

[0034] When inputting the network model, the RGB visible light image needs to be converted to the YCbCr color space first, then the Y channel is used to fuse the grayscale infrared image, and finally the fused image is converted back to the RGB space through inverse conversion.

[0035] Furthermore, the specific formula is as follows: after weighting, normalization is performed to obtain a single-channel grayscale image;

[0036]

[0037] Where W ir , W vis Two weight graphs representing the network output respectively;

[0038] The mask image is only used in the training phase of the network model. The mask image is used to participate in the calculation of the loss function value, thereby guiding the model training through gradient backpropagation. The mask image is not required in the model inference phase.

[0039] A device for real-time fusion of infrared and visible light images, comprising: at least one memory and at least one processor;

[0040] The at least one memory is configured to store a machine-readable program;

[0041] The at least one processor is used to call the machine-readable program to execute a real-time fusion method of infrared and visible light images.

[0042] Compared with the prior art, the method and device for real-time fusion of infrared and visible light images of the present invention have the following outstanding beneficial effects:

[0043] (1) The present invention retains the target information in the mask image to guide the network training, and distinguishes and enhances the target area information (mainly from the infrared image) and the background texture information (mainly from the visible light image) through a specially designed loss function, thereby effectively extracting the target characteristics.

[0044] (2) A lightweight network design is used to generate pixel-level fusion weight maps directly in the spatial domain, significantly reducing model complexity and computational complexity while ensuring the speed of model inference. This method requires only a small amount of data to complete model training and is easy to deploy. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 It is a flowchart of a method for real-time fusion of infrared and visible light images;

[0047] Figure 2 It is a schematic diagram of infrared visible light image and mask image in the real-time fusion method of infrared and visible light images;

[0048] Figure 3 This is a schematic diagram of the network architecture of a real-time fusion method of infrared and visible light images;

[0049] Figure 4 It is a schematic diagram of fused images in a real-time fusion method of infrared and visible light images;

[0050] Figure 5 This is an example of the TNO dataset test results for the real-time fusion method of infrared and visible light images;

[0051] Figure 6 This is an example of the test results of the MSRS dataset in the real-time fusion method of infrared and visible light images;

[0052] Figure 7 It is a detailed comparison chart of infrared image and fused image in the real-time fusion method of infrared and visible light images;

[0053] Figure 8 It is a schematic diagram of three sets of source images to be fused in a real-time fusion method of infrared and visible light images;

[0054] Figure 9 This is a schematic diagram comparing three image fusion methods in real-time fusion of infrared and visible light images. DETAILED DESCRIPTION

[0055] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0056] A best embodiment is given below:

[0057] Example 1:

[0058] like Figure 1 As shown, in this embodiment, in the fusion of infrared and visible light images, the most critical information is the salient target in the infrared image and the background texture structure in the visible light image.

[0059] Based on this definition, image fusion has two key points. The first is that the important information in infrared images is primarily reflected in objects that emit more heat (such as pedestrians, vehicles, and shelters). CNNs should learn to automatically detect these areas in infrared images. The second key is to accurately extract the required information from the detected areas and perform effective fusion and reconstruction.

[0060] A specific loss function and a lightweight convolutional neural network (CNN) structure were designed to solve the above two key problems.

[0061] Among them, the loss function is:

[0062] During the CNN network training phase, the loss function determines the type of information retained in the fused image and the proportional relationship between the various types of information. H ,W represent the height and width of the image, I ir ,I vis Represent the infrared image and visible light image to be fused respectively, I f Represents the fused image.

[0063] The target mask image is a binary mask image obtained by annotation. Its purpose is to highlight the objects that radiate a lot of heat in the infrared image and guide the network to focus on the salient area where the target is located during training. The pixel value of the target position in the mask image is 1, and the pixel value of the rest of the area is 0. Figure 2 shown.

[0064] The loss function of the network structure consists of two types of losses: similarity loss and target information loss;

[0065] The similarity loss uses the structural similarity index (SSIM) to calculate the similarity between the fused image and the infrared and visible light images respectively. Since the SSIM value is between -1 and 1, the following similarity loss function can be defined:

[0066] L ssim =1-ssim(I f ,I ir )+1-ssim(I f ,I vis ) (1)

[0067] The similarity loss constrains the fused image to learn the brightness, contrast, and structural similarity of infrared and visible light in order to learn the overall intensity distribution of the input image.

[0068] The fused image should be close to the infrared image in the target area and close to the visible light image in the background area. The information contained in the fused image should be measured in two dimensions: pixel and gradient. Therefore, the target information loss should be further divided into pixel loss and gradient loss, as shown below:

[0069] L sa =αL sa-pix +(1-α)L sa-grad (2)

[0070] Pixel loss constrains the pixel intensity of the fused image to be consistent with the source image, while gradient loss focuses on the edges and details of the image, with α being the balancing parameter. Since the mask divides the entire image into target and background regions, pixel loss and gradient loss can be performed separately on the two regions.

[0071] The pixel loss constrains the similarity between the fused image and the source image in the salient area and background area through the pixel-by-pixel absolute error (L1 norm). The specific formula is as follows:

[0072]

[0073] Where λ is the balancing parameter, e is the dot product, and the loss function uses l1-norm. The mask image avoids the interference of the visible light image in the target area through the dot product, retaining only the thermal target presented by infrared, while only the texture information of the visible light image is required in the background area.

[0074] Similarly, the gradient loss also consists of the following two parts:

[0075]

[0076] Where μ is the equilibrium parameter, Represents the gradient operator, using the Sobel operator.

[0077] In summary, the expression of the total loss function is as follows, where β is used to balance the similarity loss and the target information loss.

[0078] L total =L ssim +βL sa (5)

[0079] The lightweight network structure is designed as follows:

[0080] In order to meet the requirements of real-time and high efficiency, a convolutional neural network (CNN) is constructed to perform pixel-level fusion of images in the spatial domain. Given a pair of registered infrared images and visible light images, which are input into the corresponding pathways of the network respectively, the network will output two weight maps. The specific network structure is as follows: Figure 3 shown.

[0081] In order to cope with the insufficient feature extraction capabilities of shallow networks, the network structure uses two channel stacking operations (concat), and the spatial resolution of the feature map of each convolution module remains the same as that of the input image to reduce information loss.

[0082] After the input images have undergone feature extraction by one layer of convolution, they are spliced ​​and input into the shared convolution block. The purpose of this design is that the input may have different characteristics, and it is necessary to extract features suitable for their characteristics through their respective convolution blocks. The spliced ​​features are further fused through the shared convolution block. The use of shared convolution blocks takes into account that even if the imaging is in different bands, the features in the same scene must have similarities, such as common near-infrared and visible light imaging. This network structure uses shared convolution to reduce the number of parameters while enhancing the generalization ability of features.

[0083] This network structure contains a total of 8 convolution blocks, and the convolution kernel size is 3×3. The output channel of the last convolution block conv7 and conv8 is 1. In addition, the number of output channels of the remaining convolution blocks is 8. The last convolution layer is followed by a Sigmoid layer to limit the pixel value range of the weight map to [0, 1].

[0084] When inputting the network model, the RGB visible light image needs to be converted to the YCbCr color space first. Then, the Y channel is used to fuse with the grayscale infrared image. Since the structural details mainly exist in the Y channel, the fused image retains the Cb and Cr channels of the visible light image. Finally, through the inverse conversion, the fused image can be converted back to the RGB space. The calculation of the fused image is as follows Figure 4 shown.

[0085] The specific formula is as follows: After weighting, normalization is performed to obtain a single-channel grayscale image.

[0086]

[0087] Where W ir ,W vis Represent the two weight graphs of the network output respectively.

[0088] It should be noted that the mask image is only used in the training phase of the network model. The mask image is used to participate in the calculation of the loss function value, thereby guiding the model training through gradient backpropagation. The mask image is not required in the model inference phase.

[0089] It can be seen that the pixel-level fusion network of the present invention is a simple and naive design. It does not introduce the complex structure of residual blocks or attention mechanisms. The entire training process only uses 8 convolution blocks, and the loss function explicitly guides the network to distinguish between targets and backgrounds, without relying on deep features extracted by large-scale networks.

[0090] 50 pairs of infrared and visible light images were selected for training from the TNO dataset. The TNO dataset is a commonly used public dataset for the fusion of infrared and visible light images, which contains a variety of military-related scenes.

[0091] The input image and the corresponding mask image were cropped into 128×128 tiles each time. The cropping position was randomized each time, but the cropping position of all three images was the same. The training parameters were set as follows: batch size was set to 64, learning rate was set to 0.0005, iterations were set to 50, α = 0.5, β = 2, and Adam was used as the optimizer for training the model. The proposed algorithm was implemented in PyTorch. In addition, considering that the target area occupies a very small proportion of the image, λ = μ = 7 was set.

[0092] First, experiments were conducted on the TNO dataset, and the results are as follows Figure 5 As shown, from left to right are infrared, visible light and fusion images.

[0093] It can be seen that the fused image takes into account the characteristics of both infrared images and visible light images. Compared with visible light, the target in the fused image is more obvious. At the same time, due to the participation of the infrared image in the fusion, the areas with excessive brightness in the visible light band are suppressed, thereby better highlighting the target details.

[0094] In order to verify the generalization ability of the model on an untrained dataset, the MSRS dataset is selected for testing. The MSRS dataset contains several daily scenes such as streets, where the spatial resolution of the images is 480×640. The fusion results are shown in the figure below. Figure 6 As shown, from left to right are infrared, visible light and fusion images.

[0095] The MSRS dataset contains several nighttime scenes, where visible light images can hardly identify thermal targets, and the grayscale values ​​of infrared images are very low except for a few areas. For such scenes, the image fusion method should focus on retaining the infrared thermal targets with concentrated information in the image, while coordinating the fusion results with the visible light background brightness. Figure 7 As shown in the third set of fusion scenes, the left side is the infrared image and the right side is the fused image. The red rectangle marks the target's significant thermal radiation area in the image, which is effectively retained in the fusion result.

[0096] This method uses a pixel-level image fusion scheme. Most common schemes in traditional image fusion are also pixel-level fusion. Therefore, in order to show the effect of this method, we choose to compare it with the pyramid fusion method and the DCT (discrete cosine transform) fusion method. The results are as follows: Figure 8 As shown in FIG, the source image, and the same scene group is infrared image and visible light image from left to right respectively.

[0097] like Figure 9 As shown in the figure, each group shows the results of our method, pyramid fusion, and DCT methods, from left to right. It can be seen that our method effectively preserves the pedestrian target and background areas, with intensities coordinated and background texture details not lost. For example, in the first group, the pyramid fusion and DCT methods produce dim brightness and unclear details in the background tree area. Furthermore, the pyramid method produces a significant difference in brightness between the pedestrian area and the background, failing to refer to the intensity distribution of the visible light image. In the last group of images, the pyramid method also produces larger pixel intensity values ​​in the pedestrian area. Of the three methods, the fused image produced by our method is more consistent with the actual scene.

[0098] The present invention has conducted experiments on multiple data sets, proving that the proposed algorithm can retain more source image texture details and salient target information in low-light environments such as those with large differences between infrared and visible light vision, and that the fused image can have good clarity, facilitating subsequent visual tasks.

[0099] Example 2:

[0100] In the data preparation and preprocessing stage;

[0101] (1) Data collection: According to the specific scenario requirements of image fusion, obtain the registered infrared image and visible light image pairs. The data source can be actual application scenario collection or public datasets (such as TNO, MSRS).

[0102] (2) Mask image generation: For each pair of training images, manually annotate the significant thermal target areas (such as pedestrians, vehicles, etc.) in the infrared image. The pixel values ​​of the target area are marked as 1, and the pixel values ​​of the background area are marked as 0, and the corresponding binary mask image (Mask) is generated. This mask image is only used in the training phase to guide the network to focus on the target area.

[0103] (3) Visible light image conversion: If the visible light image is multi-channel, it needs to be converted to the YCbCr color space first to separate the brightness component (Y_vis) and the chrominance component (Cb, Cr).

[0104] (4) Infrared image processing: If the infrared image is multi-channel, take its grayscale image or perform grayscale processing to obtain a single-channel infrared image (I_ir_gray).

[0105] (5) Image block cropping: To facilitate training, the registered I_ir_gray, Y_vis, and Mask are randomly cropped into image blocks of a fixed size (e.g., 128×128 pixels). The cropping position must ensure that the three are spatially aligned.

[0106] Model training phase:

[0107] Network input: Paired image blocks (I_ir_gray, Y_vis) are input to the lightweight convolutional neural network of the present invention. At the same time, the corresponding mask image blocks are used for loss calculation.

[0108] Network structure: Use Figure 2 The network architecture shown in Figure 2 is a graph of the network. The network consists of two independent initial convolution paths, which output two weight maps (W_ir, W_vis) with a value range of [0, 1]. The fused image Y_fused is obtained according to formula (6).

[0109] Loss calculation and parameter optimization:

[0110] The total loss function includes: (1) Similarity loss (L_ssim): Use SSIM to calculate the similarity loss between Y_fused and I_ir_gray, Y_vis (Formula (1)) (2) Target information loss (L_tar): Includes pixel loss (L_pixel, Formula (3)) and gradient loss (L_grad, Formula (4)).

[0111] The loss uses Mask to distinguish target areas (prioritizing infrared thermal targets) from background areas (prioritizing visible light textures) and balances pixel intensity (L1) and edge details (Sobel gradient).

[0112] Use an optimizer (such as Adam) to backpropagate the calculated total loss and update the network weights. Training parameters can be set to: batch size 64, initial learning rate 0.0005, and training epochs 50. The training objective is to minimize the total loss function.

[0113] Inference phase and deployment phase:

[0114] Input: A new infrared image (I_ir) and a visible light image (I_vis) to be fused. Register them if not already registered. Convert I_vis to YCbCr space to obtain Y_vis, Cb, Cr. Grayscale the infrared image to obtain I_ir_gray. Note: A mask image is not required at this stage.

[0115] Network forward propagation: Input the aligned I_ir_gray and Y_vis (whole image or block) into the model trained in step 2. The network outputs the corresponding weight maps W_ir and W_vis.

[0116] Fused image calculation: Calculate the fused luminance component image Y_fused according to formula (6). Combine Y_fused with the chrominance components Cb and Cr of the original visible light image to form a fused YCbCr image. Finally, convert the YCbCr image back to RGB space to obtain the final fused color image (Fused_RGB).

[0117] Call the deployed model for inference and output the fusion result image / video stream for subsequent display or analysis.

[0118] Example 3:

[0119] In this embodiment, a device for real-time fusion of infrared and visible light images includes: at least one memory and at least one processor;

[0120] The at least one memory is configured to store a machine-readable program;

[0121] The at least one processor is used to call the machine-readable program to execute a real-time fusion method of infrared and visible light images.

[0122] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0123] The memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can also include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state memory devices.

[0124] The above-mentioned specific implementation methods are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementation methods. Any technical solutions that conform to the above-mentioned specific implementation methods of the present invention and any appropriate changes or substitutions made thereto by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.

[0125] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for real-time fusion of infrared and visible light images, characterized in that: First, the important information of infrared images is reflected in the presence of targets that emit more heat. A lightweight convolutional neural network (CNN) is used to automatically detect these areas from infrared images. Then, the required information is accurately extracted from the detected areas and effectively fused and reconstructed; In the lightweight convolutional neural network (CNN), a loss function is designed and a binary mask image is added to highlight objects that radiate a lot of heat in the infrared image, guiding the CNN to focus on the salient area where the target is located during training.

2. The method for real-time fusion of infrared and visible light images according to claim 2, characterized in that: During the training phase of the lightweight convolutional neural network (CNN), the loss function determines the type of information retained in the fused image and the proportional relationship between the various types of information. Define H, W to represent the height and width of the image, I ir ,I vis Represent the infrared image and visible light image to be fused, I f represents the fused image; The target mask image is a binary mask image obtained by annotation. During training, it guides the network to detect the salient area where the target is located. The pixel value of the target position in the mask image is 1, and the pixel value of the rest of the area is 0.

3. The method for real-time fusion of infrared and visible light images according to claim 2, characterized in that: The loss function of the network structure consists of two types of losses: similarity loss and target information loss; The similarity loss uses the structural similarity index (SSIM) to calculate the similarity between the fused image and the infrared and visible light images. Since the SSIM value is between -1 and 1, the following similarity loss function can be defined: L ssim =1-ssim(I f ,I ir )+1-ssim(I f ,I vis ) (1) The similarity loss constrains the fused image to learn the brightness, contrast, and structural similarity of infrared and visible light in order to learn the overall intensity distribution of the input image.

4. The method for real-time fusion of infrared and visible light images according to claim 3, characterized in that: The fused image should be close to the infrared image in the target area and close to the visible light image in the background area. The information contained in the fused image should be measured in two dimensions: pixel and gradient. Therefore, the target information loss should be further divided into pixel loss and gradient loss, as shown below: L sa =αL sa-pix +(1-α)L sa-grad (1) Pixel loss constrains the pixel intensity of the fused image to be consistent with the source image, while gradient loss focuses on the edge and detail information of the image. α is a balance parameter, and the mask divides the entire image into target and background areas. Therefore, pixel loss and gradient loss are performed separately in two areas.

5. The method for real-time fusion of infrared and visible light images according to claim 4, characterized in that: The pixel loss constrains the similarity between the fused image and the source image in the salient area and background area through the pixel-by-pixel absolute error (L1 norm). The specific formula is as follows: Among them, λ is the balance parameter, e is the dot product, and the loss function adopts l1-norm. The mask image avoids the interference of the visible light image in the target area through the dot product, retaining only the thermal target presented by infrared, while only the texture information of the visible light image is required in the background area; Similarly, the gradient loss also consists of the following two parts: Among them, μ is the balance parameter, ▽ represents the gradient operator, and the Sobel operator is used; In summary, the expression of the total loss function is as follows, where β is used to balance the similarity loss and the target information loss; L total =L ssim +βL sa (5)。 6. The method for real-time fusion of infrared and visible light images according to claim 5, characterized in that: The convolutional neural network (CNN) performs pixel-level fusion of images in the spatial domain. Given a pair of registered infrared and visible light images, they are input into their respective corresponding pathways in the network, and the network outputs two weight maps. The double channel stacking operation is used to address the insufficient feature extraction capability of shallow networks, and the spatial resolution of the feature map of each convolution module remains the same as that of the input image; After the input image has undergone feature extraction in one layer of convolution, it will be spliced ​​and input into the shared convolution block.

7. The method for real-time fusion of infrared and visible light images according to claim 6, characterized in that: The convolutional neural network (CNN) contains a total of 8 convolution blocks with a convolution kernel size of 3×3. The output channel of the last convolution block conv7 and conv8 is 1. In addition, the number of output channels of the remaining convolution blocks is 8. The last convolution layer is followed by a Sigmoid layer to limit the pixel value range of the weight map to [0, 1]. When inputting the network model, the RGB visible light image needs to be converted to the YCbCr color space first, then the Y channel is used to fuse the grayscale infrared image, and finally the fused image is converted back to the RGB space through inverse conversion.

8. The method for real-time fusion of infrared and visible light images according to claim 7, characterized in that: The specific formula is as follows: after weighting, normalization is performed to obtain a single-channel grayscale image; Where W ir ,W vis Two weight graphs representing the network output respectively; The mask image is only used in the training phase of the network model. The mask image is used to participate in the calculation of the loss function value, thereby guiding the model training through gradient backpropagation. The mask image is not required in the model inference phase.

9. A device for real-time fusion of infrared and visible light images, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Infrared target mask-based infrared and visible light image fusion method

    CN117611469A

  • Two-stage infrared and visible light image fusion method based on mask prior

    CN119477717A

Cited By

  • Domain generalization personnel re-identification method and system based on multi-modal fusion and structure perception enhancement

    CN121747157A