Infrared and visible light image fusion method and system based on LSRGF and region detection
Through least squares optimized rolling guidance filtering and multi-scale significant area detection methods, infrared and visible light images are decomposed and fused, solving the problems of insufficient spatial consistency and lack of fusion details in the decomposition process in the prior art, and achieving high-quality image fusion.
Patent Information
- Application Number
- CN202510086953.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
AI Technical Summary
In the existing infrared and visible image fusion technology, there are insufficient considerations for spatial consistency of the decomposition process, lack of fusion details, and insufficient consideration of image details correlation.
The visible light and infrared images are decomposed by least squares optimization, and the basic layer images and detailed images at each level are obtained respectively. Then, a multi-scale significant area detection method is constructed to fuse the base layer images and fuse the detailed images using an improved pulse-coupled neural network, ultimately fusing the two to obtain the final fusion result.
Effectively retaining the spatial consistency of the image, enhancing the retention of fusion details and image details relevance, and improving the image fusion quality.
Smart Images

Figure CN120070198A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of computer vision and image processing, and in particular, to an infrared and visible light image fusion method and system based on LSRGF and region detection. Background Art
[0002] Currently, the fusion techniques for infrared and visible light are divided into two categories: traditional methods and deep learning-based methods. Traditional methods use image processing means such as multi-scale transformation, sparse representation, filtering, and human visual mechanism modeling to strive to fuse the information in infrared images and visible light images while maintaining the visual effects of the images. Deep learning methods use encoding networks to extract image feature information and are used to improve the image feature fusion strategy, improving the drawbacks of manually set fusion strategies. However, the fusion effect depends on the quality of the training data set, and at the same time, the requirements for labeled data are relatively high. Compared with deep learning methods, traditional methods have the characteristics of small computational amount, high registration accuracy, and good real-time performance. Among them, the multi-scale transformation fusion method is one of the methods currently recognized by scholars and is widely used because of its simple algorithm and good performance. It includes methods such as Dual-Tree Complex Wavelet Transform (DTCWT), NonSubsampled Contourlet Transform (NSCT), and Nonsubsampled Contourlet Transform-based Super-Resolution (NSCT-SR). Due to the characteristics of filtering technology such as noise smoothing and edge preservation, in recent years, many scholars have begun to successfully apply filtering to image fusion.
[0003] In 2013, Li et al. proposed an image fusion algorithm with guided filtering (GFF), which decomposed the image using mean filtering, combined Laplacian filtering and Gaussian filtering to obtain a saliency map, and for the first time used guided filtering to construct a weight map, solving the problem of misalignment of the target edges in the initial weight map. Bavirisetti et al. borrowed the advantages of the GFF method and proposed a multi-scale guided fusion method (MGF), which greatly reduced the algorithm complexity in transferring edge structures at different scales. In 2014, Zhang et al. proposed a Rolling Guidance Filter (RGF) framework. Different from bilateral filtering and guided filtering, this filter can effectively eliminate halos and restore the edges of large-scale targets while smoothing small-scale targets.
[0004] In 2017, Ma et al. combined rolling guidance filtering and Gaussian filtering to perform multi-scale decomposition on the source image. However, RGF rarely considers spatial consistency and is prone to artifacts such as gradient reversal and halos. In the research on the mechanism of human vision, Eckhorn et al. established a pulse coupled neural network (PCNN) based on the biological phenomenon of the firing oscillation of neurons in the mammalian cerebral cortex. Since the working mechanism of the PCNN model has many similarities with that of human visual cortex neurons, many scholars use it as one of the means to study image processing methods based on the mechanism of human vision. Jason proposed a more effective simplified pulse network, and the results it produces are similar to those of the original PCNN. Subsequently, some scholars made a series of improvements to the PCNN to make it more suitable for the field of image processing. However, it was found that when using the traditional PCNN model for fusion, if PCNN is used as the fusion criterion for both the base layer and the detail layer, the high-frequency information in the low-frequency sub-images cannot be well retained. PCNN can effectively obtain local image information due to the non-linear synchronous pulse coupling firing characteristics of its neurons, so it is also widely used in the fields of image fusion, segmentation, and recognition. However, the PCNN has many parameters and a large amount of computation. Summary of the Invention
[0005] An embodiment of the present application provides an infrared and visible light image fusion method and system based on LSRGF and region detection, which uses the direction idea of first decomposing and then fusing to solve the problems of insufficient consideration of spatial consistency in the decomposition process, lack of fusion details, and insufficient consideration of the correlation degree of image details.
[0006] An embodiment of the present application provides an infrared and visible light image fusion method based on LSRGF and region detection, including:
[0007] Using the least squares optimized rolling guidance filtering method to decompose visible light and infrared images to respectively obtain the base layer image and detail images at all levels;
[0008] For the obtained base layer image, construct a multi-scale saliency region detection method and obtain the first fusion image based on the fusion rule;
[0009] Using an improved pulse coupled neural network to fuse the obtained detail images to obtain the second fusion image;
[0010] Fuse the first fusion image and the second fusion image to obtain the fusion result.
[0011] An embodiment of the present application also provides an infrared and visible light image fusion system for LSRGF and region detection, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the aforementioned infrared and visible light image fusion method for LSRGF and region detection are implemented.
[0012] In view of the current infrared and visible light image fusion technology, the embodiment of the present application uses the directional idea of first decomposing and then fusing to solve the problems of insufficient consideration of spatial consistency in the decomposition process, lack of fusion details, and insufficient consideration of the image detail correlation.
[0013] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0015] Figure 1 It is a schematic diagram of the basic process of the infrared and visible light image fusion method for LSRGF and region detection according to the embodiment of the present application;
[0016] Figure 2 It is a schematic diagram of the process of the infrared and visible light image fusion method for LSRGF and region detection according to the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Hereinafter, the exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0018] The first embodiment of the present application provides an infrared and visible light image fusion method for LSRGF and region detection, as Figure 1 、 Figure 2 shown, including the following steps:
[0019] In step S101, the least squares optimized rolling guidance filtering method is used to decompose visible light and infrared images to respectively obtain the base layer images and the detail images at each level. In a specific example, the visible light and infrared images are respectively decomposed to obtain the base layer images and the detail images at each level of the visible light and the base layer images and the detail images of the infrared images.
[0020] In step S102, for the obtained base layer images (the base layer images of visible light and the base layer images of infrared images), a multi-scale salient region detection method (Multi-scale Salient Region Detection, MSRD) is constructed, and a first fused image is obtained based on the fusion rule to preserve the image details at multiple scales.
[0021] In step S103, an improved pulse coupled neural network is used to fuse the obtained detail images (the detail images at each level of the visible light image and the detail images of the infrared image) to obtain a second fused image, which can reduce the running time while fully utilizing the correlation between images.
[0022] In step S104, the first fused image and the second fused image are fused to obtain a fusion result.
[0023] The least squares optimized rolling guidance filtering (LSRGF) image decomposition method. The rolling guidance filtering (RGF) has the characteristics of scale perception and edge protection at the same time, and its iterative convergence speed is relatively fast. Using it to decompose the source image can effectively retain the detailed edge information of the target. In some embodiments, the least squares optimized rolling guidance filtering (LSRGF) method is used to decompose visible light and infrared images. This filter includes two main steps: smoothing small-scale structures and restoring large-scale structure edges.
[0024] Specifically, it includes:
[0025] For visible light and infrared images, the small structures are removed by using Gaussian filtering respectively to filter the detailed information and edge structures of the images, where the Gaussian filter satisfies:
[0026]
[0027]
[0028] where I represents the input image, G represents the output image, p and q represent the corresponding pixel coordinates, σ s is the standard deviation, and N(p) represents the filtering window centered on the pixel point p.
[0029] For the image after Gaussian filtering, the edge information is iteratively restored using a bilateral filtering kernel, and the iterative restoration process satisfies:
[0030]
[0031] Among them, J t+1 represents the result of the t-th iteration, K p is used for normalization, σ s and σ r control the spatial range weight and the intensity difference range weight respectively. In the first iteration, J 1 is the output image G after Gaussian filtering in the first step, and I is the original image. During the iterative operation process, given the J t generated in the previous iteration of I, the result value J t+1 of the t-th iteration can be obtained.
[0032] However, the RGF method rarely considers spatial consistency, is prone to artifacts such as gradient reversal and halos, and affects the further processing of the image. In some embodiments, using the least squares optimization rolling guidance filtering method, the decomposition of visible light and infrared images further includes:
[0033] Optimizing the rolling guidance filtering method using an orthogonal direction least squares optimization model, and the optimization process satisfies:
[0034]
[0035] Among them, I out (i, j) and I in (i, j) represent the input image and the output image, and represent the gradients of the image in the x-axis and y-axis directions, and f RGF represents the rolling guidance filtering process.
[0036] The least squares optimization model (LS) consists of a data term and a regularization term. The data term is used to ensure the similarity between the input image and the output image, thereby retaining edge details. The regularization term is used to control the global properties of the output image, and the embodiments of the present application balance the proportion between the two through an optimized solution.
[0037] In some embodiments, using the least squares optimization rolling guidance filtering method, the decomposition of visible light and infrared images further includes:
[0038] For the optimized rolling guidance filtering process, the fast Fourier transform is used for solution, satisfying:
[0039]
[0040] Among them, represents the Fast Fourier Transform.
[0041] In some embodiments, when using the least squares optimized rolling guidance filtering method to decompose visible light and infrared images, it further includes:
[0042] Let the operator of the least squares optimized rolling guidance filtering be U = LSRGF(I, σ s , σ r , T), and perform multi-scale decomposition on the image using LSRGF. As Figure 2 shown, the base layer image and the detail image are obtained, and the decomposition method satisfies:
[0043]
[0044] D j = U j-1 - U j , j = 1, …, N - 1 (8)
[0045]
[0046] D j = U j-1 - U j , j = N (10)
[0047] where N is the number of decomposition layers, U 0 is the initial original image, U j is the rolling guidance filtering output image of the j-th layer, D j is the detail image of the j-th layer, U N is the new filtered image generated by combining Gaussian filtering in the last layer, U N serves as the base layer image, and D N is the calculated detail image.
[0048] This application addresses the current problem of infrared and visible light image fusion. Continuing the direction of first decomposing and then fusing, it solves the problems in the existing technology, such as insufficient consideration of spatial consistency in the decomposition process, lack of fusion details, and insufficient consideration of the image detail correlation. It proposes the Least Squares Optimization Rolling Guidance Filter (LSRGF) to decompose visible light and infrared images, preserving the spatial consistency of the images, and respectively obtaining the base layer image and the detail images at each level. For the base layer image, a Multi-Scale Region Detection Method (MSRD) is constructed to extract the significant regions of the image at multiple scales, and the fused image is obtained based on the fusion rule. For the detail layer images, an improved Pulse Coupled Neural Network (DPCNN) is used for fusion, which can reduce the running time while fully utilizing the correlation between images. Finally, the sub-fused images are superimposed to obtain the final fused image. Compared with classical algorithms for multi-scale transformation and deep learning image fusion, such as NSCT_SR, CNN, DCTWT, and IFEVIP, the method of this application is superior in the evaluation of image fusion quality.
[0049] An embodiment of this application also provides an infrared and visible light image fusion system based on LSRGF and region detection, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, it implements the steps of the infrared and visible light image fusion method based on LSRGF and region detection as described above.
[0050] It should be noted that in each embodiment of this application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0051] The serial numbers of the embodiments of this application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0052] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of this application.
[0053] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims. All of these are within the protection scope of the present application.
Claims
1. A method for fusion of infrared and visible light images based on LSRGF and region detection, characterized in that: include: The rolling guided filtering method with least square optimization is used to decompose the visible light and infrared images to obtain the base layer image and the detailed images of each level respectively. For the obtained base layer image, a multi-scale salient region detection method is constructed, and a first fused image is obtained based on a fusion rule; The obtained detail images are fused by using an improved pulse coupled neural network to obtain a second fused image; The first fused image and the second fused image are fused to obtain a fusion result.
2. The infrared and visible light image fusion method for LSRGF and regional detection as claimed in claim 1, characterized in that: The rolling guide filtering method with least squares optimization is used to decompose the visible and infrared images including: For visible light and infrared images, Gaussian filtering is used to remove small structures to filter the image details and edge structures. The Gaussian filter satisfies: Among them, I represents the input image, G represents the output image, p and q represent the corresponding image pixel coordinates, σ s is the standard deviation, N(p) represents the filter window centered at pixel p; For the image after Gaussian filtering, the edge information is iteratively restored using the bilateral filter kernel. The iterative recovery process satisfies: Among them, J t+1 represents the result of the tth iteration, K p For normalization, σ s and σ r Controls the spatial range weight and intensity difference range weight respectively.
3. The infrared and visible light image fusion method for LSRGF and regional detection as claimed in claim 2, characterized in that: The rolling guide filtering method using least squares optimization is used to decompose the visible and infrared images and also includes: The rolling guidance filtering method is optimized using the least squares optimization model in the orthogonal direction, and the optimization process satisfies: Among them, I out (i, j) and I in (i,j) represents the input image and the output image, and Represents the gradient of the image in the x-axis and y-axis directions, f RGF Represents rolling guided filtering processing.
4. The infrared and visible light image fusion method for LSRGF and regional detection as claimed in claim 3, characterized in that: The rolling guide filtering method using least squares optimization is used to decompose the visible and infrared images and also includes: For the optimized rolling guidance filtering process, the fast Fourier transform is used to solve it, satisfying: in, stands for Fast Fourier Transform.
5. The infrared and visible light image fusion method for LSRGF and regional detection as claimed in claim 4, characterized in that: The rolling guide filtering method using least squares optimization is used to decompose the visible and infrared images and also includes: Let the least squares optimized rolling guide filter operator be U = LSRGF (I, σ s ,σ r ,T), LSRGF is used to perform multi-scale decomposition on the image to obtain the base layer image and detail image. The decomposition method satisfies: D j =U j-1 -U j ,j=1,…,N-1 D j =U j-1 -U j ,j=N Among them, N is the number of decomposition levels, U 0 is the initial original image, U j is the rolling guided filter output image of the jth layer, D j is the detail image of the jth layer, U N For the last layer, a new filtered image is generated by combining Gaussian filtering, U N As the base layer image, D N is the calculated detail image.
6. An infrared and visible light image fusion system for LSRGF and regional detection, characterized in that: The method comprises a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the infrared and visible light image fusion method for LSRGF and regional detection as described in any one of claims 1 to 5 are implemented.