Method for fusing visible and near-infrared images based on regional complementary characteristics
Patent Information
- Application Number
- CN202310696297.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-12
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-06-12
AI Technical Summary
[0005]但是,对于实际场景下的室外夜间视频监控,现有的可见光与近红外光图像融合算法无法满足需求,主要存在以下问题:
Smart Images

Figure CN117237250B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a visible light and near-infrared light image fusion method based on regional complementary properties. Background Technology
[0002] Video surveillance, a technology that uses cameras to collect image and video data and then automatically processes, analyzes, and understands it, can save manpower and resources, enabling 24 / 7, multi-directional monitoring. It can also be used for reporting and collecting evidence of illegal activities. Video surveillance technology is widely used in security, traffic management, and industrial monitoring, and even more application scenarios are expected to emerge in the future.
[0003] Most existing intelligent video surveillance systems are designed based on visible light image information, which can meet the requirements when the visible light image quality is good. However, under different weather conditions, light intensity varies, causing instability in the grayscale and color values of individual pixels in the image, posing significant difficulties and challenges to video surveillance. For example, visible light images captured at night are of low quality and cannot clearly display objects and scenes; even with image signal processing (ISP), information loss may still occur. Therefore, video surveillance technology needs to adopt effective methods to improve image quality and recognition capabilities.
[0004] Near-infrared light is unaffected by visible light illumination conditions. Near-infrared supplemental lighting can clearly display target objects in dark environments, providing more object details and texture information. Therefore, near-infrared images can be fused with visible light images to obtain more effective information and improve image quality. Thus, near-infrared images can enhance nighttime surveillance capabilities and improve target detection and tracking performance.
[0005] However, existing visible light and near-infrared image fusion algorithms cannot meet the requirements for outdoor nighttime video surveillance in real-world scenarios, mainly due to the following problems:
[0006] 1) Information loss occurs in the fusion result. Due to the influence of scene depth, the useful information of near-infrared light images obtained by near-infrared supplementary lighting at night has different distributions indoors and outdoors. Indoors, usable near-infrared light information is distributed throughout the entire image, while outdoors, usable near-infrared light information is distributed in the foreground area. Current visible light and near-infrared light image fusion algorithms do not consider information distribution, causing useless near-infrared or visible light image information to affect the corresponding useful visible light or near-infrared light image information, resulting in information loss in the fusion result.
[0007] 2) The algorithm does not simultaneously consider noise removal from the visible light image and detail enhancement from the near-infrared image. Most existing fusion algorithms only consider using the near-infrared image to remove noise from the visible light image or to add detail information from the near-infrared image to the visible light image. Fusion algorithms that only consider image denoising result in a blurred fusion image and loss of detail information. While fusion algorithms that only consider adding detail information can supplement the texture information of the visible light image, noise is not removed. Summary of the Invention
[0008] To address some problems with existing visible light and near-infrared image fusion algorithms in nighttime video surveillance scenarios, this invention proposes a visible light and near-infrared image fusion method based on regional complementarity. The method includes: firstly, classifying the histogram of the near-infrared image using a Support Vector Machine (SVM) classifier to distinguish whether the current fusion scene is indoors or outdoors; then, decomposing the visible light and near-infrared images into texture layer and base layer images; next, generating corresponding indoor or outdoor texture fusion weight maps based on the fusion scene; and finally, fusing texture information according to the weights and combining it with the base layer information of the visible light image to obtain the fused result image. Specifically, this includes:
[0009] According to the fusion method based on regional complementarity of the present invention, useful information in the image is fused and useless information that will interfere with the result is discarded; selective information fusion is performed according to the information distribution of different scenes. In outdoor scenes, background information of visible light image and foreground information of near-infrared light image are fused. In indoor scenes, information of near-infrared light image is used as the main component, and visible light image information affected by noise is eliminated.
[0010] To simultaneously perform denoising and detail enhancement, the visible light and near-infrared light image fusion method based on regional complementary characteristics according to the present invention employs a texture-weighted fusion process. This method simultaneously completes the tasks of denoising the visible light image and fusing detailed information from the near-infrared light image. Based on the characteristic that near-infrared light images are essentially noise-free and rich in detailed information, the above tasks can be directly accomplished by replacing noise information with texture information. For useful information in the near-infrared light image, weighted fusion maps are obtained according to indoor and outdoor scenes, respectively. This allows the texture information of the near-infrared light image to replace noise information in the visible light image, directly removing noise from the visible light image and fusing detailed information from the near-infrared light image.
[0011] According to one aspect of the present invention, a method for image fusion of visible light and near-infrared light based on regional complementary properties is provided, characterized by comprising the following steps:
[0012] A) Perform indoor and outdoor scene segmentation, including determining the scene through the histogram distribution of near-infrared light images indoors and outdoors, and determining the current scene through an SVM classifier;
[0013] B) Extract high and low frequency information, including adaptive extraction of high and low frequency information from visible light and near-infrared light images through guided filtering;
[0014] C) Generate a weight map, including determining the generation method of the fusion weight map based on the current scene. For indoor scenes, the weight of the near-infrared light image is 1. For outdoor scenes, the fusion weight map is determined by low-frequency information.
[0015] D) Generating a fused image, including fusing high-frequency information from visible and near-infrared images using a fusion weight map to obtain a texture fusion result, and then combining the low-frequency information from the visible image with the fused high-frequency information to obtain the final fused image.
[0016] in:
[0017] Step B) includes:
[0018] B1) Convert the RGB visible light image to a LAB image, save the AB color components, and use the luminance component L as the object of subsequent processing. The visible light image mentioned below refers to the luminance component image.
[0019] B2) By comparing the differences between the filtered and unfiltered images, the system adaptively determines whether to continue or stop filtering, in order to adaptively extract high and low frequency components based on the noise level, including:
[0020] Extract low-frequency components as follows:
[0021] B v =GuidF n (V, V)
[0022] B n =GuidF n (N, N)
[0023] GuidF n This indicates n guided filtering operations, where V represents the luminance component of the visible light image, and N represents the near-infrared light image B. v and B n These represent the obtained low-frequency layer components,
[0024] The difference between the filtered and unfiltered images is measured using Peak Signal-to-Noise Ratio (PSNR), with the termination condition being:
[0025] p = PSNR(V) i V i-1 )>α
[0026] Where V i V represents the result of the i-th filtering iteration. i-1Let p represent the filtering result of the (i-1)th iteration. Filtering stops when the peak signal-to-noise ratio of the two results is greater than α, and α is set to 29.
[0027] B3) The difference between the original image and the base layer image, i.e., the high-frequency component information, is represented as:
[0028] D v =VB v
[0029] D n =NB n
[0030] Where V and N represent the original visible light and near-infrared light images, respectively, and B... v and B n D represents the low-frequency components of visible light and near-infrared light images, respectively. v and D n These represent the high-frequency layer components of visible light and near-infrared light images, respectively.
[0031] Step C) includes:
[0032] C1) Based on the scene segmentation, set the weight map W of the indoor near-infrared light image texture information to 1.
[0033] C2) For outdoor scenes, the low-frequency layer components of visible and near-infrared images are used to calculate the weight ratios and normalize them to obtain the fused weight map:
[0034]
[0035] Among them B n and B v Let represent the low-frequency layer components of the near-infrared and visible light images, respectively; norm represents the normalization operation; and W represents the resulting fusion weight map.
[0036] Step D) includes:
[0037] D1) Based on the obtained fusion weight map W, fuse the high-frequency information of the visible light and near-infrared light images. The fusion method is as follows:
[0038] F d =D n ·W+D v ·(1-W)
[0039] Where D n and D v Let F represent the high-frequency layer components of the near-infrared and visible light images, respectively, and W represent the previously obtained fusion weight map. d High-frequency information representing fusion,
[0040] D2) The low-frequency information of the visible light image is combined with the obtained fused high-frequency information to obtain the fusion result F. The fusion method is as follows:
[0041] F = B v +F d
[0042] Among them B v F represents the low-frequency information of a visible light image. d High-frequency information representing fusion,
[0043] D3) Combine the fusion result with the AB color components to obtain a new LAB image, and then convert the LAB image back to an RGB image to obtain the final fusion result image. Attached Figure Description
[0044] Figure 1 These are visible light and near-infrared light images of outdoor nighttime images processed by ISP.
[0045] Figure 2 These are visible light and near-infrared light images of indoor environments at night, processed by ISP.
[0046] Figure 3 This is a graph showing the filtering results obtained from different guided filtering iterations of a visible light image.
[0047] Figure 4 It consists of an outdoor visible light image, a near-infrared light image, and the resulting fusion weight map.
[0048] Figure 5 This is a flowchart of the fusion algorithm. Detailed Implementation
[0049] Figure 5 The diagram shown is a flowchart of a visible-near-infrared image fusion method based on regional complementarity characteristics according to an embodiment of the present invention. The method includes the following steps:
[0050] A) Determine the scene by the histogram distribution of near-infrared light images indoors and outdoors, and determine the current scene by an SVM classifier;
[0051] B) High and low frequency information of visible light and near-infrared light images are adaptively extracted through guided filtering;
[0052] C) Determine the fusion weight map generation method based on the current scene, including: setting the weight of the near-infrared light image to 1 for indoor scenes, and determining the fusion weight map based on low-frequency information for outdoor scenes;
[0053] D) The high-frequency information of the visible light and near-infrared light images is fused using the fusion weight map to obtain the texture fusion result. Then, the low-frequency information of the visible light image is combined with the fused high-frequency information to obtain the final fused image.
[0054] Step A) includes dividing the indoor and outdoor scenes.
[0055] This invention determines the relevant regions of useful near-infrared information based on the regional complementarity characteristics of visible light and near-infrared light images. Indoor and outdoor visible light and near-infrared light images exhibit different regional complementarity characteristics. Figure 1 (a) and (b) represent outdoor visible light and near-infrared light images, respectively. The foreground region of the near-infrared image contains useful information. Due to the influence of supplementary lighting and object reflectivity, the brightness of this region in the near-infrared image may be higher than that in the visible light image. Conversely, the background region lacks sufficient lighting, and its brightness is much lower than the background brightness of the visible light image. This brightness difference can effectively extract the respective high-reflectivity regions. For indoor scenes, Figure 2 (a) and (b) represent indoor visible light and near-infrared light images, respectively. The near-infrared light image contains the information of the entire image, while the corresponding visible light image is affected by noise and suffers severe information loss. Based on the brightness distribution of indoor and outdoor near-infrared light images, the histograms of the near-infrared light images in the two scenes are very different, thus effectively distinguishing different scenes.
[0056] SVM classifiers can be used for classification or regression tasks. By finding a hyperplane that maximizes the separation of different classes in the training data, they can determine which side of the hyperplane a new data point falls on, thus completing the classification. Therefore, an SVM classifier can be trained on the histograms of existing indoor and outdoor near-infrared light images to learn the differences in histogram distribution under different scenes, thereby determining the scene of a new near-infrared light image.
[0057] Step B) includes the extraction of high and low frequency information.
[0058] Visible light and near-infrared images contain different high- and low-frequency information, which needs to be selected and utilized. The low-frequency information of visible light images reflects the natural brightness of the scene, while the low-frequency information of near-infrared images is affected by supplementary lighting and object texture. Therefore, the low-frequency information of near-infrared images will affect the fusion result, causing changes in the brightness of the fused result, and further leading to color shifts in the fused image. On the other hand, the high-frequency information of visible light images is easily affected by noise, while near-infrared images obtained through supplementary lighting have less noise, and their high-frequency information can reflect the true texture information of objects.
[0059] Therefore, the availability of high- and low-frequency information differs between visible light and near-infrared images, necessitating their separation. The degree of noise impact also affects the difficulty of separating high- and low-frequency components; thus, it is necessary to determine the degree of noise impact on the visible light image and perform multiple guided filters on both the visible light and near-infrared images. If there is little or no noise impact, the number of filters is 1; the greater the noise impact, the more filters are required. Figure 3 The image shows the results of four filtering passes, demonstrating that four filters are required to completely remove noise. To adaptively extract high and low frequency components based on noise level, the process can be modified by comparing the differences between the filtered and unfiltered images to determine when to stop. Higher noise levels result in more noise removal and a greater difference between the filtered and unfiltered images. A smaller difference indicates that the high and low frequency components have been separated. The low-frequency component is extracted as follows:
[0060] B v =GuidF n (V, V)
[0061] B n =GuidF n (N, N)
[0062] GuidF n This indicates n guided filtering operations, where V represents the luminance component of the visible light image, and N represents the near-infrared light image B. v and B n These represent the obtained low-frequency layer components.
[0063] The difference between the filtered and unfiltered images is measured using Peak Signal-to-Noise Ratio (PSNR), with the termination condition being:
[0064] p = PSNR(V) i V i-1 )>α
[0065] Where V i V represents the result of the i-th filtering iteration. i-1 This represents the filtering result of the (i-1)th iteration. When the peak signal-to-noise ratio p of the two results is greater than α, filtering is stopped and α is set to 29.
[0066] The difference between the original image and the base layer image is the high-frequency component, characterized as follows:
[0067] D v =VB v
[0068] D n =NB n
[0069] Where V and N represent the original visible light and near-infrared light images, respectively, and B... v and B n D represents the low-frequency components of visible light and near-infrared light images, respectively. v and D n These represent the high-frequency layer components of visible light and near-infrared light images, respectively.
[0070] Step C) includes weight graph generation.
[0071] Based on the complementary characteristics of indoor and outdoor image regions, different fusion weight maps need to be generated. Of the separated low-frequency information, only the low-frequency information from the visible light image needs to be fused, while the high-frequency information needs to be determined according to different scenes. For indoor scenes, the near-infrared image contains all the usable information of the entire image, while the high-frequency information of the visible light image is affected by noise; therefore, the fusion weight for the near-infrared image is 1, and for the visible light image, it is 0. Outdoor scenes have greater depth of field, and near-infrared light often only captures foreground information, resulting in high foreground brightness and low background brightness. Visible light images, on the other hand, have less noise outdoors, and after ISP processing, the overall brightness of the image is relatively high. Utilizing the regional brightness characteristics represented by low-frequency information, a fusion weight map can be generated. The weight map generated for outdoor scenes is shown below. Figure 4 As shown in (c), the formula for calculating the weighted graph is as follows:
[0072]
[0073] Among them B n and B v These represent the low-frequency information of near-infrared and visible light images, respectively, and norm represents the normalization operation.
[0074] Step D) includes fused image generation.
[0075] Based on the obtained high and low frequency information and the fusion weight map W, the high frequency information is first fused according to the weights, and then the high frequency information is combined with the low frequency information to obtain the fused image. First, the high frequency information of the visible light and near-infrared light images is fused, and the fusion method is as follows:
[0076] F d =D n ·W+D v ·(1-W)
[0077] Where D n and D v Let F represent the high-frequency layer components of the near-infrared and visible light images, respectively, and W represent the previously obtained fusion weight map. d This represents the high-frequency information obtained through fusion. The low-frequency information from the visible light image is combined with the fused high-frequency information to obtain the fusion result. The fusion method is as follows:
[0078] F = B v +F d
[0079] Among them B v F represents the low-frequency information of a visible light image. d This represents the high-frequency information of the fusion process. The F and AB color components are combined to obtain a new LAB image. Converting this LAB image back to an RGB image yields the final fused image. The overall fusion process is as follows: Figure 5 As shown.
Claims
1. A method for image fusion of visible light and near-infrared light based on regional complementary properties, characterized in that... Includes the following steps: A) Perform indoor and outdoor scene segmentation, including determining the scene through the histogram distribution of near-infrared light images indoors and outdoors, and determining the current scene through an SVM classifier; B) Extract high and low frequency information, including adaptive extraction of high and low frequency information of visible light images and near-infrared light images through guided filtering; C) Generate a weight map, including determining the generation method of the fusion weight map based on the current scene. For indoor scenes, the weight of the near-infrared light image is 1. For outdoor scenes, the fusion weight map is determined by low-frequency information. D) Generating a fused image, including fusing high-frequency information from the visible light image and the near-infrared light image using a fusion weight map to obtain a texture fusion result, and then combining the low-frequency information from the visible light image with the fused high-frequency information to obtain the final fused image. in: Step B includes: B1) Convert the RGB visible light image to a LAB image, save the AB color components, and use the luminance component L as the object of subsequent processing. The visible light image mentioned below refers to the luminance component image. B2) By comparing the differences between the filtered and unfiltered images, the system adaptively determines whether to continue or stop filtering, in order to adaptively extract high and low frequency components based on the noise level, including: Extract low-frequency components as follows: in Indicates to proceed Secondary guided filtering. Represents the luminance component of a visible light image. This represents the brightness component of a near-infrared light image. This represents the low-frequency components of a visible light image. These represent the low-frequency components of the near-infrared light image. The difference between the filtered and unfiltered images is measured using Peak Signal-to-Noise Ratio (PSNR), with the termination condition being: in Indicates the first The result of the second filtering step Indicates the first The result of the second filtering step B3) The difference between the original image and the base layer image, i.e., the high-frequency component information, is represented as: in Represents the high-frequency components of a visible light image. Represents the high-frequency components of near-infrared light images. Step C includes: C1) Based on the segmented scenes, weight the texture information of the indoor near-infrared light image. Set to 1, C2) For outdoor scenes, the low-frequency components of visible and near-infrared light images are used to calculate the weight ratios and normalize them to obtain a fused weight map: , in This represents the normalization operation, and W represents the resulting fused weight graph. Step D includes: D1) According to the obtained fusion weight map The high-frequency information from visible light and near-infrared light images is fused using the following method: in High-frequency information representing fusion, D2) The low-frequency information of the visible light image is combined with the obtained fused high-frequency information to obtain the fusion result F. The fusion method is as follows: D3) Combine the F and AB color components to obtain a new LAB image, and convert the LAB image back to an RGB image to obtain the final fusion result image.
2. The visible light and near-infrared light image fusion method based on regional complementary characteristics according to claim 1, characterized in that... It was set to 29.
Citation Information
Patent Citations
Regional characteristic-based fusion method of infrared and visible light images
CN106204509A
Infrared and visible light image fusion method
CN113628151A