Image fusion method and device based on adaptive noise processing and visual enhancement

By employing an image fusion method that incorporates adaptive noise processing and visual enhancement, the problems of edge blurring and noise interference in image fusion are solved, achieving high-quality image fusion results and enhancing scene adaptability and readability of the fused image.

CN120876248APending Publication Date: 2025-10-31BEIJING AERONAUTIC SCI & TECH RES INST OF COMAC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510935896.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing image fusion techniques, mean filtering leads to blurred image edges and details, visual saliency strategies struggle to accurately capture subtle structures, and simple weighting methods cannot balance the contrast and brightness differences between different source images, resulting in a decrease in the quality and readability of the fused image.

Method used

An image fusion method employing adaptive noise processing and visual enhancement is proposed. By constructing detail and base layer images through multi-scale decomposition, and combining noise estimation and saliency analysis, adaptive weighting and enhancement techniques are used for fusion.

Benefits of technology

While ensuring a short fusion delay, it effectively overcomes detail blurring and noise interference, improves the fusion effect, conforms to human vision, and enhances scene adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876248A_ABST
    Figure CN120876248A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image fusion method based on adaptive noise processing and visual enhancement, and the method comprises the steps: constructing the multi-scale representation of any image in a group of images to be fused, and determining a detail layer image and a basic layer image of the image through the multi-scale representation of the image; performing noise estimation on the detail layer image of each image in the group of images, determining a detail layer fusion strategy of the group of images according to a noise estimation result, and fusing the same-layer detail layer images of each image in the group of images through the determined detail layer fusion strategy; determining the saliency of each image in the group of images, and fusing the basic layer images of each image in the group of images according to the saliency of each image in the group of images; and obtaining a fused image of the group of images according to the detail layer fusion result and the basic layer fusion result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image fusion method and apparatus based on adaptive noise processing and visual enhancement. Background Technology

[0002] Image fusion technology has been applied in many fields. For example, in cockpit infrared visual enhancement systems, image fusion is key to improving pilots' visual perception in complex environments. Image fusion typically combines images from multiple bands, including visible light, long-wave infrared, and short-wave infrared, each with its own characteristics. Visible light images can present rich details and high resolution under good lighting conditions, but their information acquisition capability drops significantly in low-light or inclement weather scenarios. Infrared images, while providing target information in low-light and smoke environments, have relatively low resolution and often contain coarse-scale structural information and noise, lacking fine-scale detail. Effectively integrating these images with different characteristics to obtain a high-quality fused image has become a core challenge in the field of image fusion.

[0003] Currently, image fusion algorithms generally consist of several key stages: detail layer-base layer decomposition, detail layer fusion, base layer fusion, and image reconstruction. Each stage presents a series of technical solutions and corresponding problems.

[0004] 1. Multi-scale decomposition often employs mean filtering. Mean filtering is based on the simple principle of pixel averaging, calculating the average value of each pixel's neighboring pixels to update that pixel's value. This method is fast, but it doesn't consider the local feature differences between pixels. When processing an image, it treats all pixels equally, causing excessive blurring of image edges and details during smoothing, resulting in reduced contrast. This makes it difficult for the decomposed image to accurately retain the key information of the original image, severely impacting subsequent image fusion and providing users with insufficiently clear and accurate visual information.

[0005] 2. Detail-level fusion strategies often employ visual saliency, first applying mean filtering and median filtering to the source images, then subtracting the two results and taking the absolute value to determine visual saliency. However, mean filtering itself blurs image edges and details, and subsequent calculations further exacerbate the loss of detail information. This makes it difficult for the algorithm to accurately capture subtle structures and important features in the image. During the fusion process, visually significant details may be misclassified as noise and weakened or removed, significantly reducing the quality and readability of the fused image and hindering the user's rapid understanding of the image content.

[0006] 3. Basic layer fusion strategies often employ simple weighting methods, linearly combining pixel values ​​only. In practical applications, image content and features are extremely complex, and simple weighting cannot consider the local structure and detailed features of the image. When processing the boundaries of different source images, it is difficult to effectively handle differences, easily producing unnatural transition regions in the basic layer of the fused image, resulting in obvious stitching marks or color abrupt changes. Furthermore, this method cannot adaptively adjust according to the characteristics of different regions of the image, making it difficult to balance the contrast and brightness differences between different source images. If the weights are not set reasonably, the fused image may be too bright, too dark, or have distorted contrast, severely affecting the readability and analyzability of the image and making visual judgment difficult for users. Summary of the Invention

[0007] This specification provides an image fusion method and apparatus based on adaptive noise processing and visual enhancement to solve the technical problem of how to improve the image fusion effect.

[0008] To address the aforementioned technical problems, the embodiments in this specification provide the following technical solutions:

[0009] This specification provides an image fusion method based on adaptive noise processing and visual enhancement, the method comprising:

[0010] For any image in a set of images to be fused, construct a multi-scale representation of the image, and use the multi-scale representation of the image to determine the detail layer image and the base layer image of the image;

[0011] Noise is estimated for the detail layer images of each image in the image set. Based on the noise estimation results, a detail layer fusion strategy for the image set is determined. The same level detail layer images of each image in the image set are then fused using the determined detail layer fusion strategy. The saliency of each image in the image set is determined. Based on the saliency of each image in the image set, the base layer images of each image in the image set are then fused.

[0012] The fused image of this set of images is obtained based on the fusion results of the detail layer and the basic layer.

[0013] Preferably, a multi-scale representation of the image is constructed, and the detail layer image and the base layer image of the image are determined using the multi-scale representation, including:

[0014] Perform a Gaussian filter operation on the image and construct the corresponding Gaussian pyramid.

[0015] For each non-top layer image of the Gaussian pyramid, the predicted image of the previous layer is subtracted from the image of that layer, and the result is used as the detail layer image; and the top layer image of the Gaussian pyramid is used as the base layer image.

[0016] For any given image layer, the predicted image for that layer is the image obtained by upsampling the image of that layer and then performing a Gaussian convolution.

[0017] Preferably, noise estimation is performed on the detail layer images of each image in the set of images, including:

[0018] For any image in the set, determine the significant difference ratio corresponding to each layer of the multi-scale representation of the image, and estimate the noise of the detail layer image based on the significant difference ratio corresponding to each layer.

[0019] Specifically, for any layer of the multi-scale representation of the image, determining the significant difference ratio corresponding to that layer includes:

[0020] Filter the image of this layer to obtain the filtered image corresponding to this layer;

[0021] Calculate the absolute difference between the image of this layer and the corresponding filtered image, obtain the significant difference portion based on the absolute difference, and determine the significant difference ratio corresponding to this layer image based on the significant difference portion.

[0022] As a preferred approach, the detail layer fusion strategy for this set of images, determined based on the noise estimation results, includes:

[0023] For any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images does not exceed the first condition, then the detail layer fusion strategy of the pair of detail layer images is a fusion strategy based on the maximum absolute value.

[0024] As a preferred approach, the detail layer fusion strategy for this set of images, determined based on the noise estimation results, includes:

[0025] For any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images exceeds the second condition, the detail layer fusion strategy of the pair of detail layer images includes: firstly, using a fusion strategy based on the maximum absolute value to perform preliminary fusion, and then processing the noise of the preliminary fusion result to obtain the target fusion result.

[0026] Preferably, the preliminary fusion result is subjected to noise processing to obtain the target fusion result, including:

[0027] The spatial adaptive weights are calculated based on one of the detail layer images in the pair. The spatial adaptive matrix is ​​then calculated based on the spatial adaptive weights. Finally, the target fusion result is calculated based on the spatial adaptive matrix and the other detail layer image in the pair.

[0028] Preferably, determining the saliency of each image in this set of images includes:

[0029] For any image in the set, the saliency of the image is determined by the contrast between a single pixel in that image and the other pixels in that image.

[0030] Preferably, the fusion of the base layer images of each image in the set of images based on the saliency of each image includes:

[0031] For a pair of base layer images to be fused, the base layer fusion weights are determined based on the saliency of each base layer image, and the base layer images are fused using the base layer fusion weights.

[0032] Preferably, the set of images to be fused includes visible light images, shortwave infrared images, and images in other bands;

[0033] The fusion of the same level detail images in this image set includes:

[0034] First, perform image fusion of the same layer detail layer of the visible light image and the shortwave infrared image, and then perform a fusion weight enhancement operation during the image fusion of the same layer detail layer of the visible light image and the shortwave infrared image;

[0035] And / or,

[0036] The set of images to be fused includes visible light images, short-wave infrared images, and mid-to-long-wave infrared images;

[0037] The fusion of the base layer images of each image in this set of images includes:

[0038] First, the basic layer images of the visible light image and the short-wave infrared image are fused. Then, the fusion result of the basic layer images of the visible light image and the short-wave infrared image is fused with the basic layer image of the mid- and long-wave infrared image. Image enhancement operations are performed during the fusion process of the basic layer images of the visible light image and the short-wave infrared image, and during the fusion process of the basic layer images of the visible light image and the short-wave infrared image and the mid- and long-wave infrared image.

[0039] This specification provides an image fusion apparatus based on adaptive noise processing and visual enhancement, the apparatus comprising:

[0040] The image decomposition module is used to construct a multi-scale representation of any image in a set of images to be fused, and to use the multi-scale representation to determine the detail layer image and the base layer image of the image.

[0041] The detail layer and base layer fusion module is used to estimate the noise of the detail layer images of each image in the group of images, determine the detail layer fusion strategy of the group of images based on the noise estimation results, and fuse the detail layer images of the same layer of each image in the group of images according to the determined detail layer fusion strategy; determine the saliency of each image in the group of images, and fuse the base layer images of each image in the group of images according to the saliency of each image in the group of images.

[0042] The image reconstruction module is used to obtain a fused image of the set of images based on the detail layer fusion result and the base layer fusion result.

[0043] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:

[0044] By employing multi-scale decomposition, an innovative fusion strategy for detail and base layers is developed. Adaptive noise processing and original image enhancement are introduced, and the real-time performance of image fusion is optimized. This approach effectively overcomes fusion problems such as detail blurring, halo, and noise interference while ensuring a short fusion delay. It boasts advantages such as good fusion effect, better conformity to human vision, and strong scene adaptability. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings used in the description of the embodiments of this specification or the prior art will be briefly described below. Obviously, the drawings used in some embodiments of this application are only described below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of the image fusion framework in the first embodiment of this specification.

[0047] Figure 2 This is a schematic diagram of the image fusion effect in the first embodiment of this specification. Detailed Implementation

[0048] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments involved in the specific implementation are only a part of the embodiments of this application, and not all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments in the specific implementation without creative effort should fall within the protection scope of this application.

[0049] The first embodiment of this specification (hereinafter referred to as "Embodiment 1") provides an image fusion method based on adaptive noise processing and visual enhancement. The execution subject of Embodiment 1 includes, but is not limited to, a terminal, a server, an operating system, or an application; that is, the execution subject can be diverse and can be set, used, or changed as needed. Alternatively, a third-party application can assist the execution subject in executing Embodiment 1. For example, a server can execute the image fusion method based on adaptive noise processing and visual enhancement in Embodiment 1, and a corresponding application can be installed on a terminal (which may be held by a user). Data transmission can occur between the terminal or application and the server, thereby assisting the server in executing the image fusion method based on adaptive noise processing and visual enhancement in Embodiment 1.

[0050] refer to Figure 1 The image fusion method based on adaptive noise processing and visual enhancement provided in Example 1 includes:

[0051] S101: For any image in the set of images to be fused, construct a multi-scale representation of the image, and use the multi-scale representation of the image to determine the detail layer image and the base layer image of the image;

[0052] In Example 1, when multiple images need to be fused, these multiple images can be used as a set of images to be fused. Typically, the set of images to be fused includes images in multiple wavelength bands, such as visible light images, short-wave infrared images, and long-wave infrared images (or mid-to-long-wave infrared images).

[0053] For any image in a set of images to be fused, a multi-scale representation of the image can be constructed, and this multi-scale representation can be used to determine the detail layer image (also called the detail layer) and the base layer image (also called the base layer or foundation layer). Constructing the multi-scale representation and using it to determine the detail and base layers can include: performing a Gaussian filter on the image to construct a corresponding Gaussian pyramid; for each non-top layer image in the Gaussian pyramid, subtracting the predicted image of the previous layer from that layer image, and using the result as the detail layer image; and using the top layer image of the Gaussian pyramid as the base layer image; wherein, for any layer image, the predicted image is the image obtained by upsampling and Gaussian convolving the layer image.

[0054] In summary, Example 1 employs a multi-scale decomposition method based on an improved image pyramid, constructing a multi-scale representation based on a Gaussian pyramid to obtain the base layer image and detail layer image of the image.

[0055] The following section further explains how to obtain the detail layer image and the base layer image:

[0056] 1. Constructing the Gaussian Pyramid

[0057] For any image in the set of images to be fused (this image can also be called the input image. Example 1 can run on the system itself or as an application, accepting image input), a Gaussian filtering operation is performed on that image, for example, using a two-dimensional Gaussian function as the filter. The image is then convolved with the filter to obtain or output the filtered image. Afterwards, the latest obtained or output filtered image is used as the input for the next filtering, and the filtering process is repeated to continue obtaining or outputting new filtered images.

[0058] The above filtering process can be repeated multiple times, and the specific number of repetitions can be set arbitrarily. Furthermore, during each filtering or iteration, the size of the filtered output image is halved, and downsampling is achieved through alternating row and column sampling.

[0059] Through continuous iteration, the Gaussian pyramid G corresponding to this image is constructed. i (as a multi-scale representation of the image), where i is the number of decomposition layers.

[0060] 2. Obtain the detail layer and the basic layer

[0061] The detail layer of an image is obtained by subtracting the predicted image (i.e., the predicted image of the previous layer) from each layer of the Gaussian pyramid (excluding the top layer image). This process can be represented as follows:

[0062] L i+1 =G i -UP(G i+1 ), i = 0, 1, ..., N-1;

[0063] Among them, G i Let Gi be the image of the i-th layer of the Gaussian pyramid, and G0 be the image of the bottom layer of the Gaussian pyramid. n This is the top layer image of the Gaussian pyramid; UP(·) represents the upsampling operation; L i This represents the i-th layer image after multi-scale decomposition and computation; N represents the number of decomposition layers.

[0064] In Example 1, the intermediate layer images L1, L2, ..., Ln obtained through the above calculations are used as the detail layer images of the image, and the top layer image of the Gaussian pyramid (i.e., the top layer filtered image) is used as the base layer image of the image. In this way, the detail layer images and the base layer images of the image are obtained.

[0065] Through the above, the detail layer retains fine-scale details and texture features, with the feature scale gradually increasing, while the base layer contains coarse-scale structural features that affect the contrast and overall appearance of the final fused image.

[0066] Each image in a set of images to be fused can have its detail layer image and base layer image obtained using the above method.

[0067] S103: Perform noise estimation on the detail layer images of each image in the group of images, determine the detail layer fusion strategy of the group of images based on the noise estimation results, and fuse the detail layer images of the same layer of each image in the group of images using the determined detail layer fusion strategy; determine the saliency of each image in the group of images, and fuse the base layer images of each image in the group of images based on the saliency of each image in the group of images.

[0068] Detail layer fusion

[0069] In Example 1, noise estimation can be performed on the detail layer images of each image in the set of images to be fused. For example, if the set of images to be fused includes a visible light image, a short-wave infrared image, and a long-wave infrared image, then noise estimation can be performed on the detail layer images of the visible light image, the short-wave infrared image, and the long-wave infrared image, respectively.

[0070] In Example 1, performing noise estimation on the detail layer images of each image in the set of images to be fused may include:

[0071] For any image in the set, determine the significant difference ratio corresponding to each layer of the multi-scale representation of the image, and estimate the noise of the detail layer image based on the significant difference ratio corresponding to each layer.

[0072] Specifically, determining the significant difference ratio corresponding to any layer of the multi-scale representation of the image can include:

[0073] Filter the image of this layer to obtain the filtered image corresponding to this layer;

[0074] Calculate the absolute difference between the image of this layer and the corresponding filtered image, obtain the significant difference portion based on the absolute difference, and determine the significant difference ratio corresponding to this layer image based on the significant difference portion.

[0075] The noise estimation process is explained further below:

[0076] For any image in the set of images to be fused, it is assumed that the multi-scale decomposition of the image has 3 layers (experiments show that when the multi-scale decomposition has 3 layers, the fusion effect and calculation speed of Example 1 can reach the optimal solution, but this does not constitute a limitation on the number of layers of multi-scale decomposition).

[0077] The first layer image of the multi-scale decomposition of the image is filtered, and the absolute difference image G1 between the first layer image of the multi-scale decomposition and its filtered image is calculated. This process is represented as follows:

[0078] G1 = abs(M1 - MT1);

[0079] Where M1 is the first layer image of the multi-scale decomposition of the image, MT1 is the filtered image of the first layer image of the multi-scale decomposition of the image, and abs(·) represents the absolute value operation.

[0080] Then, a thresholding process is performed on the absolute difference image G1. Specifically, the portion of the absolute difference image G1 that is greater than a preset value (the preset value can be, for example, 10) is taken to obtain the significantly different portion B1.

[0081] Then calculate the significant difference ratio a1 using the following formula:

[0082] a1 = sum(B1) / 4 3-i ;

[0083] Where sum(·) represents the summation operation, and i represents the number of layers in the multi-scale decomposition (i is the number of layers corresponding to which the significant difference ratio of the image is calculated; for example, i is 1 when calculating the significant difference ratio of the first image).

[0084] Similarly, calculate the significant difference ratios a2 and a3 for the second and third layers of the multi-scale decomposition of the image.

[0085] Based on the significant difference ratios corresponding to each layer of the multi-scale decomposition of the image, the estimated noise level e of the image can be calculated, as shown below:

[0086]

[0087] Based on the estimated noise level of the image (which can be used as the noise level estimate), the noise level of the image can be determined (the noise level of the image is equivalent to the noise level of the detail layer image). For example, if the estimated noise level of the image is less than a certain threshold, the image is judged to have a low noise level; if the estimated noise level of the image is greater than or equal to a certain threshold, the image is judged to have a high noise level.

[0088] For each image in the set of images to be fused, noise estimation of its detail layer can be performed using the above method.

[0089] In Example 1, for each image in a set of images to be fused, a detail layer fusion strategy for the set of images can be determined based on the noise estimation results (i.e., the noise estimation results of the detail layer images of each image in the set). Determining the detail layer fusion strategy based on the noise estimation results may include:

[0090] For any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images does not exceed a first condition, then the detail layer fusion strategy for the pair of detail layer images is a fusion strategy based on the maximum absolute value; and / or, for any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images exceeds a second condition, then the detail layer fusion strategy for the pair of detail layer images includes: firstly performing preliminary fusion using a fusion strategy based on the maximum absolute value, and then processing the noise of the preliminary fusion result to obtain the target fusion result. Wherein, processing the noise of the preliminary fusion result to obtain the target fusion result may include: calculating spatial adaptive weights based on one of the detail layer images in the pair, calculating a spatial adaptive matrix based on the spatial adaptive weights, and calculating the target fusion result based on the spatial adaptive matrix and the other detail layer image in the pair.

[0091] The first and second conditions mentioned above can be the same or different, and can be freely changed, as is the case in Example 1.

[0092] The following section provides a further explanation of the detail layer blending process:

[0093] First, in Example 1, regardless of the number of images in a set of images to be fused, the detail layer fusion process is performed sequentially. That is, the detail layer images of two images in the set are first fused, then the latest fused detail layer image is fused with the detail layer image of the next image, and so on, each time fusing the latest fused detail layer image with the detail layer image of the next image, until the detail layer images of all images in the set are fused. For example, if the set of images to be fused includes a visible light image, a short-wave infrared image, and a mid-to-long-wave infrared image, the detail layer images of the visible light image and the short-wave infrared image are first fused, then the fused result (i.e., the fused image obtained after fusing the detail layer images of the visible light image and the short-wave infrared image) is fused with the detail layer image of the mid-to-long-wave infrared image, thereby obtaining a multi-layer three-light fused image.

[0094] Furthermore, as can be seen from the aforementioned process of determining detail layer images, each image actually has multiple detail layer images. Therefore, fusing the detail layer images of two images involves fusing the same layer images of each of the two images pairwise, and the fusion result is also in the form of a multi-layered image (each layer image in the fusion result is also called a detail layer image). Similarly, fusing the fusion result with the detail layer image of the next image involves fusing the fusion result with the same layer images of the detail layer image of the next image pairwise, and the fusion result is also in the form of a multi-layered image.

[0095] The different detail layer fusion strategies are explained below, case by case.

[0096] Scenario 1

[0097] In Example 1, for any pair of detail layer images to be fused (as can be seen above, this pair of detail layer images are in the same layer, the same below), if the noise estimation result of the pair of detail layer images does not exceed the first condition (meaning the noise level is low), then the detail layer fusion strategy of the pair of detail layer images is a fusion strategy based on the maximum absolute value (or a fusion strategy that takes the maximum absolute value).

[0098] Specifically, in this fusion strategy, for any pair of detail layer images to be fused, assuming that the pair of detail layer images are both at layer j, the weighting coefficient W used for fusion is... j Defined as:

[0099]

[0100] In the formula, d1 is the first source image (i.e., the source image of the first detail layer image in the pair of detail layer images), and d2 is the second source image (i.e., the source image of the second detail layer image in the pair of detail layer images).

[0101] The fusion result M of this pair of detail layer images at level j j Defined as:

[0102]

[0103] Scenario 2

[0104] In Example 1, for any pair of detail layer images to be fused (assuming that the pair of detail layer images are both at layer j), if the noise estimation result of the pair of detail layer images exceeds the second condition (meaning that the noise level is high; the second condition and the first condition may be the same or different), the detail layer fusion strategy for the pair of detail layer images includes: firstly, using a fusion strategy based on the maximum absolute value for preliminary fusion, and then processing the noise of the preliminary fusion result to obtain the target fusion result (the target fusion result is the final fusion result of the pair of detail layer images).

[0105] Specifically, the pair of detail layer images are first initially fused using the maximum absolute value-based fusion strategy described in Case 1. The initial fusion result is M. j .

[0106] Spatial adaptive weights Defined as:

[0107]

[0108] In the formula, p represents the spatial position of the pixel, ω p It is a square window centered on pixel p. Embodiment 1 proposes that a satisfactory fusion result can be obtained when the window size is 7×7, but Embodiment 1 does not limit the window size.

[0109] To facilitate program code implementation, a mean filter with a kernel size of 7×7 can be used to filter image d1. Dividing the filtered result by the kernel size yields the spatial adaptive weight 'a' for the entire image. j .

[0110] Calculate the spatial adaptive matrix A using the following formula j :

[0111]

[0112] In the formula, ε is a very small constant (set to 0.0001 in Example 1, but its specific value is not limited in Example 1) to prevent the denominator from being 0; the parameter λ is responsible for controlling the fusion effect. After experimentation, the fusion effect is more in line with human visual observation habits when λ is set to 0.8 (but its specific value is not limited in Example 1).

[0113] Target fusion result D j The calculation method is as follows:

[0114]

[0115] As mentioned earlier, when two detail layer images of layer j are fused, the resulting fusion is still considered as layer j.

[0116] The noise estimation result can be determined based on the overall noise level of the pair of detail layer images, or based on the noise level of one of the detail layer images. For example, considering the processing speed, the noise level of the detail layer image of the short-wave infrared image can be used as the standard. If the noise level of the detail layer image of the short-wave infrared image does not exceed the first condition (or exceeds the second condition), then the noise estimation result does not exceed the first condition (or exceeds the second condition).

[0117] During the detail layer fusion process, the fusion strategy is determined based on the actual situation. For example, when fusing the j-th detail layer image of a visible light image and a short-wave infrared image, since both have high noise levels, the fusion strategy in case two is suitable. When fusing the image fusion result of the j-th detail layer image of the visible light image and the short-wave infrared image with the same layer (i.e., the j-th layer) detail layer image of the long-wave infrared image, since the detail layer image of the long-wave infrared image has less noise, the fusion strategy in case one can be used.

[0118] In Example 1, the detail layer fusion strategy is determined based on the noise estimation results. This approach can fully consider the characteristics of different images, resulting in higher fusion quality. For example, compared to visible light images, infrared images often exhibit coarse-scale structural information but lack detail information conducive to visual perception. Furthermore, infrared images may contain significant noise and details detrimental to visual perception, all of which can reduce the quality of the fused image. In this case, It will become smaller, and the corresponding a j This will increase, thus increasing the visible light image detail layer. More detailed information will be integrated into D. j This makes the fused image more suitable for human perception. When the infrared detail layer contains some high-contrast and coarse-scale details, a fusion strategy of "taking the maximum absolute value" can satisfy the requirements. In this case, the weight a... j Will because The value decreases as the value increases. In this case, the fused detail layer D... j The detail layer M obtained by combining the "maximum absolute value" strategy j The images are closer together, thus meeting the requirements. The fused images retain more details beneficial to visual perception and are less affected by noise. Therefore, in Example 1, the detail layer fusion strategy is determined based on the noise estimation results, achieving adaptive noise processing during the fusion process, including noise processing for images with high noise levels, so that as little noise as possible is fused.

[0119] Specifically, to make the fusion results of the detail layers easier for the human eye to observe, the fusion weight of specific bands was enhanced in Example 1, including enhancing the fusion weight of visible light and shortwave infrared. That is, when a set of images to be fused includes visible light images, shortwave infrared images, and other band images (e.g., mid- and long-infrared images), the fusion of the same-layer detail layers of each image in the set of images may include: first performing the fusion of the same-layer detail layers of the visible light image and the shortwave infrared image, and then performing a fusion weight enhancement operation during the fusion of the same-layer detail layers of the visible light image and the shortwave infrared image.

[0120] Specifically, the fusion weight enhancement operation can include: For visible light images, the first layer result of the visible light image pyramid decomposition is denoted as M1, and a coefficient (called the first coefficient) is calculated using the following coefficient calculation formula. Similarly, for shortwave infrared images, the first layer result of the shortwave infrared image pyramid decomposition is denoted as M1, and a coefficient (called the second coefficient) is calculated using the following coefficient calculation formula. When fusing the same level detail layer of the visible light image and the shortwave infrared image, the detail layer of the visible light image is multiplied by the first coefficient to obtain a result, and the detail layer of the shortwave infrared image is multiplied by the second coefficient to obtain a result. Using these two results as input, the various fusion weights involved in the fusion process of the same level detail layer images of the visible light image and the shortwave infrared image (including the weighting coefficient W in Case 1) are calculated using the methods shown in Case 1 and Case 2. j This also includes the spatial adaptive weight 'a' in case two. j and the spatial adaptive matrix A j In other words, when calculating the fusion weights, the detail layer images of the same layer in the visible light image and the shortwave infrared image are not used directly. Instead, the two results mentioned above are used to replace the detail layer images of the visible light image and the shortwave infrared image in the calculation of the fusion weights. The coefficient calculation formula is as follows:

[0121] σ=exp(τ(255-M1)).

[0122] In the formula, M1 is the first layer result of the image pyramid decomposition, and τ is used to control the enhancement effect (it is set to 0.0015 in Example 1, but its specific value is not limited in Example 1).

[0123] Basic layer fusion

[0124] In Example 1, for each image in a set of images to be fused, the saliency (or visual saliency) of each image in the set of images can be determined, and the basic layer images of each image in the set of images are fused based on the saliency of each image in the set of images.

[0125] Determining the saliency of each image in the set of images may include: determining the saliency of any image in the set by comparing the contrast of a single pixel in that image with the other pixels in that image.

[0126] The process of fusing the base layer images of each image in the set of images based on their saliency may include: for a pair of base layer images to be fused, determining the base layer fusion weights based on the saliency of each of the pair of base layer images, and fusing the pair of base layer images using the base layer fusion weights.

[0127] The image fusion process at the base layer is explained in further detail below:

[0128] First, in Example 1, regardless of the number of images in a set of images to be fused, the basic layer fusion process is performed sequentially. That is, the basic layer images of two images in the set are first fused together. Then, the fused result of the latest basic layer image is fused with the basic layer image of the next image, and so on, each time fused with the basic layer image of the next image, until the basic layer images of all images in the set are fused. For example, if the set of images to be fused includes a visible light image, a short-wave infrared image, and a mid-to-long-wave infrared image, the basic layer images of the visible light image and the short-wave infrared image are first fused together. Then, the fused result (i.e., the fused image obtained after fusing the basic layer images of the visible light image and the short-wave infrared image) is fused with the basic layer image of the mid-to-long-wave infrared image, thereby obtaining a multi-layer three-light fused image.

[0129] 1. Example 1 proposes a method to define pixel-level saliency based on the contrast between a pixel and other pixels in an image, that is, to define pixel-level saliency based on the contrast between a pixel and other pixels in an image.

[0130] Specifically, I p Let V(p) represent the intensity value of pixel p in image I. The salience value of pixel p is defined as:

[0131] V(p)=|I p -I1|+|I p -I2|+…+|I p -I N |;

[0132] In the formula, N represents the total number of pixels in image I. If two pixels have the same intensity value, then their salience values ​​are also the same, and the above formula can be expressed as:

[0133]

[0134] Where j represents the pixel intensity value, M j V(p) represents the number of pixels with intensity j, and L is the total number of pixel intensity values ​​(256). V(p) is then normalized to [0,1]. Furthermore, this method of defining pixel-level saliency can be accelerated using an intensity histogram.

[0135] For two images, assuming their saliency maps are represented as V1 and V2, and their base layers are represented as B1 and B2, then the fusion result of the base layers of the two images is B. F It can be represented as follows:

[0136] B F =W bB1+(1-W b B2;

[0137] Wherein, weight W b Defined as:

[0138]

[0139] The above formula is a weighted averaging technique that takes into account visual saliency. If V1 equals V2 at some locations in the two images, then the weight W... b This will degenerate into an "average" weight. If V1 is greater than V2 at some position, then the weight W... b If the value is greater than 0.5, then more information from the basic layer B1 will be integrated into B. F Otherwise, when V1 is less than V2, more information from the basic layer B2 will be integrated into B. F .

[0140] Specifically, to make the fusion result more consistent with human observation habits, in Example 1, image enhancement operations are performed on the basic layer fusion process for specific wavelength bands. This includes the basic layer image fusion process of visible light images and short-wave infrared images, and the image enhancement operation performed on the fusion result of the basic layer images of visible light images and short-wave infrared images with the basic layer image process of mid-to-long-wave infrared images. That is, when a set of images to be fused includes visible light images, short-wave infrared images, and mid-to-long-wave infrared images, the fusion of the basic layer images of each image in the set of images may include: first, performing basic layer image fusion of visible light images and short-wave infrared images; then, fusing the fusion result of the basic layer images of visible light images and short-wave infrared images (i.e., the dual-light fusion result) with the basic layer image of mid-to-long-wave infrared images; and performing image enhancement operations during the fusion process of the basic layer images of visible light images and short-wave infrared images, and during the fusion process of the fusion result of the basic layer images of visible light images and short-wave infrared images with the basic layer image process of mid-to-long-wave infrared images.

[0141] Specifically, image enhancement operations may include: during the fusion of the base layer images of visible light and shortwave infrared images, normalizing the visual saliency of both the visible light and shortwave infrared images before calculating the fusion weight W. b This ensures the rationality of the weighting ratio, making it more suitable for weighted calculations and resulting in a more natural image appearance. When fusing the base layer fusion results of visible light and short-wave infrared images with the base layer image of long-wave infrared images, the weight of the visible light and short-wave infrared fusion results is amplified by 2.5 times (the amplification factor is not limited in Example P), thereby enhancing the proportion of visible light and short-wave infrared in the fusion process. This makes the background, such as the sky, and the entire image brighter and easier for the human eye to observe.

[0142] S105: Obtain the fused image of this set of images based on the detail layer fusion result and the basic layer fusion result.

[0143] Assuming a single image has N detail layers, the final result of fusing the detail layers of any set of images (i.e., the fused detail layer result of the set of images) also has N layers, denoted as D. 1 To D N The result of fusing the base layer images in this set of images (i.e., the base layer fusion result of this set of images) is represented as B. F .

[0144] In Example 1, the fused image of the set of images can be obtained based on the detail layer fusion result and the base layer fusion result. Specifically, the fused image F of the set of images can be reconstructed by adding the detail layer fusion result and the base layer fusion result, as shown below:

[0145] F = B F +D 1 +D 2 +…+D N .

[0146] In actual calculations, the corresponding pixel values ​​are directly added together to obtain the final value of each pixel in the fused image F, thus completing the reconstruction from the decomposed image layers (including detail layer images and base layer images) to the fused image.

[0147] Example 1 has been practically verified, as detailed below:

[0148] A prototype of a multi-band enhanced visual system, composed of visible light, short-wave infrared, and long-wave infrared sensors, was used to acquire external visual images of practical application scenarios such as airport approach lights under low-visibility weather conditions. The acquired three-light images were then used to verify Example 1. Experimental results show that the image fusion method provided in Example 1 can effectively fuse image information from various channels while meeting real-time requirements, with high feature preservation, comfortable visual experience for the human eye, significantly reduced noise, and compliance with design requirements, demonstrating the feasibility of the engineering implementation of Example 1. Specific experimental results are as follows... Figure 2 As shown.

[0149] Example 1 can achieve the following beneficial effects:

[0150] Example 1 employs multi-scale decomposition, innovates the fusion strategy between the detail layer and the base layer, introduces adaptive noise processing and original image enhancement functions, and optimizes the real-time performance of image fusion. While ensuring a short fusion latency, it effectively overcomes fusion problems such as detail blurring, halo effects, and noise interference, possessing advantages such as good fusion results, better conformity to human visual perception, and strong scene adaptability. Compared to traditional image fusion methods, it solves problems such as the inability to effectively fuse key information from the original image for human visual perception and the indiscriminate amplification of noise in the image.

[0151] Regarding multi-scale decomposition, Implementation 1 proposes an improved scheme based on image pyramids. By performing differential and upsampling operations on adjacent layers of the Gaussian pyramid, it enhances edge preservation capabilities and highlights high-frequency information, while also being computationally simple and meeting real-time requirements. For detail layer fusion, Implementation 1 fully considers differences in image characteristics and performs adaptive noise processing to improve the fusion effect. In the basic layer fusion method, Implementation 1 introduces a weighted averaging technique based on visual saliency to make the fusion more natural. Regarding original image enhancement, Implementation 1 performs adaptive enhancement within the fusion framework to address noise and brightness saturation issues. In terms of noise processing, Implementation 1 proposes an adaptive noise processing scheme. Unlike traditional fixed-mode noise processing methods that struggle to cope with complex and varied real-world shooting scenarios, Implementation 1 takes a unique approach by estimating noise at the detail layer, achieving accurate processing of images with different noise levels.

[0152] Specifically, considering that camera-captured scenes often contain blank areas with limited image detail, such as the sky, road surface, and walls, Implementation 1 uses this as a starting point. First, the entire image is filtered. By calculating the absolute difference between the original image (the image after multi-scale representation) and the filtered image, areas where noise may exist are accurately located. Then, the difference image is thresholded, retaining only the portion exceeding a preset value, thus obtaining the significantly different portion. This information is closely related to the noise level. Furthermore, based on the proportion of the significantly different portion in the image, an estimate of the noise level can be accurately calculated. Based on this estimate, the image noise level can be automatically determined. If the estimate is less than the threshold, the image noise level is low; if it is greater than the threshold, the image noise level is high.

[0153] For images with high noise levels, Example 1 eschews traditional image preprocessing. Instead, it cleverly incorporates noise reduction during the detail layer fusion process, effectively reducing the degree of noise fusion and minimizing noise interference with subsequent image processing and applications. The detail layer fusion strategy employed in Example 1 is an improvement upon the weighted least squares optimization method. Traditional weighted least squares fusion strategies combine complementary information from different modalities through optimization. Example 1, however, improves the calculation process, fully leveraging the potential of the detail layer fusion strategy in suppressing source image noise. Through mathematical derivation and reasonable approximations, the calculation of fusion weights eliminates the need for a cumbersome optimization process, allowing for a more straightforward and direct calculation.

[0154] This noise processing step not only significantly improves the accuracy and relevance of noise handling, but also effectively preserves key information and details in complex scenes, providing high-quality image data for subsequent image processing and analysis. Furthermore, the scheme's computational process is simple and efficient. Through a well-designed noise estimation and processing flow, it meets real-time requirements while ensuring effective noise processing, avoiding the impact of complex calculations on operational efficiency.

[0155] Example 1 proposes an image enhancement strategy based on human visual observation habits, overcoming the limitations of traditional image enhancement methods that fail to fully consider the characteristics of human vision, leading to deviations between the enhancement effect and human expectations. Example 1 achieves a deep alignment between the image enhancement effect and human observation habits by simulating the laws of human visual perception. Specifically, human vision is caused by visible light stimulating the optic nerve. Due to the unique physiological mechanisms of the optic nerve, humans also possess some inherent characteristics when observing images. Research shows that when the background brightness is too strong or too weak, the resolution of the human eye decreases, meaning that the overall brightness of the image should be moderate. The human eye has a higher resolution for brightness information than for color information; that is, the human eye is more sensitive to brightness information. Based on these characteristics of human vision, Example 1 innovatively performs a normalization operation on the visual saliency calculated from the two source images (i.e., the visible light image and the shortwave infrared image) during the basic layer fusion process of the visible light image and the shortwave infrared image, and then calculates the fusion weight accordingly. This processing method precisely ensures the rationality of the weight ratio, making the weighted calculation more scientific, and thus allowing the fused image to present an extremely natural visual effect, fully meeting the human eye's expectations for the visual presentation of real scenes.

[0156] When fusing the results of dual-light fusion with a long-wave infrared image, Example 1 innovatively amplifies the weighting of the visible light and short-wave infrared fusion results. This significantly enhances the proportion of visible light and short-wave infrared in the fusion process, making the brightness of background areas such as the sky and the entire image more comfortable to view. This greatly optimizes the image's visual friendliness, allowing the human eye to capture key information in the image more easily and efficiently.

[0157] Example 1 not only significantly improves the quality of image fusion and achieves visual optimization and upgrades in image enhancement, but more importantly, it closely revolves around the human visual mechanism in its image enhancement design, effectively promoting the efficient integration of image information and human perception. Simultaneously, through a carefully designed weight adjustment strategy and a simple and reasonable processing flow, it ensures the high efficiency of the algorithm while achieving excellent image enhancement results, making it widely applicable to various application scenarios that demand both high image quality and processing efficiency.

[0158] Through the innovations described in Example 1, the image fusion effect and performance can be comprehensively improved, providing users with higher quality visual information.

[0159] Example 1 can be applied to the field of avionics technology, including cockpit infrared vision enhancement systems (EVS) and various devices and scenarios requiring high-quality image fusion. It enhances pilots' situational awareness in nighttime and low-visibility scenarios, meeting practical application requirements for image clarity, detail preservation, natural contrast, and real-time algorithm performance. It effectively overcomes fusion problems caused by sensor characteristics, such as detail blurring, halos, and noise interference, significantly improving the clarity and accuracy of pilots' visual perception in complex environments, thus ensuring flight safety. Example 1 can provide guidance for the design of image fusion algorithms for next-generation EVS devices.

[0160] The second embodiment of this specification provides an image fusion apparatus based on adaptive noise processing and visual enhancement, corresponding to the method described in the embodiments. The apparatus includes:

[0161] The image decomposition module is used to construct a multi-scale representation of any image in a set of images to be fused, and to use the multi-scale representation to determine the detail layer image and the base layer image of the image.

[0162] The detail layer and base layer fusion module is used to estimate the noise of the detail layer images of each image in the group of images, determine the detail layer fusion strategy of the group of images based on the noise estimation results, and fuse the detail layer images of the same layer of each image in the group of images according to the determined detail layer fusion strategy; determine the saliency of each image in the group of images, and fuse the base layer images of each image in the group of images according to the saliency of each image in the group of images.

[0163] The image reconstruction module is used to obtain a fused image of the set of images based on the detail layer fusion result and the base layer fusion result.

[0164] Preferably, a multi-scale representation of the image is constructed, and the detail layer image and the base layer image of the image are determined using the multi-scale representation, including:

[0165] Perform a Gaussian filter operation on the image and construct the corresponding Gaussian pyramid.

[0166] For each non-top layer image of the Gaussian pyramid, the predicted image of the previous layer is subtracted from the image of that layer, and the result is used as the detail layer image; and the top layer image of the Gaussian pyramid is used as the base layer image.

[0167] For any given image layer, the predicted image for that layer is the image obtained by upsampling the image of that layer and then performing a Gaussian convolution.

[0168] Preferably, noise estimation is performed on the detail layer images of each image in the set of images, including:

[0169] For any image in the set, determine the significant difference ratio corresponding to each layer of the multi-scale representation of the image, and estimate the noise of the detail layer image based on the significant difference ratio corresponding to each layer.

[0170] Specifically, for any layer of the multi-scale representation of the image, determining the significant difference ratio corresponding to that layer includes:

[0171] Filter the image of this layer to obtain the filtered image corresponding to this layer;

[0172] Calculate the absolute difference between the image of this layer and the corresponding filtered image, obtain the significant difference portion based on the absolute difference, and determine the significant difference ratio corresponding to this layer image based on the significant difference portion.

[0173] As a preferred approach, the detail layer fusion strategy for this set of images, determined based on the noise estimation results, includes:

[0174] For any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images does not exceed the first condition, then the detail layer fusion strategy of the pair of detail layer images is a fusion strategy based on the maximum absolute value.

[0175] As a preferred approach, the detail layer fusion strategy for this set of images, determined based on the noise estimation results, includes:

[0176] For any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images exceeds the second condition, the detail layer fusion strategy of the pair of detail layer images includes: firstly, using a fusion strategy based on the maximum absolute value to perform preliminary fusion, and then processing the noise of the preliminary fusion result to obtain the target fusion result.

[0177] Preferably, the preliminary fusion result is subjected to noise processing to obtain the target fusion result, including:

[0178] The spatial adaptive weights are calculated based on one of the detail layer images in the pair. The spatial adaptive matrix is ​​then calculated based on the spatial adaptive weights. Finally, the target fusion result is calculated based on the spatial adaptive matrix and the other detail layer image in the pair.

[0179] Preferably, determining the saliency of each image in this set of images includes:

[0180] For any image in the set, the saliency of the image is determined by the contrast between a single pixel in that image and the other pixels in that image.

[0181] Preferably, the fusion of the base layer images of each image in the set of images based on the saliency of each image includes:

[0182] For a pair of base layer images to be fused, the base layer fusion weights are determined based on the saliency of each base layer image, and the base layer images are fused using the base layer fusion weights.

[0183] Preferably, the set of images to be fused includes visible light images, shortwave infrared images, and images in other bands;

[0184] The fusion of the same level detail images in this image set includes:

[0185] First, perform image fusion of the same layer detail layer of the visible light image and the shortwave infrared image, and then perform a fusion weight enhancement operation during the image fusion of the same layer detail layer of the visible light image and the shortwave infrared image;

[0186] And / or,

[0187] The set of images to be fused includes visible light images, short-wave infrared images, and mid-to-long-wave infrared images;

[0188] The fusion of the base layer images of each image in this set of images includes:

[0189] First, the basic layer images of the visible light image and the short-wave infrared image are fused. Then, the fusion result of the basic layer images of the visible light image and the short-wave infrared image is fused with the basic layer image of the mid- and long-wave infrared image. Image enhancement operations are performed during the fusion process of the basic layer images of the visible light image and the short-wave infrared image, and during the fusion process of the basic layer images of the visible light image and the short-wave infrared image and the mid- and long-wave infrared image.

[0190] The contents not described in detail in Embodiment 1 and Embodiment 2 can be referred to each other. Embodiment 2 can achieve the same beneficial effects as Embodiment 1. The above embodiments can be used in combination.

[0191] The above description is merely an embodiment of this specification and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. An image fusion method based on adaptive noise processing and visual enhancement, characterized in that, The method includes: For any image in a set of images to be fused, construct a multi-scale representation of the image, and use the multi-scale representation of the image to determine the detail layer image and the base layer image of the image; Noise is estimated for the detail layer images of each image in the image set. Based on the noise estimation results, a detail layer fusion strategy for the image set is determined. The same level detail layer images of each image in the image set are then fused using the determined detail layer fusion strategy. The saliency of each image in the image set is determined. Based on the saliency of each image in the image set, the base layer images of each image in the image set are then fused. The fused image of this set of images is obtained based on the fusion results of the detail layer and the basic layer.

2. The method as described in claim 1, characterized in that, Constructing a multi-scale representation of the image, and using this multi-scale representation to determine the detail layer image and the base layer image, including: Perform a Gaussian filter operation on the image and construct the corresponding Gaussian pyramid. For each non-top layer image of the Gaussian pyramid, the predicted image of the previous layer is subtracted from the image of that layer, and the result is used as the detail layer image; and the top layer image of the Gaussian pyramid is used as the base layer image. For any given image layer, the predicted image for that layer is the image obtained by upsampling the image of that layer and then performing a Gaussian convolution.

3. The method as described in claim 1, characterized in that, Noise estimation was performed on the detail layer images of each image in this set of images, including: For any image in the set, determine the significant difference ratio corresponding to each layer of the multi-scale representation of the image, and estimate the noise of the detail layer image based on the significant difference ratio corresponding to each layer. Specifically, for any layer of the multi-scale representation of the image, determining the significant difference ratio corresponding to that layer includes: Filter the image of this layer to obtain the filtered image corresponding to this layer; Calculate the absolute difference between the image of this layer and the corresponding filtered image, obtain the significant difference portion based on the absolute difference, and determine the significant difference ratio corresponding to this layer image based on the significant difference portion.

4. The method as described in claim 1, characterized in that, Based on the noise estimation results, the detail layer fusion strategy for this group of images is determined as follows: For any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images does not exceed the first condition, then the detail layer fusion strategy of the pair of detail layer images is a fusion strategy based on the maximum absolute value.

5. The method as described in claim 1 or 4, characterized in that, Based on the noise estimation results, the detail layer fusion strategy for this group of images is determined as follows: For any pair of detail layer images to be fused, if the noise estimation result of the pair of detail layer images exceeds the second condition, the detail layer fusion strategy of the pair of detail layer images includes: firstly, using a fusion strategy based on the maximum absolute value to perform preliminary fusion, and then processing the noise of the preliminary fusion result to obtain the target fusion result.

6. The method as described in claim 5, characterized in that, The initial fusion results are subjected to noise processing to obtain the target fusion results, including: The spatial adaptive weights are calculated based on one of the detail layer images in the pair. The spatial adaptive matrix is ​​then calculated based on the spatial adaptive weights. Finally, the target fusion result is calculated based on the spatial adaptive matrix and the other detail layer image in the pair.

7. The method as described in claim 1, characterized in that, Determining the saliency of each image in this set of images includes: For any image in the set, the saliency of the image is determined by the contrast between a single pixel in that image and the other pixels in that image.

8. The method as described in claim 1 or 7, characterized in that, The base layer images of each image in the image set are fused based on their saliency, including: For a pair of base layer images to be fused, the base layer fusion weights are determined based on the saliency of each base layer image, and the base layer images are fused using the base layer fusion weights.

9. The method as described in claim 1, characterized in that, The set of images to be fused includes visible light images, shortwave infrared images, and images in other bands; The fusion of the same level detail images in this image set includes: First, perform image fusion of the same layer detail layer of the visible light image and the shortwave infrared image, and then perform a fusion weight enhancement operation during the image fusion of the same layer detail layer of the visible light image and the shortwave infrared image; And / or, The set of images to be fused includes visible light images, short-wave infrared images, and mid-to-long-wave infrared images; The fusion of the base layer images of each image in this set of images includes: First, the basic layer images of the visible light image and the short-wave infrared image are fused. Then, the fusion result of the basic layer images of the visible light image and the short-wave infrared image is fused with the basic layer image of the mid- and long-wave infrared image. Image enhancement operations are performed during the fusion process of the basic layer images of the visible light image and the short-wave infrared image, and during the fusion process of the basic layer images of the visible light image and the short-wave infrared image and the mid- and long-wave infrared image.

10. An image fusion device based on adaptive noise processing and visual enhancement, characterized in that, The device includes: The image decomposition module is used to construct a multi-scale representation of any image in a set of images to be fused, and to use the multi-scale representation to determine the detail layer image and the base layer image of the image. The detail layer and base layer fusion module is used to estimate the noise of the detail layer images of each image in the group of images, determine the detail layer fusion strategy of the group of images based on the noise estimation results, and fuse the detail layer images of the same layer of each image in the group of images according to the determined detail layer fusion strategy; determine the saliency of each image in the group of images, and fuse the base layer images of each image in the group of images according to the saliency of each image in the group of images. The image reconstruction module is used to obtain a fused image of the set of images based on the detail layer fusion result and the base layer fusion result.