Multi-exposure image fusion method based on perceptual enhancement structural block decomposition

By decomposing image patches into perceptual gain, signal strength, and structure, and combining logarithmic stretching and multi-scale decomposition, the problem of insufficient information and lack of realism in existing multi-exposure image fusion methods under poor exposure conditions is solved, achieving efficient and realistic image fusion results.

CN117315416BActive Publication Date: 2025-10-31CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311006314.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-04-04
Filing Date
2023-08-10
Publication Date
2025-10-31
Estimated Expiration
2043-08-10

AI Technical Summary

Technical Problem

Existing multi-exposure image fusion methods cannot effectively recover rich information under poor exposure conditions, and the reconstructed results are not realistic and have high computational costs.

Method used

Image patches are decomposed into perceptual gain, signal strength, signal structure, and average strength. Perceptual gain is estimated by logarithmic stretching, and fusion is performed using Hadamard product operation and multi-scale decomposition strategy to reconstruct the fused image.

Benefits of technology

By generating information-rich and perceptually realistic fused images under different exposure conditions, the efficiency and quality of image fusion are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315416B_ABST
    Figure CN117315416B_ABST
Patent Text Reader

Abstract

This invention relates to image processing methods, providing a multi-exposure image fusion method based on perceptual enhancement structured block decomposition, comprising the following steps: S1. Acquiring multiple exposure source images, and decomposing each image block in each exposure source image into perceptual gain, signal intensity, signal structure, and average intensity; S2. Estimating the perceptual gain using a logarithmic stretching method; S3. Fusing the perceptual gain, signal intensity, signal structure, and average intensity respectively to obtain fused perceptual gain, fused signal intensity, fused signal structure, and fused average intensity, and reconstructing a fused image block; S4. Integrating all the fused image blocks to obtain the final fused image. This invention's multi-exposure image fusion method based on perceptual enhancement structured block decomposition can obtain information-rich and perceptually realistic results under different exposure rates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to image processing methods, specifically to a multi-exposure image fusion method based on perceptual enhancement structure block decomposition. Background Technology

[0002] The dynamic range of a scene is much wider than that of an imaging device. Therefore, traditional imaging devices struggle to capture high dynamic range (HDR) scenes well in a single shot, often resulting in underexposure or overexposure. HDR imaging technology was developed to address this problem and has attracted significant attention from both academia and industry.

[0003] To reconstruct HDR images, a common approach is to use the inverted camera response function (CRF) followed by tone mapping in the luminance domain. Due to the difficulty in designing CRF estimation and tone mappers, multi-exposure image fusion (MEF) is an economical and convenient method for HDR imaging. MEF aims to accurately represent HDR scenes by integrating images with different exposures. Current MEF algorithms can be broadly categorized into traditional methods and deep learning-based methods. One existing technique is Structured Block Decomposition-based MEF (SPD-MEF), a representative algorithm in traditional methods. In SPD-MEF, image patches are decomposed into signal intensity, signal structure, and average intensity. Experimental results show that SPD-MEF achieves state-of-the-art performance. However, this algorithm is computationally expensive. To accelerate SPD-MEF, a fast multi-scale version, MSPD-MEF, was proposed, ignoring the normalization operator. MSPD-MEF has demonstrated its advantage over SPD-MEF. However, under poor exposure conditions, MSPD-MEF, by ignoring perceptual factors, fails to recover rich information from the original source image, resulting in unrealistic reconstructions.

[0004] In recent years, deep learning has been successfully applied to image fusion tasks such as infrared and visible light image fusion, multi-focus image fusion, polarization image fusion, and multi-exposure image fusion. Therefore, a fast multi-exposure image fusion network (MEFNet) for static scenes has emerged. MEFNet is trained end-to-end by optimizing perceptually calibrated MEF structural similarity (MEF-SSIM), and this network can fuse images of arbitrary spatial resolution and exposure numbers. DIFNet is a deep image fusion network (DIFNet) that uses structural tensor representation as the loss function instead of maximizing MEF-SSIM. DIFNet consists of three subnetworks: feature extraction, feature fusion, and image reconstruction. Furthermore, the average of all source images is used as a hypothetical HDR image to guide the training of DIFNet. Existing technologies also include a unified unsupervised image fusion network (U2Fusion), which can solve different fusion problems through continuous learning. It employs perceptual loss and structural similarity constraints to ensure that different fusion tasks are unified within the same framework. While these networks perform well in general multi-exposure image fusion tasks, the results are not ideal under poor exposure conditions because they only use information from the source images. To achieve information-rich and visually realistic results, a Deep Perception Enhancement Network (DPE-MEF) was proposed. DPE-MEF consists of two sub-modules: detail enhancement and color enhancement. The detail enhancement module aims to generate enhanced images by seeking optimal local exposure. Color enhancement is used to learn the relationship between color and brightness. In extreme exposure image fusion, the DPE-MEF method outperforms other methods. However, the fused results are not visually realistic enough, and DPE-MEF tends to reconstruct noisy fused results. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a multi-exposure image fusion method based on perceptual enhancement structure block decomposition, which can obtain information-rich and perceptually realistic results under different exposure rates.

[0006] To address the aforementioned technical problems, this invention provides a multi-exposure image fusion method based on perceptual enhancement structure block decomposition, comprising the following steps:

[0007] S1. Acquire multiple exposure source images, and decompose each image block in each exposure source image into sensing gain, signal strength, signal structure, and average intensity;

[0008] S2. Estimate the perceived gain using the logarithmic stretching method;

[0009] S3. The sensing gain, signal strength, signal structure and average intensity are fused respectively to obtain fused sensing gain, fused signal strength, fused signal structure and fused average intensity, and the fused image block is reconstructed;

[0010] S4. Integrate all the fused image blocks to obtain the final fused image.

[0011] Specifically, in step S1, K images of the exposure source are acquired, and the image blocks of size p×p×3 are rearranged into a vector, represented as follows:

[0012]

[0013] in, It is the set of real numbers;

[0014] The image block decomposition is represented by the following formula:

[0015]

[0016] Where, x k Let r represent the normalized vector of the k-th image (k is an integer, ranging from 1 to K), ||·|| represents the l2 norm of the vector, and r k For the image patch after perceptual gain of the k-th image, i k For the perceptual gain of the k-th image, c k Let s be the signal strength of the k-th image. k For the signal structure of the k-th image, μ k The average intensity of the k-th image.

[0017] Specifically, in step S2, the perceived gain i is estimated using the logarithmic stretching method. k The expression is as follows:

[0018]

[0019] Among them, l kc For l k The correction value, l k Represents the vector x k The mean, m k This represents the maximum mean value among the image blocks of the k-th image;

[0020] According to the following formula, l k Make corrections:

[0021] l kc =max(min(l k ,0.5),ξ)

[0022] Where ξ is set to 5 / 255.

[0023] Specifically, in step S3, the maximum mean value l in each of the image blocks is used. k Estimate the fusion sensing gain i f The formula is as follows:

[0024]

[0025] Among them l k (j) represents the j-th image block of the k-th image.

[0026] Specifically, the fused signal strength c f The maximum value of the signal intensity across all the image blocks is given by the following formula:

[0027]

[0028] Specifically, the fused signal structure s f This is the optimal representation of the structure in the exposure source image, and its formula is as follows:

[0029]

[0030] Where q is the exponential parameter, q≥0.

[0031] Specifically, the fusion average intensity μ f The calculations based on local bias and good exposure of the image patch yielded the following results:

[0032]

[0033] Wherein, weight α k Represented as:

[0034]

[0035] Where λ is a constant, σ k That is the standard deviation.

[0036] Specifically, in step S3, the fused image block xf reconstruction formula is as follows:

[0037]

[0038] in,

[0039] In step S4, the fused image X f The following formula can be used to synthesize the results:

[0040]

[0041] Where ⊙ represents the Hadamard product operation, A k For all α k I is obtained by integrating the source images based on their size and arrangement. f For i f U is derived by integrating the source images based on their size and arrangement. k For μ k B is derived by integrating the size and arrangement of the source images. k For β k The algorithm integrates the source images based on their size and arrangement. The operator φ(·) represents a mean filter with a kernel size of p×p, R. k For r k The result of mean filtering after aggregation.

[0042] Preferably, the method further includes step S5: performing multi-scale decomposition on the fused image obtained in step S4 to re-obtain the final fused image.

[0043] Specifically, step S5 includes the following steps:

[0044] A) Let E k =I f ⊙A k F k =I f ⊙B k Obtain detailed information H at the original scale. (1) :

[0045]

[0046] in, and U, respectively, at the original scale k E k and F k ;

[0047] B) Perform downsampling by a factor of n to obtain information at the nth scale. based on Calculate the nth scale and The detailed information H at the nth scale is obtained using the following formula. (n) :

[0048]

[0049] Where n is an integer, ranging from 2 to Z, and the decomposition process satisfies the iterative condition:

[0050]

[0051] Where h and w are the height and width of the image, respectively;

[0052] C) Calculate basic information L at the Z-scale (Z) :

[0053]

[0054] D) Using the bilinear upsampling operator U p (·), adding the base layer and detail layer together, yields a new blended image, as shown in the following formula:

[0055] X f =U p (…U p (U p (L (Z) +H (Z) )+H (Z-1) )+…)+H (1) .

[0056] The beneficial effects of the present invention through the above solution are as follows:

[0057] This invention presents a multi-exposure image fusion method based on perceptual enhancement structure block decomposition. First, the image block is decomposed into four components: perceptual gain, signal intensity, signal structure, and average intensity. By using logarithmic stretching to estimate the perceptual gain, the intensity of the underexposed image can be effectively increased to reveal the rich information hidden in the underexposed image. The perceptual gain, signal intensity, signal structure, and average intensity are then fused separately to reconstruct the fused image block. All the fused image blocks are then integrated to obtain the final fused image, which is rich in information and perceptually realistic.

[0058] Other features and advantages of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0059] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the following detailed description to explain the invention, but do not constitute a limitation thereof. In the drawings:

[0060] Figure 1 This is a flowchart illustrating the steps of the multi-exposure image fusion method based on perceptual enhancement structure block decomposition of the present invention;

[0061] Figure 2 This is a comparison of the visual effects of different methods used to fuse building images;

[0062] Figure 3 This is a visual comparison of the results of different methods for fusing images at the center of the activity. Detailed Implementation

[0063] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the present invention, and the scope of protection of the present invention is not limited to the specific embodiments described below.

[0064] This invention provides a multi-exposure image fusion method based on perceptual enhancement structured block decomposition, see [link to relevant documentation]. Figure 1 The method includes the following steps:

[0065] S1. Acquire multiple exposure source images, and decompose each image block in each exposure source image into sensing gain, signal strength, signal structure, and average intensity;

[0066] S2. Estimate the perceived gain using the logarithmic stretching method;

[0067] S3. The sensing gain, signal strength, signal structure, and average intensity are fused separately to obtain fused sensing gain, fused signal strength, fused signal structure, and fused average intensity, and the fused image patch is reconstructed.

[0068] S4. Integrate all the fused image blocks to obtain the final fused image.

[0069] Specifically, in step S1, K exposure source images are acquired, and image blocks of size p×p×3 are rearranged into vectors, represented as follows:

[0070]

[0071] in, It is the set of real numbers;

[0072] Each image block is decomposed into four parts: perceptual gain, signal strength, signal structure, and average strength, which are represented by the following formula:

[0073]

[0074] Where, x k Let r represent the normalized vector of the k-th image (k is an integer, ranging from 1 to K), ||·|| represents the l2 norm of the vector, and r k For the image patch after perceptual gain of the k-th image, i k Let c be the perceptual gain of the k-th image. k =||r k -μ k || represents the signal strength of the k-th image, s k =(r k -μ k ) / c k Let μ be the signal structure of the k-th image.k Let μ be the average intensity of the k-th image, and μ k For vector r k The mean.

[0075] To enhance the details of the underexposed image and maintain the highest brightness of the overexposed image, the cutoff point is set to 0.5 based on the characteristics of the histogram. In step S2, the perceptual gain i is estimated using the logarithmic stretching method. k Its expression is as follows:

[0076]

[0077] Among them, l kc For l k The correction value, l k Represents vector x k The mean, m k This represents the maximum mean value among the image patches of the k-th image;

[0078] According to the following formula, l k Make corrections:

[0079] l kc =max(min(l k ,0.5),ξ)

[0080] Among them, ξ is preferably set to 5 / 255.

[0081] Since the image patches were captured under the same imaging conditions, it can be assumed that the required perceptual gain for the entire image is uniform, thus improving the brightness of the fused image. In step S3, the maximum average value l among the image patches is used. k Estimate the fusion sensing gain i f The formula is as follows:

[0082]

[0083] Among them l k (j) represents the j-th image block of the k-th image.

[0084] Signal strength c k Related to local contrast, the higher the contrast, the better the visual perception. Specifically, the fused signal strength c f The maximum signal strength across all image patches is given by the following formula:

[0085]

[0086] Signal structure s k Given a unit-length vector, the fused signal structure s f It is the optimal representation of the structure in the exposure source image, and its formula is as follows:

[0087]

[0088] Where q is the exponential parameter, q≥0.

[0089] Specifically, the fusion average intensity μ f The calculations are based on local bias and good exposure of image patches:

[0090]

[0091] Wherein, weight α k Represented as:

[0092]

[0093] Where λ is a constant, σ k That is the standard deviation.

[0094] Using a larger λ in the arctan function helps to better preserve global brightness. Specifically, in step S3, the image patch x is fused. f The reconstruction formula is as follows:

[0095]

[0096] in,

[0097] In step S4, all the fused image blocks are integrated, and the fused image X is formed. f The following formula can be used to synthesize the results:

[0098]

[0099] Where ⊙ represents the Hadamard product operation, A k For all α k I is obtained by integrating the source images based on their size and arrangement. f For i f U is derived by integrating the source images based on their size and arrangement. k For μ k B is derived by integrating the size and arrangement of the source images. k For β k The algorithm integrates the source images based on their size and arrangement. The operator φ(·) represents a mean filter with a kernel size of p×p, R. k For r k The aggregated mean filter result is used to average overlapping image patches.

[0100] In order to improve the fusion performance, the present invention applies a multi-scale decomposition strategy to the proposed model. Specifically, it also includes step S5: performing multi-scale decomposition on the fused image obtained in step S4 to obtain the final fused image again.

[0101] Step S5 includes the following steps:

[0102] A) To simplify the expression, let E k =I f ⊙A k F k =I f ⊙B k Obtain detailed information H at the original scale. (1) :

[0103]

[0104] in, and U, respectively, at the original scale k E k and F k ;

[0105] B) Perform downsampling by a factor of n to obtain information at the nth scale. based on Calculate the nth scale and The detailed information H at the nth scale is obtained using the following formula. (n) :

[0106]

[0107] Where n is an integer, ranging from 2 to Z, and the decomposition process satisfies the iterative condition:

[0108]

[0109] Where h and w are the height and width of the image, respectively;

[0110] C) Calculate basic information L at the Z-scale (Z) :

[0111]

[0112] D) Using the bilinear upsampling operator U p (·), adding the base layer and detail layer together, yields a new blended image, as shown in the following formula:

[0113] X f =U p (…U p (Up (L (Z) +H (Z) )+H (Z-1) )+…)+H (1) .

[0114] The following visual and quantitative comparisons demonstrate that the multi-exposure image fusion method based on perceptual enhancement structured block decomposition of this invention is superior to existing image processing methods. In a preferred embodiment, p = 9, q = 4, and λ = 20 in this comparison. The fusion results of the building image sequence are as follows: Figure 2 As shown, (a) and (b) display underexposed and overexposed source images, while (c) to (l) list the fusion results of different methods, in order: MEFAW, MEFDSIFT, MTI, MGFF, MSPD-MEF, MEFNet, DIFNet, U2Fusion, DPE-MEF, and the method provided in this invention. It can be seen that the MEFAW method produces local dark areas, while the results of MTI and MSPD-MEF alleviate these dark areas. Traditional methods MEFDSIFT and MGFF can produce continuous results but lose details of walls and interiors. Among deep learning-based methods, MEFNet produces the worst results. DIFNet and U2Fusion improve the fusion results; however, the results are visually unrealistic and do not recover much detail. DPE-MEF is designed to solve the problem of extremely exposed image fusion, and it produces better and more realistic results than the others. However, referring to Figure (l), which shows the method proposed in this invention, compared with DPE-MEF, this invention reconstructs more details, such as walls and interiors, as shown in the boxes in the figure. Therefore, compared with other prior art, this invention restores interior details that other methods cannot restore.

[0115] The fusion result on the activity center image is as follows Figure 3As shown, (a) and (b) display underexposed and overexposed source images, while (c) to (l) list the fusion results of different methods, in order: MEFAW, MEFDSIFT, MTI, MGFF, MSPD-MEF, MEFNet, DIFNet, U2Fusion, DPE-MEF, and the method provided in this invention. It can be seen that MEFAW, MEFDSIFT, and MSPD-MEF produce overly bright results. MTI and MGFF preserve indoor details, but roads and cars are difficult to see. MEFNet also produces locally overexposed areas. DIFNet produces noisy results, and U2Fusion produces unrealistic colors. DPE-MEF recovers details better than other methods. However, due to non-smooth detail enhancement, it also produces noise in the sky and road areas. The comparison shows that the method provided in this invention preserves color and denoises well, resulting in a more natural result; indoor information is recovered, and sky color is well preserved.

[0116] Quantitative comparison was conducted, and the quality of the fusion result was comprehensively evaluated using eight indicators. Among them, the comprehensive indicator based on image features included Q. AB / F AG, EI, SF; structural similarity-based indicators include Q c Q w And MEF-SSIM, the PEM metric is designed based on human visual perception. Specifically, see Table 1, which lists the average values ​​of the above metric for 55 scenarios.

[0117] Table 1: Quantitative results of different methods; the top three best results are shown in bold.

[0118]

[0119] It can be seen that the present invention is in Q AB / F AG, EI, Q c Q w Both SF and MEF-SSIM showed high results. The top three in each metric are shown in bold. Except for MEF-SSIM, all other metrics of this invention are in the top three, and MEF-SSIM is only slightly different from the top three. This also shows that the method provided by this invention is effective in preserving details and color information.

[0120] By comparing the above with existing methods, the multi-exposure image fusion method based on perceptual enhancement structure block decomposition in this invention decomposes image blocks into four components: perceptual gain, signal intensity, signal structure, and average intensity. Furthermore, it proposes an effective perceptual gain estimation function, designs a fusion strategy, and proposes a multi-scale fusion framework to improve fusion performance. As a result, the method provided by this invention is significantly superior to existing methods in terms of perceptual realism, and the information in the fused image is richer.

[0121] The preferred embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the specific details of the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.

[0122] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.

[0123] Furthermore, various different embodiments of the present invention can be combined in any way, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed by the present invention.

Claims

1. A multi-exposure image fusion method based on perceptual enhancement structured block decomposition, characterized in that, Includes the following steps: S1. Acquire multiple exposure source images, and decompose each image block in each exposure source image into sensing gain, signal strength, signal structure, and average intensity; S2. Estimate the perceived gain using the logarithmic stretching method; S3. The sensing gain, signal strength, signal structure and average intensity are fused respectively to obtain fused sensing gain, fused signal strength, fused signal structure and fused average intensity, and the fused image block is reconstructed; S4. Integrate all the fused image blocks to obtain the final fused image; In step S1, obtain K The image of the source of the exposure described by Zhang will be the size of p × p The image patches of ×3 are rearranged into vectors, represented as follows: in, It is the set of real numbers; The image block decomposition is represented by the following formula: in, Indicates the first k The vector after normalization of the image ( k It is an integer, ranging from 1 to... K ), Representing vectors l 2-norm, For the first k Image patches after perceptual gain of the image. For the first k The perceptual gain of the image. For the first k The signal strength of the image, For the first k The signal structure of the image, For the first k The average intensity of the images; In step S2, the perceived gain is estimated using the logarithmic stretching method. The expression is as follows: in, for The correction value, Represents the vector The mean, Indicates the first k The maximum mean value among the image blocks of the image; According to the following formula Make corrections: in, Set to 5 / 255.

2. The multi-exposure image fusion method based on perceptual enhancement structure block decomposition according to claim 1, characterized in that, In step S3, the maximum mean value among each of the image blocks is used. Estimate the fusion sensing gain The formula is as follows: in Indicates the first k The first image j Image blocks.

3. The multi-exposure image fusion method based on perceptual enhancement structure block decomposition according to claim 2, characterized in that, The fused signal strength The maximum value of the signal intensity across all the image blocks is given by the following formula: 。 4. The multi-exposure image fusion method based on perceptual enhancement structure block decomposition according to claim 3, characterized in that, The fused signal structure This is the optimal representation of the structure in the exposure source image, and its formula is as follows: in, It is an exponential parameter. .

5. The multi-exposure image fusion method based on perceptual enhancement structure block decomposition according to claim 4, characterized in that, The average fusion intensity The calculations based on local bias and good exposure of the image patch yielded the following results: Among them, weight Represented as: in, It is a constant. That is the standard deviation.

6. The multi-exposure image fusion method based on perceptual enhancement structure block decomposition according to claim 5, characterized in that, In step S3, the fused image block The reconstruction formula is as follows: in, ; In step S4, the fused image The following formula can be used to synthesize the results: in, This represents the Hadamard product operation. For all It is derived by integrating the source images based on their size and arrangement. for It is derived by integrating the source images based on their size and arrangement. for It is derived by integrating the source images based on their size and arrangement. for The operator is derived by integrating the size and arrangement of the source images. Indicates the kernel size as p × p Mean filter, for The result of mean filtering after aggregation.

7. The multi-exposure image fusion method based on perceptual enhancement structure block decomposition according to claim 6, characterized in that, It also includes step S5: performing multi-scale decomposition on the fused image obtained in step S4 to re-obtain the final fused image.

8. The multi-exposure image fusion method based on perceptual enhancement structure block decomposition according to claim 7, characterized in that, Step S5 includes the following steps: A) Let , Obtain detailed information at the original scale. : in, , and They are respectively at the original scale , and ; B) Perform downsampling by a factor of n to obtain information at the nth scale. ,based on Calculate the nth scale , and The detailed information at the nth scale can be obtained using the following formula. : Where n is an integer, ranging from 2 to Z, and the decomposition process satisfies the iterative condition: in, h and w These are the height and width of the image, respectively; C) Calculate basic information at the Z-scale : D) Using the bilinear upsampling operator The base layer and detail layer are added together to obtain a new blended image, as shown in the following formula: 。