A Multi-Exposure Image Fusion Method Based on Attention Generative Adversarial Network

By introducing attention generation adversarial networks into multi-exposure image fusion, image weights are adaptively selected, and the problem of different methods of manually defining weight calculation in the prior art is solved, and a high-quality multi-exposure image fusion effect is achieved.

CN111429433BActive Publication Date: 2025-06-27XIAN SHUHE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010219045.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-25
Publication Date
2025-06-27
Estimated Expiration
2040-03-25

AI Technical Summary

Technical Problem

The existing multi-exposure image fusion method relies on manually defining the fusion weights, and the calculation methods are different, making it difficult to effectively solve the problem of image dynamic range enhancement.

Method used

The multi-exposure image fusion method based on attention generation adversarial network is adopted. By introducing visual attention mechanisms and residual blocks, the weight of the input image is adaptively selected to realize data-driven adaptive end-to-end multi-exposure image fusion.

Benefits of technology

Adaptive multi-exposure image fusion is realized, the generated image is clear in detail, strong contrast, enhanced dynamic range and improved visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111429433B_ABST
    Figure CN111429433B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-exposure image fusion method based on an attention generative adversarial network. The idea of the attention mechanism highly matches the problem of detail weighting in multi-exposure fusion. Channel attention can be applied to adaptively select the weights of each input image, and spatial attention is used to adaptively select the weights at different spatial positions. This technology has broad application prospects in various multimedia vision fields. The algorithm designs a new attention generative adversarial network for the multi-exposure image fusion task. By introducing the visual attention mechanism into the generative network, it can help the network adaptively learn the weights of different input images and different spatial positions to achieve a better fusion effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital image / video signal processing, and particularly relates to a multi-exposure image fusion method based on an attention generative adversarial network. Background Art

[0002] With the development of computer and multimedia technologies, various multimedia applications have put forward extensive requirements for high-quality images. High-quality images can provide rich information and a realistic visual experience. However, during the image acquisition process, affected by factors such as image acquisition devices, acquisition environments, and noise, the images presented on display terminals are often of low quality. Therefore, how to reconstruct high-quality images from low-quality images has always been a challenge in the field of image processing.

[0003] From bright sunlight to dim starlight, the illumination intensity in natural scenes can span a very large dynamic range, and the brightness contrast can exceed 14 orders of magnitude. Ordinary digital cameras can only capture 8-bit images on each color channel. The existing image brightness level does not match the dynamic range of natural scene brightness. The limited brightness dynamic range restricts the display ability of digital images for high-contrast natural scenes, and problems such as overexposure in bright areas or underexposure in dark areas will occur. Enhancing the image brightness dynamic range can effectively improve the performance ability of images for high-contrast scenes and improve the visual quality of images.

[0004] Various multimedia applications have put forward extensive requirements for high dynamic range images and videos. For example, online video companies hope to improve the subjective quality of video content by increasing the dynamic range of videos, and mobile phone manufacturers will promote the high dynamic range content shooting performance of cameras as a selling point. It can be seen that high dynamic range images have extensive application requirements and important commercial value in the field of visual media.

[0005] Regarding the problem of enhancing the dynamic range of images, the multi-exposure fusion method generates a high dynamic range image with enhanced details by fusing the detail information of different exposure images. Among them, how to select detail information from images with different exposures is a challenging problem.

[0006] By quickly scanning the global image, the human visual system can obtain the target area that needs to be focused on (usually called the focus of attention), and then invest more attention resources in this area to obtain more detail information of the focus area. This is a human ability to quickly screen out high-value information from a large amount of information with limited resources. Human visual attention greatly improves the efficiency and accuracy of visual information processing.

[0007] Inspired by human visual attention, the concept of attention mechanism has been introduced into deep learning. In recent years, deep learning-based methods have achieved great success in many computer vision tasks and some low-level image processing problems. Among them, the attention mechanism has played a role in various applications. We observe that the attention mechanism is suitable for solving the multi-exposure fusion problem and can use the attention mechanism to adaptively select weights.

[0008] The present invention proposes a multi-exposure image fusion method based on an attention generative adversarial network. The idea of the attention mechanism highly matches the problem of detail weighting in multi-exposure fusion. Channel attention can be applied to adaptively select the weights of each input image, and spatial attention is used to adaptively select the weights of different spatial positions. This technology has broad application prospects in various multimedia visual fields. Summary of the Invention

[0009] The purpose of the present invention is to overcome the dependence of existing multi-exposure image fusion methods on manually defined fusion weights and different calculation methods. A multi-exposure fusion method based on an attention generative adversarial network is provided for the problem of image dynamic range enhancement in multi-exposure image fusion methods. The generative adversarial network can use the channel attention mechanism to adaptively determine the weight of each input, and use the spatial attention mechanism to adaptively determine the weight of any spatial position, realizing data-driven adaptive end-to-end multi-exposure image fusion.

[0010] The present invention is implemented by the following technical means:

[0011] A multi-exposure image fusion method based on an attention generative adversarial network. This method adopts the generative adversarial network framework. First, multiple images with different exposures are input into the generative network integrated with the attention mechanism to obtain a multi-exposure fusion image. Then, the fusion image and the target Ground-truth image are sent into the discriminative network for judgment. In the mutual game between the generative network and the discriminative network, a multi-exposure fusion generative network with enhanced details and dynamic range is trained. The overall network of this method is as shown in the appendix Figure 1 shown, which is divided into two parts: the generative network and the discriminative network. The specific generative network is as shown in the appendix Figure 2 shown.

[0012] The present invention introduces a visual attention mechanism into the designed generative network to adaptively select weights and uses residual blocks to extract image detail information.

[0013] The present invention is implemented by the following technical means: A multi-exposure image fusion method based on an attention generative adversarial network, which includes three parts: the construction of the generative adversarial network structure based on the attention mechanism, the adversarial training of the multi-exposure image fusion generative network and the discriminative network, and the multi-exposure image fusion test.

[0014] First, the first part is to build a generative adversarial network based on the attention mechanism. The overall network consists of two parts: a generative network and a discriminative network. The attention mechanism is introduced into the generative network. The specific network construction includes the following steps:

[0015] 1) Construction of the generative network structure

[0016] The structure of the generative network is illustrated in the figure. Figure 2 As shown, it is composed of the fusion of feature extraction and the attention mechanism. The feature extraction part consists of a 3×3 convolution with 32 output channels, a PReLU activation operation, 5 residual block modules with both input and output channels of 32. Then, a 3×3 convolution with 32 output channels and a PReLU activation operation are performed. The resulting feature map is added to the corresponding position of the feature map obtained after the first convolution and activation operation, that is, the feature extraction operation for an image is completed, and 32 feature maps of an image are obtained. At the same time, the same feature extraction is performed on each of the N images in a training pair, and 32 feature maps of N images are obtained. They are concatenated to obtain N×32 feature maps.

[0017] Among them, each residual block operation includes a sequential 1-layer 3×3 convolution, a batch normalization operation, and a PReLU activation. Then, a 1-layer 3×3 convolution and a batch normalization operation are performed. Finally, the resulting feature map of the above operations is added to the corresponding position of the input feature map to obtain the result of one residual module.

[0018] The attention module in the present invention is designed as a cascaded hybrid attention module, that is, first, channel attention operation is performed on the input feature map. The channel attention weights are multiplied by the channel feature map channel by channel to complete the channel attention operation. Then, spatial attention operation is performed on the feature map adjusted by channel attention to calculate the weights of each spatial position, and the weights are multiplied by the feature map element by element to complete the spatial attention operation. Through the sequential operations of channel attention and spatial attention, the hybrid attention operation is completed.

[0019] Among them, the channel attention operation mainly extracts attention parameters based on two pooling operations on the channel plane. The global average pooling (average value) and max pooling (maximum value) of each channel of the input feature map are calculated respectively to obtain feature vectors with the same scale and number of channels as the input feature map. Then, the two feature vectors pass through a multi-layer perceptron with shared weights respectively. After the two feature vectors are linearly added and passed through a sigmoid activation operation, the channel attention result is obtained, that is, the weight of each feature map is obtained. The channel attention weights are multiplied by the corresponding channels to obtain the feature map adjusted by channel attention.

[0020] The spatial attention operation performs Average pooling (average value) and Max pooling (maximum value) on each feature map of all channels in units of spatial positions, and splices them together along the channel dimension to obtain two weight matrices with the same scale as the input feature map; then, performs a 7×7 convolution operation on the obtained feature map to obtain a spatial attention weight matrix with the same scale as the input feature map, that is, obtains the weight of each spatial position. After the channel attention operation, perform an element-wise multiplication of the feature map and the spatial attention weight to complete the hybrid attention operation.

[0021] After the attention operation, perform a 3×3 convolution operation, and obtain the output fusion result through the tanh activation function.

[0022] 2) Construction of the discriminant network structure

[0023] The discriminant network is connected to the generator network. It receives the result of the generator network and the Ground-truth corresponding to the input image of the generator network to determine the authenticity of the two input images. The structural parameters of the discriminant network are shown in Table 1 of the attached drawing description. The discriminant network contains 10 convolutional layers, the filter size is 3×3 for all, and the number of filters increases continuously, from 64 to 1024, doubling every 2 times. In the 2nd - 8th convolutional operation layers, each convolutional layer contains 1 convolutional operation, 1 batch normalization, and 1 LeakyReLU activation. Only the first convolutional layer does not have the batch normalization operation. Next, perform average pooling operation, convolutional operation, LeakyReLU activation, and another convolutional operation on 512 feature maps in sequence, and finally activate the output discriminant result with the sigmoid function.

[0024] The second part is the adversarial training between the multi-exposure image fusion generator network and the discriminant network.

[0025] First is the preparation of training data. The present invention processes the publicly available multi-exposure image dataset to obtain the training dataset. The original dataset contains a total of 589 samples, each sample contains 2 to 7 low-dynamic-range images with different exposures and a corresponding Ground-truth image, and the image size is about 3000×5000 pixels. Based on this dataset, 440 image pairs are selected as the training set, and 3 low-dynamic-range images with different exposures and the corresponding Ground-truth image in each sample are selected as a training sample pair. Downsample and then segment, divide each pair of images into 6 pieces, and finally obtain 2640 pairs of image patches as the training dataset.

[0026] The specific adversarial training method is to alternately train the generator network and the discriminator network. First, perform one training of the generator network using the generator loss, backpropagate, then perform one training of the discriminator network using the discriminator loss, and then backpropagate again, and keep alternating training like this. The total loss function is shown in formula (1):

[0027] min G max D f(G,D), (1)

[0028] To achieve Nash equilibrium and complete the training.

[0029] The loss function of the generator network designed in the present invention consists of four parts, namely the image loss (l mse ), the perceptual loss (l pe ), the adversarial loss (l ad ), and the TV loss (l tv ).

[0030] Adding the four losses in a certain proportion gives the generator network loss. The specific loss function is shown in formula (2):

[0031] l mef =αl mse +βl pe +γl ad +δl tv , (2)

[0032] Finally, it is the multi-exposure image fusion test part.

[0033] The test dataset selects the remaining 97 pairs of data after the training part is selected, performs downsampling processing without cropping, inputs the test program, and generates multi-exposure fusion images. The test program applies the results of the adversarial training of the two-part multi-exposure image fusion generator network and discriminator network, inputs the parameters obtained from the adversarial training into the test program for multi-exposure image fusion, generates multi-exposure fusion images, and evaluates them using subjective visual effects and objective peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) indices.

[0034] The algorithm designs a new attention generative adversarial network for the multi-exposure image fusion task. By introducing the visual attention mechanism into the generator network, it can help the network adaptively learn the weights of different input images and different spatial positions to achieve better fusion effects; BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 、Overall network architecture diagram;

[0036] Figure 2 、Generator network structure diagram;

[0037] Figure 3 1. Structure diagram of the attention module;

[0038] Figure 4 2. Comparison chart of the subjective results between the method of the present invention and the existing methods; in the upper row from left to right are the low, medium, and high exposure input and the label images, and in the lower row from left to right are the results of Li, Ma, Kou, and the method of the present invention Specific implementation manners

[0039] The following describes the embodiments of the present invention with reference to the accompanying drawings of the specification:

[0040] The present invention consists of the following parts. First, a generative adversarial network based on the attention mechanism is built, which is composed of a generative network and a discriminative network, and the attention mechanism is introduced into the generative network. Second, the multi-exposure image fusion generative network and the discriminative network are trained adversarially. The generative network and the discriminative network are alternately trained adversarially with a training sample set to obtain the network parameters of the generative network. Finally, in the test stage, three different exposure images are used as the input in the multi-exposure fusion stage, and the trained generative network is passed through to achieve multi-exposure image fusion. The following introduces the specific process.

[0041] (1) Network construction

[0042] a) Generative network

[0043] The generative network is mainly divided into two stages. The first half is the feature extraction and connection network, and the second half is the attention operation and fusion network.

[0044] In the feature extraction stage, first, a 3×3 convolution is performed, then a PReLU activation operation is carried out, and then five residual block operations are performed. Each residual block operation includes one 3×3 convolution, one batch normalization operation, one PReLU activation operation, another 3×3 convolution and one batch normalization operation, and then the feature map initially input to the residual block and the feature map of the last batch normalization operation are added to obtain the output result of the residual block. After five residual block operations, one 3×3 convolution and one PReLU activation operation are sequentially performed on the obtained 32 feature maps, and then the obtained 32 feature maps are added to the feature map obtained after the first convolution operation, thus completing the feature extraction of each input image and obtaining 32 feature maps.

[0045] In the connection stage, the feature maps obtained from the three different exposure input images input to the network after the feature extraction stage are connected to obtain 3×32 feature maps.

[0046] The attention operation stage is a hybrid attention operation, that is, the channel attention and spatial attention operations are sequentially completed. Among them, the channel attention operation performs pooling (Average pooling and Max pooling operations) on the input feature map for each channel, then passes through a multi-layer perceptron, adds the two feature vectors, and finally obtains the channel attention result through the sigmoid operation. Multiply this result by the input feature map to complete the channel attention operation and obtain a 3×32 feature map. Then, perform spatial pooling on the feature map after the channel attention operation, use the Average pooling and Max pooling operations to compress the feature map in the channel dimension, obtain the spatial attention result for each spatial position, and multiply this result by the input feature map to complete the spatial attention operation. Completing the above two parts of the operation completes the attention operation and obtains 3×32 feature maps.

[0047] In the fusion stage, the 3×32 feature maps after the attention operation are convolved again with a 3×3 convolution, and finally activated by the tanh function for output.

[0048] b) Discriminative network

[0049] The discriminative network is connected behind the generative network and receives the multi-exposure fusion result of the generative network and the Ground-truth corresponding to the input end of the generative network as two inputs, with a size of 400×400×3 pixels, to determine whether the generated image belongs to the same category as the Ground-truth;

[0050] The discrimination network performs the following operations on the input image. First, it performs a convolution with a 3×3 convolution kernel and a stride of 1, followed by a LeakyReLU operation to obtain 64 feature maps. Then, it performs 1 convolution with a 3×3 convolution kernel and a stride of 2, 1 batch normalization operation, and 1 LeakyReLU operation to obtain 64 feature maps. Again, it performs 1 convolution with a 3×3 convolution kernel and a stride of 1, 1 batch normalization operation, and 1 LeakyReLU operation to obtain 128 feature maps. It performs 1 convolution with a 3×3 convolution kernel and a stride of 2, 1 batch normalization operation, and 1 LeakyReLU operation to obtain 128 feature maps. It performs 1 convolution with a 3×3 convolution kernel and a stride of 1, 1 batch normalization operation, and 1 LeakyReLU operation to obtain 256 feature maps. It performs 1 convolution with a 3×3 convolution kernel and a stride of 2, 1 batch normalization operation, and 1 LeakyReLU operation to obtain 256 feature maps. It performs 1 convolution with a 3×3 convolution kernel and a stride of 1, 1 batch normalization operation, and 1 LeakyReLU operation to obtain 512 feature maps. It performs 1 convolution with a 3×3 convolution kernel and a stride of 2, 1 batch normalization operation, and 1 LeakyReLU operation to obtain 512 feature maps. Then, it performs 1 average pooling operation, again performs 1 convolution with a 3×3 convolution kernel and a stride of 1, followed by a LeakyReLU operation, a convolution with a 3×3 convolution kernel and a stride of 1, and finally uses the sigmoid function to obtain the discrimination result. The specific network parameters are shown in Table 1 in the appendix.

[0051] (2) Adversarial training

[0052] a) Training dataset preparation

[0053] The present invention uses a publicly available multi-exposure image dataset, and preprocesses it to obtain a training dataset and a test dataset. The original dataset contains a total of 589 samples, each sample contains 2 to 7 low-dynamic-range images with different exposures and a corresponding Ground-truth image, and the image size is approximately 3000×5000 pixels. Based on this dataset, 440 sample pairs are selected to constitute the training dataset of the present invention. For each sample, 3 low-dynamic-range images with different exposures and the corresponding Ground-truth image are selected as a training sample pair. Then, the spatial resolution of the samples is uniformly downsampled to 1200×800 pixels to retain more details and maximize the contrast of the input images. Each image is correspondingly segmented into image blocks of 400×400 pixels. Therefore, the training dataset contains a total of 2640 pairs of image blocks.

[0054] b) Loss function

[0055] During the training process of the network, the total loss function is shown in formula (1):

[0056] min G max D f(G, D), (1)

[0057] The generation network and the discriminant network are alternately trained to achieve Nash equilibrium.

[0058] The definition of the loss function is crucial for the generation network. The loss function of the generation network designed in the present invention consists of four parts, namely the image loss (l mse ), the perceptual loss (l pe ), the adversarial loss (l ad ), and the TV loss (l tv ).

[0059] Adding the four losses in a certain proportion gives the loss of the generation network. The specific loss function is shown in formula (2):

[0060] l mef = αl mse + βl pe + γl ad + δl tv , (2)

[0061] Among them, α = 1, β = 6×10 -3 , γ = 10 -3 , δ = 2×10 -8 .

[0062] Specifically, l mse is used to calculate the mean square loss between the fusion result generated by the generation network and the Ground-truth, while l pe is the perceptual loss, which is used to calculate the mean square loss between the feature maps obtained after passing the fusion result generated by the generation network and the Ground-truth through the pre-trained VGG network, as shown in formulas (3) and (4):

[0063]

[0064]

[0065] Among them, W and H respectively refer to the width and height dimensions of the input image, F i refers to the fusion result generated by the generation network, GT refers to the Ground-truth corresponding to the input, and V gg (·) corresponds to the operation of the pre-trained VGG network. In the present invention, the output results of the first 30 layers of the pre-trained VGG network are selected for calculation.

[0066] l tv usually takes a very small proportion (not greater than 10 -8This magnitude is used in combination with other losses to suppress noise during the generation process, l tv As shown in Equation (5):

[0067]

[0068] Wherein, u is used to refer to the image for calculation, and D u refers to the support domain of the image.

[0069] c) Adversarial training process

[0070] The specific training process is that the generation network and the discriminator network are alternately performed. That is, for every 1 training of the generation network, followed by backpropagation, and then the discriminator network is updated 1 time, followed by backpropagation. When training, the batch size is set to 1, and the training for about 100 rounds can achieve a convergence effect.

[0071] (3) Multi-exposure fusion test

[0072] The test dataset used in the test process is 97 pairs of data selected from the remaining part of the original dataset after removing the training dataset. It is downsampled to a size of 400×400×3, without cropping. Each pair contains 3 input images with different exposures, and 1 label Ground-truth image used for comparison and evaluation with the generated multi-exposure fusion image.

[0073] During the multi-exposure fusion test process, the generation network part of the generative adversarial network is intercepted. Using the network parameters obtained after the last training in the adversarial training stage, they are loaded into the generation network part. By inputting the test dataset into the network, the multi-exposure fusion image can be obtained. The specific process is to input 3 images with different exposures and a size of 400×400×3 into the generation network. After feature extraction, connection, attention operation, and fusion, a fused image with clear details and enhanced contrast of 400×400×3 is obtained.

[0074] To verify the effectiveness of the present invention, we use subjective visual effects and objective numerical indicators to evaluate the generation fusion effect. The subjective visual effect comparison between the method of the present invention and other existing methods is as attached Figure 4 shown. And the objective indicators use two commonly used image quality evaluation indicators, namely peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). The results are shown in Table 2. From the subjective visual effect, the results of the present invention have clearer details and stronger contrast. From the objective numerical values, the values of the present invention are also higher. Therefore, from both subjective and objective aspects, the results of the present invention are better than the existing methods.

[0075] Table 1 Network parameter table of the discriminator network

[0076]

[0077] Table 2 Comparison Table of Objective Results between the Method of the Present Invention and Existing Methods

[0078] Average PSNR (dB) Average SSIM Li 15.9827 0.5234 Ma 15.9915 0.5350 Kou 16.4513 0.5469 This invention 17.6994 0.5489

Claims

1. A multi-exposure image fusion method based on an attention generative adversarial network, characterized in that: It includes three parts: building a generative adversarial network structure based on the attention mechanism, adversarial training of the multi-exposure image fusion generation network and the discriminative network, and multi-exposure image fusion testing; First, the first part is to build a generative adversarial network based on the attention mechanism. The overall network consists of two parts: a generative network and a discriminative network. The attention mechanism is introduced into the generative network. The specific network construction includes the following steps: 1) Building the generative network structure The generative network structure is composed of feature extraction and attention mechanism fusion. The feature extraction part consists of a 3×3 convolution with 32 output channels, a PReLU activation operation, 5 residual block modules with both input and output channels of 32, and then a 3×3 convolution with 32 output channels and a PReLU activation operation. The resulting feature map is added to the feature map obtained after the first convolution and activation operation at the corresponding positions to complete the feature extraction operation for an image, obtaining 32 feature maps of an image. At the same time, the same feature extraction is performed on each of the N images in a training pair to obtain 32 feature maps of N images, and they are concatenated to obtain N×32 feature maps; Among them, each residual block operation includes a sequential 3×3 convolution layer, a batch normalization operation, and a PReLU activation. Then, a 3×3 convolution layer and a batch normalization operation. Finally, the resulting feature map of the above operations is added to the input feature map at the corresponding positions to obtain the result of one residual module; The attention module is designed as a cascaded hybrid attention module. First, channel attention operation is performed on the input feature map, and the channel attention weights are multiplied by the channel feature map channel by channel to complete the channel attention operation. Then, spatial attention operation is performed on the feature map adjusted by channel attention to calculate the weights of each spatial position, and the weights are multiplied by the feature map element by element to complete the spatial attention operation. Through the sequential operations of channel attention and spatial attention, the hybrid attention operation is completed; Among them, the channel attention operation is to perform two pooling operations based on the channel plane to extract attention parameters. The global average pooling (average value) and max pooling (maximum value) of each channel of the input feature map are calculated respectively to obtain feature vectors with the same scale and number of channels as the input feature map. Then, the two feature vectors pass through a multi-layer perceptron with shared weights respectively. After the two feature vectors are linearly added and passed through a sigmoid activation operation, the channel attention result is obtained, and the weight of each feature map is obtained. The channel attention weights are multiplied by the corresponding channels to obtain the feature map adjusted by channel attention; The spatial attention operation performs Average pooling (average value) and Max pooling (maximum value) on each feature map of all channels in units of spatial positions, and concatenates them together along the channel dimension to obtain two weight matrices with the same scale as the input feature map; then, performs a 7×7 convolution operation on the obtained feature map to obtain a spatial attention weight matrix with the same scale as the input feature map, and obtains the weight of each spatial position; after the channel attention operation, performs an element-wise multiplication of the feature map and the spatial attention weight to complete the hybrid attention operation; After the attention operation, a 3×3 convolution operation is performed, and the output fusion result is obtained through the tanh activation function; 2) Construction of the discriminative network structure The discriminative network is connected to the generative network. It receives the result of the generative network and the Ground-truth corresponding to the input image of the generative network to judge the authenticity of the two input images; the discriminative network contains 10 convolutional layers, the filter size is 3×3 for all, and the number of filters increases continuously, from 64 to 1024, doubling every 2 times; in the 2nd - 8th convolutional operation layers, each convolutional layer contains 1 convolutional operation, 1 batch normalization, and 1 LeakyReLU activation, and only the 1st convolutional layer does not have batch normalization operation; next, performs average pooling operation, convolutional operation, LeakyReLU activation, and another convolutional operation on 512 feature maps in sequence, and finally activates the output discriminative result with the sigmoid function; The second part is the adversarial training of the multi-exposure image fusion generative network and the discriminative network; First is the preparation of training data, downsampling and then segmentation, and each pair of images is divided into 6 pieces; The adversarial training method is to alternately train the generative network and the discriminative network. First, perform one training of the generative network using the generative loss, backpropagation, then perform one training of the discriminative network using the discriminative loss, and then backpropagation, and keep alternating training like this; the total loss function is shown in formula (1): min G max D f(G, D), (1) To achieve the Nash equilibrium and complete the training; The loss function of the designed generation network consists of four parts, namely, the image loss (l mse ), the perceptual loss (l pe ), the adversarial loss (l ad ), and the TV loss (l tv ); Adding the 4 losses in a certain proportion is the generative network loss, and the specific loss function is shown in formula (2): l mef = αl mse + βl pe + γl ad + δl tv , (2) The test dataset selects the remaining data after the training part is selected, performs downsampling processing without cropping, inputs it into the test program to generate multi-exposure fusion images; the test program applies the result of the adversarial training of the second part of the multi-exposure image fusion generative network and the discriminative network, and inputs the parameters obtained from the adversarial training into the test program for multi-exposure image fusion to generate multi-exposure fusion images.