Generative adversarial network for fusion of images, image fusion method and terminal device

By designing a generative adversarial network for multimodal image fusion, using adversarial learning and multimodal hierarchical wavelet fusion modules, the problem of poor multimodal image fusion in the prior art is solved, and high-quality image fusion is achieved.

CN114926382BActive Publication Date: 2025-05-02SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210539173.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-05-02
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

The existing method of fused image has poor fusion effect on multimodal images. The fused image features tend to be single modal features, and high-quality fusion images cannot be obtained.

Method used

A generative adversarial network is designed, including a generator, an edge sensing module, an edge detection unit and multiple discriminators, to constrain the generation of fusion images by adversarial learning relationships, and feature fusion is performed through multimodal hierarchical wavelet fusion module and specific modal information fusion module.

Benefits of technology

The fusion of feature information of different levels and different frequency levels is realized, which avoids the loss of intermediate layer information of different modes, and integrates texture details information through the edge sensing module to obtain higher quality fusion images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926382B_ABST
    Figure CN114926382B_ABST
Patent Text Reader

Abstract

The present invention discloses a generative adversarial network for fusion of images, an image fusion method and a terminal device, wherein the generative adversarial network includes a generator, an edge perception module, a first edge detection unit, a first discriminator, a second discriminator and a third discriminator. In the design of the generator, the present invention constructs a hierarchical multimodal wavelet fusion module, and adaptively performs fusion processing on different feature attributes to achieve fusion of feature information of different levels and frequencies, thereby avoiding the loss of information in the middle layer of different modalities; by constructing an edge perception module, the edge information of different modal data is integrated, the texture detail information representation capability is increased, and by jointly learning the adversarial relationship between the fused image and the two original input images and the fused edge image and the original edge image, the final fused image is prompted to not only contain the intensity information of the original image, but also avoid the loss of edge texture detail information, thereby obtaining a higher quality fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and in particular to a generative adversarial network for fusion of images, an image fusion method and a terminal device. Background Art

[0002] Image fusion technology can make up for the limitations of a single sensor, and integrate the complementary information of multiple source images captured by different sensors or optical devices to obtain an enhanced image containing rich information, which promotes the rapid development of this technology in the fields of industry, civil, military, and medical. For example, in medical imaging, the fusion of different modal images (such as PET and MRI) can improve the accuracy of medical diagnosis, and in the military field, the fusion of infrared and visible light images can achieve the clarity of night surveillance. According to the imaging characteristics of medical images and infrared and visible light images, the characteristics of infrared or PET images can be described by the intensity distribution of pixels, and the detailed texture information of visible light or MRI images can be characterized by image gradients. Therefore, medical images and infrared images have similar modal characteristics, and infrared and visible light images and medical images can be classified as multimodal images, thereby constructing a multimodal image fusion framework with strong generalization and improving its practical value.

[0003] However, existing image fusion methods have poor fusion effects on multimodal images, and the fused image features tend to be single-modal features, so high-quality fused images cannot be obtained.

[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide a generative adversarial network, an image fusion method and a terminal device for fusing images, aiming to solve the problem that the existing image fusion method has poor fusion effect on multimodal images, the fused image features tend to be single modal features, and high-quality fused images cannot be obtained.

[0006] The technical solution of the present invention is as follows:

[0007] A generative adversarial network for fusion of images, comprising a generator, an edge perception module, a first edge detection unit, a first discriminator, a second discriminator and a third discriminator;

[0008] The generator is used to fuse the first modality image and the second modality image to generate a fused image;

[0009] The edge perception module is used to extract and combine edge information of the first modality image and the second modality image to generate an original edge image;

[0010] The first edge detection unit is used to extract and combine edge information of the fused image to generate a fused edge image;

[0011] The first discriminator is used to obtain a probability P1 that the fused image is the first modality image;

[0012] The second discriminator is used to obtain a probability P2 that the fused image is the second modality image;

[0013] The third discriminator is used to obtain the probability P3 that the fused edge image is the original edge image;

[0014] The generator is further configured to perform image fusion on the first modality image and the second modality image again when one or more of the probabilities P1, P2 and P3 are less than a threshold probability, and to output the fused image when the probabilities P1, P2 and P3 are all greater than a threshold probability.

[0015] The generative adversarial network for image fusion, wherein the generator includes a specific modality information fusion module and a multimodal hierarchical wavelet fusion module, the specific modality information fusion module includes a first encoder and a first decoder; the multimodal hierarchical wavelet fusion module includes a second encoder and a second decoder; the first encoder includes four convolution blocks, each convolution block consists of two convolution sub-blocks, each convolution sub-block consists of a convolution layer, a batch normalization processing layer and a leaky ReLU function layer; the first decoder includes 6 convolution blocks, wherein the first 5 convolution blocks consist of a convolution layer, a pooling layer and a Relu layer, and the last convolution block consists of a convolution layer and a Tanh function; the second encoder includes four inter-layer feature fusion modules, each inter-layer feature fusion module consists of a DWT unit, a feature selection unit, an IDWT unit, and a feature superposition unit; the second decoder includes 6 convolution blocks, wherein the first 5 convolution blocks consist of a convolution layer, a pooling layer and a Relu layer, and the last convolution block consists of a convolution layer and a Tanh function.

[0016] The generative adversarial network for fusion of images, wherein the feature selection unit consists of a cascade layer, a first 1*1 convolutional layer, a global average pooling layer, a fully connected layer, a Softmax function layer, two weighted layers and a second 1*1 convolutional layer.

[0017] The generative adversarial network for image fusion, wherein the edge perception module includes a Gaussian filtering unit, a second edge detection unit and a Max function.

[0018] The generative adversarial network for image fusion, wherein the first modality image is an infrared image, and the second modality image is a visible light image; or, the first modality image is a PET image, and the second modality image is an MRI image.

[0019] An image fusion method based on a generative adversarial network, comprising the steps of:

[0020] Inputting the first modality image and the second modality image into a generator for image fusion to generate a fused image;

[0021] Inputting the first modality image and the second modality image into an edge perception module to extract and combine edge information to generate an original edge image;

[0022] Inputting the fused image into a first edge detection unit to extract and combine edge information to generate a fused edge image;

[0023] Inputting the fused image and the first modality image into a first discriminator to obtain a probability P1 that the fused image is the first modality image;

[0024] Inputting the fused image and the second modality image into a second discriminator to obtain a probability P2 that the fused image is the second modality image;

[0025] Inputting the fused edge image and the original edge image into a third discriminator to obtain a probability P3 that the fused edge image is the original edge image;

[0026] Compare the probability P1, probability P2 and probability P3 with the threshold probability. If at least one of the probability P1, probability P2 and probability P3 is less than the threshold probability, adjust the generator parameters, and input the first modality image and the second modality image into the generator again to train the generator until the probability P1, probability P2 and probability P3 are all greater than the threshold probability, and complete the training of the generator to obtain a trained generator;

[0027] The original image to be fused is input into the trained generator, and the fused image is input.

[0028] A storage medium, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the image fusion method of the present invention.

[0029] A terminal device, comprising: a processor, a memory and a communication bus; the memory stores a computer-readable program executable by the processor;

[0030] The communication bus realizes the connection and communication between the processor and the memory;

[0031] When the processor executes the computer-readable program, the steps in the image fusion method of the present invention are implemented.

[0032] Beneficial effect: The generative adversarial network provided by the present invention forms an adversarial learning relationship between the generator and the three discriminators, which is used to constrain the generation of fused images. In the design of the generator, a hierarchical multimodal wavelet fusion module is constructed according to discrete wavelet transform, and different feature attributes are adaptively fused to achieve the fusion of feature information of different levels and frequencies, thereby avoiding the loss of information in the middle layers of different modalities; in addition, by constructing an edge perception module, the edge information of different modal data is integrated to increase the ability to represent texture detail information. By combining the fused image with the two original input images and the adversarial learning relationship between the fused edge image and the original edge image, the final fused image not only contains the intensity information of the original image, but also avoids the loss of edge texture detail information, thereby obtaining a higher quality fused image. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 The framework diagram of the generative adversarial network of the present invention, wherein I1 is the first modality image, I2 is the second modality image, and I f is the fused image generated by the generator, I e To fuse edge images, I ef is the original edge image, I e1 and I e2 is the edge detection mapping, and B() is the edge detection operator.

[0034] Figure 2 This is a framework diagram of the generator in the generative adversarial network of the present invention, in which DWT (discrete wavelet transform) is used for feature decomposition and IDWT (inverse discrete wavelet transform) is used for feature reconstruction. are the image features of different convolutional layers, is the intra-layer multimodal fusion feature, is the fusion feature of the encoding network, r = 1, 2, 3, 4, I f is the final fusion image of multimodal hierarchical wavelet features, I′ f It is an auxiliary fusion image in the modality-specific information fusion network.

[0035] Figure 3 It is a framework diagram of the feature selection unit in the present invention, wherein, Hook multi-modal low-frequency features, is the low-frequency fusion feature of the rth layer, ω1 and ω2 are weight mappings.

[0036] Figure 4 It is a schematic diagram of the terminal device of the present invention. DETAILED DESCRIPTION

[0037] The present invention provides a generative adversarial network, an image fusion method and a terminal device for fusion images. In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] In the design of image fusion algorithm, feature extraction and fusion rule design are two key factors affecting the quality of fused image. Currently, researchers mainly use two algorithms for feature extraction and fusion: traditional fusion algorithm and deep learning algorithm. Traditional framework algorithm is the mainstream fusion algorithm (such as conventional multi-scale decomposition algorithm, sparse representation algorithm, saliency algorithm, edge-preserving filtering algorithm, etc.). Although this type of algorithm can improve the fusion effect of the algorithm, it requires manual design of complex fusion rules and salient feature extraction schemes in the construction of the algorithm, which increases the complexity and limitations of the algorithm to a certain extent.

[0039] At present, the fusion algorithms based on deep learning generally include CNN-based algorithms, codec-based algorithms, and generative adversarial network-based algorithms. Among them, CNN-based algorithms are mostly used for multi-focus and multi-exposure image fusion, and the latter two are mostly used for multi-modal image fusion. Although the fusion performance of this type of algorithm is better than that of traditional algorithms, there are still some shortcomings in the implementation of the algorithm. Most CNN-based methods extract weight mappings through the design of convolutional layers. This type of algorithm is usually combined with a multi-scale decomposition algorithm as a feature extraction tool for traditional algorithms, and cannot achieve end-to-end network design. In contrast, the codec-based algorithm obtains deep features through the encoder and fuses multi-modal features through simple weighting and / or cascade operations, and then reconstructs the fusion output according to the decoder. However, this type of algorithm only generates fusion results through simple loss function design constraints, which reduces the quality of the fused image. At present, GAN-based algorithms usually use multi-modal cascade data as network input, and use single-modal images as real images and generated images for discrimination processing. This type of algorithm often causes the generated image features to be biased towards single-modal features, and cannot obtain high-quality fused images.

[0040] Based on this, the present invention provides a generative adversarial network for fusion images, such as Figure 1-2 As shown, the generative adversarial network includes a generator 10, an edge perception module 20, a first edge detection unit 30, a first discriminator 40, a second discriminator 50 and a third discriminator 60;

[0041] The generator 10 is used to fuse the first modality image and the second modality image to generate a fused image;

[0042] The edge perception module 20 is used to extract and combine edge information of the first modality image and the second modality image to generate an original edge image;

[0043] The first edge detection unit 30 is used to extract and combine edge information of the fused image to generate a fused edge image;

[0044] The first discriminator 40 is used to obtain a probability P1 that the fused image is the first modality image;

[0045] The second discriminator 50 is used to obtain a probability P2 that the fused image is the second modality image;

[0046] The third discriminator 60 is used to obtain the probability P3 that the fused edge image is the original edge image;

[0047] The generator 10 is further configured to perform image fusion on the first modality image and the second modality image again when one or more of the probability P1, probability P2 and probability P3 is less than a threshold probability, and to output the fused image when the probability P1, probability P2 and probability P3 are all greater than a threshold probability.

[0048] Specifically, the purpose of image fusion technology is to obtain a fused image containing specific targets or detail information by learning data of different modalities. However, most current networks based on deep learning focus more on the intensity information between image pixels, and consider less the texture detail information of the image. In addition, the influence of the intermediate layer features is ignored during the algorithm implementation process, resulting in the loss of inter-layer detail information, which makes the fused image quality poor. The generative adversarial network provided by the present invention forms an adversarial learning relationship between the generator and the three discriminators, which is used to constrain the generation of the fused image. In the design of the generator, a hierarchical multimodal wavelet fusion module is constructed according to the discrete wavelet transform, and different feature attributes are adaptively fused to achieve the fusion of feature information of different levels and frequencies, thereby avoiding the loss of intermediate layer information of different modalities; in addition, the edge perception module is constructed to achieve the integration of edge information of different modal data, increase the representation capability of texture detail information, and promote the final fused image to contain not only the intensity information of the original image, but also the loss of edge texture detail information by combining the fused image with the two original input images and the adversarial learning relationship between the fused edge image and the original edge image, thereby obtaining a higher quality fused image. The present invention conducts experiments on two types of public multimodal image datasets and shows that the proposed generative adversarial network is superior to the existing fusion network in both subjective and objective evaluations and has good application potential.

[0049] In some embodiments, the generator includes a specific modal information fusion module and a multimodal layered wavelet fusion module, the specific modal information fusion module includes a first encoder and a first decoder; the multimodal layered wavelet fusion module includes a second encoder and a second decoder; the first encoder includes four convolution blocks, each convolution block consists of two convolution sub-blocks, and each convolution sub-block consists of a convolution layer, a batch normalization processing layer, and a Leaky ReLU function layer; the first decoder includes 6 convolution blocks, among which the first 5 convolution blocks consist of a convolution layer, a pooling layer, and a Relu layer, and the last convolution block consists of a convolution layer and a Tanh function; the second encoder includes four inter-layer feature fusion modules, each inter-layer feature fusion module consists of a DWT unit, a feature selection unit, an IDWT unit, and a feature superposition unit; the second decoder includes 6 convolution blocks, among which the first 5 convolution blocks consist of a convolution layer, a pooling layer, and a Relu layer, and the last convolution block consists of a convolution layer and a Tanh function.

[0050] Specifically, if Figure 2 As shown, the network structure of the generator uses the encoder-decoder network as the basic construction route, which mainly includes a specific modality information fusion module and a multimodal hierarchical wavelet fusion module, wherein the output of the multimodal hierarchical wavelet fusion module is used as the fusion output of the multimodal image, and the specific modality information fusion module is used as a constraint item to correct the fusion output. The two module branches are interconnected through the encoder network and are independent through the decoder network. The difference lies in the design of the encoder network.

[0051] In the specific modality information fusion module, the first encoder can be used to obtain different modal images. In the convolution sub-blocks of the first encoder, the kernel size of each convolution layer is 3 and the step size is 1. In order to avoid information loss during the sampling process, the pooling parameters of all convolution layers are set to 1 in the network settings. At the same time, in order to avoid the gradient diffusion problem, batch normalization is performed after the convolution layer, and Leaky ReLU is used as the activation function to avoid gradient sparsity. After obtaining the image features through the first encoder, the features are globally averaged pooled to obtain the feature vectors corresponding to the different modal data, and after a 1*1 convolution operation, the sigmoid function is used to calculate the weight vectors corresponding to the different modal data, and the fusion features of the multi-modal data are obtained by adaptive weighting and weighting. Finally, the first decoder performs 6 layers of convolution block processing on the obtained fusion features to reconstruct the fused image I′. f ,The first five convolutional blocks are composed of convolutional layers, pooling layers, and ReLU layers, and the last convolutional block is composed of convolutional layers and Tanh function layers.

[0052] In the multimodal hierarchical wavelet fusion module, four inter-layer feature fusion modules are constructed to achieve the fusion of information flows between different modal layers. By simply setting adaptive fusion rules, inter-layer connections between multimodal features are established to make up for the problem of information loss in a single information flow. Specifically, this embodiment obtains the image features corresponding to each convolution layer through the second encoder part in the specific modality information fusion module. Where r = 1, 2, 3, 4, and the features within the layer Execute DWT once respectively to transform the features within the layer Decomposed into a corresponding low-frequency component feature And 3 high frequency component characteristics (m represents the modal number, and its value is 1, 2), realizing the conversion of spatial domain features to transform domain features.

[0053] Because different frequency domain features focus on the expression of different information, this embodiment adopts the principle of taking the largest absolute value for the high-frequency features within the layer. At the same time, this embodiment sets a feature selection unit (such as Figure 3 The feature selection unit is composed of a cascade layer, a first 1*1 convolution layer, a global average pooling layer, a fully connected layer, a Softmax function layer, two weighted layers and a second 1*1 convolution layer. The feature selection unit combines the low-frequency features of the rth layer into a single layer. As the input of the inter-layer feature fusion module, the preliminary fusion of multi-modal low-frequency features is achieved through cascading and 1*1 convolution operations; the preliminary fused features are globally averaged pooled to obtain the corresponding global information, and the compressed information is fully connected twice to reduce the number of parameters and calculations while fitting the complex correlation between channels; then two weight maps ω1 and ω2 are generated through the Softmax function and combined with Perform weighted sum and 1*1 convolution operations to obtain the final low-frequency fusion features IDWT is then used to integrate the low-frequency and high-frequency component features to achieve effective fusion of multimodal features within the layer.

[0054] After multimodal fusion of layered features, this study uses a 1*1 convolution kernel to fusion features of the previous layer. Perform dimension expansion and combine with the current layer features The superposition is performed to achieve information fusion of inter-layer features, avoid the loss of intermediate layer information, and obtain the final encoder fusion features And the final fusion feature Perform the same decoding operation to obtain the final fused image. f with I′ fThe mean square error loss between them constrains the training of the hierarchical wavelet fusion network to optimize the final fused image.

[0055] In some embodiments, Figure 1 As shown, for the multimodal image fusion task, the present invention uses the original multimodal data I1 and I2 containing different targets and structures as the input of the generator G to generate a fused image I that can be used to fool the discriminator. f Then I1 and I f As the input of the discriminator D1, I2 and I f As the input of the discriminator D2, it is used to generate two probabilities that can reflect whether the input data is from real data or generated data. At the same time, in order to use I f The edge structure information of the original image may be included. This embodiment realizes the extraction of the joint edge information of the original image by constructing an edge perception module B, and uses it and the edge information of the fused image as a set of adversarial constraints. Through adversarial learning, the generated image can retain the rich edge features in the original image. This embodiment realizes the integration of edge information of different modal data by constructing an edge perception module, increases the ability to represent texture detail information, and promotes the final fused image to contain not only the intensity information of the original image, but also avoid the loss of edge texture detail information by jointly combining the fused image with the two original input images and the adversarial learning relationship between the fused edge image and the original edge image, thereby obtaining a higher quality fused image.

[0056] Specifically, if Figure 1 As shown, the edge perception module includes a Gaussian filter unit, a second edge detection unit and a Max function. Edge perception module I ef The algorithm implementation process is as follows:

[0057] Input data: Input images I1, I2

[0058] Output data: Joint edge data I ef .

[0059] for n≤max_epoch do

[0060] #Original image base layer

[0061] 1:I s1 =Gaussian_filter(I1);

[0062] I s2 =Gaussian_filter(I2);

[0063] #Original image detail layer

[0064] 2:Id1 =I1-I s1 ;

[0065] I d2 =I2-I s2 ;

[0066] # Original image edge layer

[0067] 3:I e1 =Sobel_filter(I d1 );

[0068] I e2 =Sobel_filter(I d2 );

[0069] #Joint edge data of the original image

[0070] 4:I ef =Lmax_function(I e1 , I e2 );

[0071] end for

[0072] return I ef .

[0073] In some embodiments, the first modality image is an infrared image, and the second modality image is a visible light image; or, the first modality image is a PET image, and the second modality image is an MRI image.

[0074] In some embodiments, the generative adversarial network includes three discriminators, which respectively implement adversarial training with the generator from the perspective of the original image and edge detail information. The first discriminator D1 is used to distinguish the fused image from the first modality data (such as infrared, PET, SPECT), the second discriminator D2 is used to distinguish the fused image from the second modality data (such as visible light, MRT1, MRGAD), and the third discriminator D3 is used to discriminate the edge map. The three discriminators included use the same structure but do not share parameters. Since the task difficulty of the discriminator is lower than the difficulty of generating images using neural networks, the structure of the discriminator selected in this embodiment is relatively simple. The discriminator used includes 5 convolutional layers, of which the first four convolutional blocks include convolutional layers, batch normalization layers, and LeakyReLU activation function layers with a slope of 0.2. The step size of the first four convolutional layers is 2, and the step size of the last convolutional layer is 1.

[0075] In some embodiments, a method for image fusion based on a generative adversarial network is also provided, which comprises the steps of:

[0076] Inputting the first modality image and the second modality image into a generator for image fusion to generate a fused image;

[0077] Inputting the first modality image and the second modality image into an edge perception module to extract and combine edge information to generate an original edge image;

[0078] Inputting the fused image into a first edge detection unit to extract and combine edge information to generate a fused edge image;

[0079] Inputting the fused image and the first modality image into a first discriminator to obtain a probability P1 that the fused image is the first modality image;

[0080] Inputting the fused image and the second modality image into a second discriminator to obtain a probability P2 that the fused image is the second modality image;

[0081] Inputting the fused edge image and the original edge image into a third discriminator to obtain a probability P3 that the fused edge image is the original edge image;

[0082] Compare the probability P1, probability P2 and probability P3 with the threshold probability. If at least one of the probability P1, probability P2 and probability P3 is less than the threshold probability, adjust the generator parameters, and input the first modality image and the second modality image into the generator again to train the generator until the probability P1, probability P2 and probability P3 are all greater than the threshold probability, and complete the training of the generator to obtain a trained generator;

[0083] The original image to be fused is input into the trained generator, and the fused image is input.

[0084] The generative adversarial network provided by the present invention forms an adversarial learning relationship between the generator and the three discriminators, which is used to constrain the generation of fused images. In the design of the generator, a hierarchical multimodal wavelet fusion module is constructed according to discrete wavelet transform, and different feature attributes are adaptively fused to achieve the fusion of feature information of different levels and frequencies, thereby avoiding the loss of information in the middle layer of different modalities; in addition, by constructing an edge perception module, the edge information of different modal data is integrated to increase the ability to represent texture detail information, and by combining the fused image with the two original input images and the adversarial learning relationship between the fused edge image and the original edge image, the final fused image not only contains the intensity information of the original image, but also avoids the loss of edge texture detail information, thereby obtaining a higher quality fused image.

[0085] The present invention conducts experiments on two types of public multimodal image datasets and shows that the proposed generative adversarial network is superior to the existing fusion network in both subjective and objective evaluations and has good application potential. The specific experiments are as follows:

[0086] Dataset 1: Two public datasets, TNO Human Factors dataset1 and RoadScene2, are selected to verify the effectiveness of the proposed method in infrared and visible light images. The TNO dataset contains matched multispectral night images of different military-related scenes (such as near infrared, long-wave infrared, and thermal infrared). RoadScene is a registered dataset that contains rich scenes such as roads, vehicles, and pedestrians. In the TNO dataset, we select a total of 137 pairs of images from video sequences and pictures for training. In RoadScene, a total of 43 pairs of grayscale images and 43 pairs of color images are selected for network training. In addition, we collect 232 sets of data through infrared and visible light devices, including 11 scenes (such as construction sites, villages, urban roads, rural roads, etc.) as training data, and register all training data for fusion experiments. In the experiment, all input data are normalized and randomly cropped to 88*88 image size as input for the training process.

[0087] Dataset 2: The Whole Brain Atlas dataset is selected to verify the effectiveness of the proposed algorithm in multimodal medical images. The Whole Brain Atlas dataset contains MRI images of the normal brain and multimodal brain atlas data of different pathological diseases (cerebrovascular disease, brain tumors, Alzheimer's disease). We selected 1063 pairs of brain images of different diseases as training datasets, including CT, MRI, PET, SPECT and other modal data. In order to increase the number of training samples, we first expanded the dataset by rotating the image direction, and normalized all input data in the experiment, and randomly cropped them to 128*128 image size as input for the training process.

[0088] The parameters in the multi-discriminator generative adversarial network are updated by AdamOptimizer. In the network setting, the original learning rate is set to 0.001 in the first 100 echops and linearly decays in the remaining epochs. The batch size is set to k, k=4, and the number of epochs is set to 500. The implementation platform of the proposed algorithm is Intel(R)Xeon(R)CPU F5-2620v4@2.10GHz and GPU NVIDIATitan Xp. The training and testing of the network are implemented on PyTorch.

[0089] This embodiment uses structural similarity (SSIM), weighted fusion quality index (Q W ), nonlinear correlation information entropy (NCIE) is used to evaluate the performance of the generative adversarial network. Its calculation method is as follows:

[0090]

[0091] Where C1 and C2 are constants. and is the variance of the fused image F and the input image X, σ FX is the covariance of the fused image F and the input image X.

[0092]

[0093] Where c(ω) is the importance weight of each local window, λ(ω) is the local window weight, and ω is the local window.

[0094] NCIE(X,Y)=H'(X)+H'(Y)-H'(X,Y)#(3)

[0095]

[0096]

[0097] where p i,l Represents the normalized joint grayscale histogram between the original image and the fused image.

[0098] In some embodiments, a storage medium is further provided, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the image fusion method as described in the present invention.

[0099] In some implementations, the present application also provides a terminal device, such as Figure 4 As shown, it includes at least one processor (processor) 20; display screen 21; and memory (memory) 22, and may also include a communications interface (Communications Interface) 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22 and the communication interface 23 can communicate with each other through the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logic instructions in the memory 22 to execute the method in the above embodiment.

[0100] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0101] The memory 22 is a computer-readable storage medium that can be configured to store software programs, computer executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, implementing the methods in the above embodiments.

[0102] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, a variety of media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, may also be a transient storage medium.

[0103] In addition, the specific process of loading and executing multiple instruction processors in the storage medium and the terminal device has been described in detail in the above method and will not be described here one by one.

[0104] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A generative adversarial network for image fusion, characterized in that: It includes a generator, an edge perception module, a first edge detection unit, a first discriminator, a second discriminator and a third discriminator; The generator is used to fuse the first modality image and the second modality image to generate a fused image; The edge perception module is used to extract and combine edge information of the first modality image and the second modality image to generate an original edge image; The first edge detection unit is used to extract and combine edge information of the fused image to generate a fused edge image; The first discriminator is used to obtain a probability P1 that the fused image is the first modality image; The second discriminator is used to obtain a probability P2 that the fused image is the second modality image; The third discriminator is used to obtain the probability P3 that the fused edge image is the original edge image; The generator is further configured to perform image fusion on the first modality image and the second modality image again when one or more of the probability P1, probability P2 and probability P3 is less than a threshold probability, and output the fused image when the probability P1, probability P2 and probability P3 are all greater than a threshold probability; The generator includes a specific modal information fusion module and a multimodal hierarchical wavelet fusion module, the specific modal information fusion module includes a first encoder and a first decoder; the multimodal hierarchical wavelet fusion module includes a second encoder and a second decoder; the first encoder includes four convolution blocks, each of which consists of two convolution sub-blocks, and each convolution sub-block consists of a convolution layer, a batch normalization processing layer, and a leaky ReLU function layer; the first decoder includes 6 convolution blocks, wherein the first 5 convolution blocks consist of a convolution layer, a pooling layer, and a Relu layer, and the last convolution block consists of a convolution layer and a Tanh function; the second encoder includes four inter-layer feature fusion modules, each of which consists of a DWT unit, a feature selection unit, an IDWT unit, and a feature superposition unit; the second decoder includes 6 convolution blocks, wherein the first 5 convolution blocks consist of a convolution layer, a pooling layer, and a Relu layer, and the last convolution block consists of a convolution layer and a Tanh function.

2. The generative adversarial network for image fusion according to claim 1, characterized in that: The feature selection unit consists of a cascade layer, a first 1*1 convolutional layer, a global average pooling layer, a fully connected layer, a Softmax function layer, two weighted layers and a second 1*1 convolutional layer.

3. The generative adversarial network for image fusion according to claim 1, characterized in that: The edge perception module includes a Gaussian filtering unit, a second edge detection unit and a Max function.

4. The generative adversarial network for image fusion according to any one of claims 1 to 3, characterized in that: The first modality image is an infrared image, and the second modality image is a visible light image; or, the first modality image is a PET image, and the second modality image is an MRI image.

5. An image fusion method based on the generative adversarial network described in any one of claims 1 to 4, characterized in that: Includes steps: Inputting the first modality image and the second modality image into a generator for image fusion to generate a fused image; Inputting the first modality image and the second modality image into an edge perception module to extract and combine edge information to generate an original edge image; Inputting the fused image into a first edge detection unit to extract and combine edge information to generate a fused edge image; Inputting the fused image and the first modality image into a first discriminator to obtain a probability P1 that the fused image is the first modality image; Inputting the fused image and the second modality image into a second discriminator to obtain a probability P2 that the fused image is the second modality image; Inputting the fused edge image and the original edge image into a third discriminator to obtain a probability P3 that the fused edge image is the original edge image; Compare the probability P1, probability P2 and probability P3 with the threshold probability. If at least one of the probability P1, probability P2 and probability P3 is less than the threshold probability, adjust the generator parameters, and input the first modality image and the second modality image into the generator again to train the generator, until the probability P1, probability P2 and probability P3 are all greater than the threshold probability, the training of the generator is completed, and a trained generator is obtained; The original image to be fused is input into the trained generator, and the fused image is input.

6. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the image fusion method as claimed in claim 5.

7. A terminal device, characterized in that: include: A processor, a memory and a communication bus; the memory stores a computer-readable program that can be executed by the processor; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, the steps in the image fusion method according to claim 5 are implemented.

Citation Information

Patent Citations

  • Multi-modal image fusion method based on generative adversarial network and super-resolution network

    CN109325931A

  • Image enhancement method fusing two-dimensional discrete wavelet transform and generative adversarial network

    CN111275640A