A Color Polarization Image Fusion Method Improved Based on PFNET

By improving the PFNET model, the shift window transform feature extraction network and dual discriminator module are introduced, and the problem of unbalanced feature extraction and information retention in color polarized image fusion is solved, and the fusion image generation with high polarization resolution and rich texture is achieved, which improves image quality and information volume.

CN119850444BActive Publication Date: 2025-07-04CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328911.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing color polarized image fusion methods have limitations in feature extraction capabilities, uneven information retention and loss function design, resulting in poor quality of the fusion image and making it difficult to achieve the combination of high polarization resolution and rich textures.

Method used

Improve the PFNET model, introduce the encoder module and dual discriminator module built by shift window transformation feature extraction network, combine the total loss function of the Gan adversarial network, and iteratively train the generator module to achieve deep fusion of global and local features and information retention equalization.

Benefits of technology

It improves the integrity of feature expression and the quality of the fusion image, can better retain multi-source complementary information, improves the visual effect and information volume of the image, and meets the high-quality fusion needs of multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850444B_ABST
    Figure CN119850444B_ABST
Patent Text Reader

Abstract

A Color Polarization Image Fusion Method Improved Based on PFNET. It relates to the field of color polarization image fusion, specifically to the technical field of color polarization image fusion improved based on PFNET. While realizing the fusion of color polarization images, the present invention improves the integrity of feature expression, and has both high polarization resolution and rich texture. The method includes the following steps: dividing the data set into a training set, a test set and a validation data set; improving the PFNET model: introducing the encoder module in the fusion image generator constructed by the shifted window transform feature extraction network into the PFNET model; introducing the dual discriminator module constructed by the Gan adversarial network into the PFNET model; introducing the total loss of the Gan adversarial network into the PFNET model; inputting the training set into the improved PFNET model to obtain a target model: inputting the test set into the target model to obtain an executable PFNET improved model; inputting the validation data set into the executable PFNET improved model to obtain a fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of color polarization image fusion, and specifically to the technical field of color polarization image fusion improved based on PFNET. Background Art

[0002] In the current process of technological development, color polarization image fusion technology has shown significant important application value in key fields such as remote sensing detection and target recognition. It provides strong technical support for many complex tasks and can help obtain more comprehensive and accurate image information.

[0003] However, most of the currently adopted color polarization image fusion methods are built on the basis of traditional convolutional neural networks (CNNs) or basic generative adversarial networks (GANs). Although these methods have achieved certain results to some extent, there are still some significant defects that need to be solved urgently:

[0004] 1. Limited feature extraction ability: Traditional CNNs mainly rely on local convolution operations, which makes it difficult for them to capture global context information of images. The lack of global information directly leads to difficulties in achieving good integration of polarization characteristics (such as DoLP) and texture details in the fusion results, seriously affecting the quality of the fused images.

[0005] 2. Unbalanced information retention: Most existing GANs adopt a single discriminator design, which is prone to bias towards a certain type of source image during training, such as intensity images or degree of linear polarization images. Once there is a bias, it will cause the loss of complementary information, making the fused image unable to retain the advantageous features of multi-source data simultaneously and reducing the application value of the fused image.

[0006] 3. Limitations of the loss function: Traditional loss functions, typically represented by mean square error, overly focus on pixel-level alignment but lack high-level semantic constraints such as structural similarity and visual fidelity. This results in problems such as low contrast and blurred edges in the fused images, greatly affecting the visual effect and subsequent analysis of the images. Summary of the Invention

[0007] In view of the above problems, the object of the present invention is to propose a color polarization image fusion method improved based on PFNET. By improving the PFNET model, while achieving the fusion of color polarization images, the integrity of feature expression is enhanced, with both high polarization resolution and rich texture.

[0008] The method includes the following steps:

[0009] S1. Obtain a color polarization image dataset and perform preprocessing. Divide the preprocessed color polarization image dataset into a training set, a test set, and a validation dataset according to the ratio of 6:2:2.

[0010] S2. Improve the PFNET model. Specifically:

[0011] S21. Replace the original generator module with an improved generator module. The improved generator module is obtained by introducing the encoder module in the fusion image generator constructed by the shifted window transform feature extraction network into the original generator module of the PFNET model. The improved generator module includes an upper encoder module and a lower encoder module.

[0012] S22. Introduce a dual discriminator module constructed by the Gan adversarial network into the PFNET model. The dual discriminator module includes discriminator 1 and discriminator 2.

[0013] S23. Introduce the total loss of the Gan adversarial network into the loss function of the PFNET model to obtain an improved loss function. The improved loss function is specifically: where represents the total loss of the improved generator module, and represent the losses of discriminator 1 and discriminator 2 respectively.

[0014] S3. Input the training set into the improved PFNET model for iterative training. When the number of training times is greater than 100, and the total loss of the generator module and the losses of discriminator 1 and discriminator 2 no longer decrease significantly, obtain the target model.

[0015] S4. Input the test set into the target model for testing, and obtain performance test metrics: information entropy, standard deviation, visual information fidelity, structural similarity index, and mutual information. When the performance test metrics reach a fixed threshold, obtain an executable improved PFNET model.

[0016] S5. Input the validation dataset into the executable improved PFNET model to obtain an RGB fusion image.

[0017] Furthermore, the preprocessing is specifically: using the Stokes vector to describe the polarization characteristics of the color polarization image dataset; the preprocessed color polarization image dataset includes: an intensity image and a degree of linear polarization image DOLP.

[0018] Furthermore, the upper encoder module, from input to output, sequentially passes through an image block module, a first-stage layer module, a second-stage layer module, a third-stage layer module, and a fourth-stage layer module. The structure of the lower encoder module is the same as that of the upper encoder module.

[0019] The first-stage layer module includes: a linear embedding module and two shifted window transformation modules; both the second-stage layer module and the fourth-stage layer module include: a block merging module and two shifted window transformation modules; the third-stage layer module includes: a block merging module and six shifted window transformation modules.

[0020] Furthermore, the shifted window transformation module, from input to output, sequentially passes through a multi-head self-attention module and a multi-layer perceptron module, and a normalization layer is set before each multi-head self-attention module and multi-layer perceptron module; residual connections are used between each normalization layer;

[0021] The setting method of the multi-head self-attention module is as follows: the multi-head self-attention module includes a window multi-head self-attention module and a shifted window multi-head self-attention module; in several consecutive shifted window transformation modules, from input to output, the multi-head self-attention module in the odd-positioned shifted window transformation modules is: the window multi-head self-attention module, and the multi-head self-attention module in the even-positioned shifted window transformation modules is: the shifted window multi-head self-attention module.

[0022] Furthermore, the discriminator 1, from input to output, sequentially passes through a first convolutional module, a second convolutional module, a third convolutional module, and a fully connected layer; both the second convolutional module and the third convolutional module use batch normalization layers; the first convolutional module, the second convolutional module, and the third convolutional module all use the LeakyReLU activation function; the fully connected layer uses the tanh activation function; the structure of discriminator 2 is the same as that of discriminator 1.

[0023] Furthermore, the total loss of the improved generator module includes an adversarial loss and a content loss; the calculation formula for the total loss of the improved generator module is: , where represents the adversarial loss, represents the training parameters, represents the content loss.

[0024] Furthermore, during iterative training, the learning rate of the generator network of the improved PFNET model is set to 0.0001, the optimizer is AdamW, the learning rate of the dual discriminator network is 0.00001, the optimizer is Adam, the random rotation angle of the training set images is 5 degrees, the probability of random horizontal flipping is 5, and the probability of random vertical flipping is 5.

[0025] Furthermore, when the performance test metrics reach a fixed threshold, specifically: information entropy: 7.0, standard deviation: , visual information fidelity: , structural similarity index: Mutual information: 。

[0026] Furthermore, the implementation process of inputting the verification dataset into the executable improved PFNET model to obtain the RGB fusion image includes the following steps:

[0027] S91. Convert the degree of linear polarization image DOLP and the intensity image in the training set from the RGB color space to the YCrCb color space to obtain the luminance information and the Cr and Cb channel information;

[0028] S92. The upper encoder module and the lower encoder module respectively extract features from the luminance information, and splice the feature information extracted by the upper encoder module and the lower encoder module to obtain the luminance feature;

[0029] S93. The luminance feature passes through the fusion module and the decoder module in sequence to obtain the reconstructed image;

[0030] S94. Input the reconstructed image and the luminance information into discriminator 1 and discriminator 2 for adversarial training to obtain the feedback signal;

[0031] S95. Splice the reconstructed image and the Cr and Cb channel information, and convert the spliced image to the RGB color space to obtain the RGB fusion image.

[0032] Furthermore, the feedback signal is used to optimize the improved generator module.

[0033] The beneficial effects of the method of the present invention are as follows:

[0034] 1. Achieve deep fusion of global and local features: The present invention introduces an encoder constructed by a shifted window transform feature extraction network (swintransformer), and at the same time, with the help of the window multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA) mechanisms, realizes their coordinated operation. This method can effectively extract the global polarization features and local texture details of the image, breaks the limitations of traditional methods in feature extraction, greatly improves the integrity of feature expression, provides more comprehensive and accurate feature information for subsequent image fusion, and significantly enhances the description ability of the fusion image for complex scenes.

[0035] 2. Adversarial Optimization Based on Dual Discriminators: The present invention designs an independent dual discriminator module. During the training process, the two discriminators conduct adversarial training on the intensity image and the degree of linear polarization image respectively. This unique design can effectively prompt the generator to evenly retain multi-source complementary information and avoid the problem of information bias. Through this method, it is ensured that the fused image not only has high polarization resolution and can accurately present information related to polarization characteristics, but also has rich textures, greatly improving the quality and practicality of the fused image and meeting the requirements for high-quality fused images in multiple fields.

[0036] 3. Composite Loss Function: By combining and improving the adversarial loss, structural similarity loss, and mean square error loss, the network is comprehensively optimized from multiple dimensions such as pixel alignment, structural consistency, and adversarial stability. Compared with traditional methods, the fused images trained based on this composite loss function perform better in key indicators such as information entropy (EN) and visual information fidelity (VIF). In terms of visual effects, the clarity, contrast, and detail expressiveness of the images are significantly improved. At the information level, the amount of information contained in the images is more abundant, providing a better data basis for subsequent image analysis and applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flowchart of the method described in the embodiment of the present invention;

[0038] Figure 2 It is a schematic diagram of the improved PFNET model in the embodiment of the present invention;

[0039] Figure 3 It is a schematic diagram of the upper encoder module and the lower encoder module structures in the improved generator module in the embodiment of the present invention;

[0040] Figure 4 It is a schematic diagram of the continuous shifted window transform module structure in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] This embodiment provides a color polarization image fusion method based on the improvement of PFNET. The flowchart of the method is as Figure 1 shown, and the method includes the following steps:

[0043] S1. Obtain a color polarization image dataset and perform preprocessing. Divide the preprocessed color polarization image dataset into a training set, a test set, and a validation dataset according to the ratio of 6:2:2.

[0044] The relevant operations in step S1 are introduced with specific examples:

[0045] The color polarization image dataset includes color polarization images of different objects. In this embodiment, 40 color polarization images of the Tokyo Institute of Technology Color-Polarization Dataset are incorporated into the training set to enhance the diversity of the training set data.

[0046] Natural light is unpolarized light, which becomes partially polarized light after being reflected by the surface of an artificial target. In this embodiment, the Stokes vector is used to describe the polarization characteristics of the color polarization image dataset;

[0047] The preprocessed color polarization image dataset includes: intensity image and degree of linear polarization (DOLP) image.

[0048] The intensity image has the following calculation formula: , where represents the Stokes vector, represents the total light intensity, and represent the light intensity differences of the horizontal and vertical linearly polarized components respectively, and represent the light intensity differences of the 45° and 135° linearly polarized components respectively, and represent the light intensity differences of the right-handed and left-handed circularly polarized components respectively, represent the light intensity values of the 0°, 45°, 90°, and 135° polarized components respectively, represent the light intensity values of the right-handed and left-handed polarized components respectively; if , and any of the parameter values is not equal to 0, it means that the incident light has polarization characteristics.

[0049] The calculation formula for the degree of linear polarization (DOLP) image is: .

[0050] S2. Improve the PFNET model, specifically:

[0051] S21. Replace the original generator module with an improved generator module, which is obtained by introducing the encoder module in the fusion image generator constructed by the shifted window transform feature extraction network into the original generator module of the PFNET model. The improved generator module includes an upper encoder module and a lower encoder module;

[0052] S22. Introduce the dual discriminator module constructed by the Gan adversarial network into the PFNET model. The dual discriminator module includes discriminator 1 and discriminator 2;

[0053] S23. Introduce the total loss of the Gan adversarial network into the loss function of the PFNET model to obtain an improved loss function. The improved loss function is specifically: , where represents the total loss of the improved generator module, and represent the losses of discriminator 1 and discriminator 2 respectively.

[0054] Introduce the relevant operations of step S2 with a specific example:

[0055] As Figure 3 shown, from input to output, the upper encoder module sequentially passes through an image patch partition module (patchpartition), a first-stage layer module, a second-stage layer module, a third-stage layer module, and a fourth-stage layer module; the structure of the lower encoder module is the same as that of the upper encoder module.

[0056] The first-stage layer module includes: a linear embedding module (liner embedding) and two shifted window transform modules; both the second-stage layer module and the fourth-stage layer module include: a patch merging module (patchmerging) and two shifted window transform modules; the third-stage layer module includes: a patch merging module and six shifted window transform modules.

[0057] The role of the encoder in the improved generator module is to extract features from the input image, capture more advanced features by increasing the number of channels (depth). Therefore, the quality of feature extraction affects the generation quality of the subsequent fusion image. Introducing the encoder module of the shifted window transform feature extraction network to construct the fusion image generator here can enhance the ability of the fusion image generator to extract local and global information in the image, thereby improving the quality of the fusion image generated by the generator.

[0058] As Figure 3 and 4As shown, the shifted window transformation module, from input to output, sequentially passes through the multi-head self-attention module (Multi-Head Self-Attention, MSA) and the multi-layer perceptron module (Multilayer Perceptron, MLP). A normalization layer (Layer Normalization, LN) is set before each multi-head self-attention module and multi-layer perceptron module; residual connections are made between each normalization layer.

[0059] The setting method of the multi-head self-attention module is as follows: The multi-head self-attention module includes a window multi-head self-attention module and a shifted window multi-head self-attention module; in several consecutive shifted window transformation modules, from input to output, the multi-head self-attention module in the odd-numbered shifted window transformation modules is: the window multi-head self-attention module (Windows Multi-Head Self-Attention, WMSA), and the multi-head self-attention module in the even-numbered shifted window transformation modules is: the shifted window multi-head self-attention module (Shifted Windows Multi-Head Self-Attention, SW-MSA).

[0060] The image input to the shifted window transformation module needs to be first processed by layer normalization and then input into the multi-head self-attention module. During this process, the attention weights between different positions are adjusted according to the pre-computed attention mask to ensure that only elements within the same window can pay attention to each other. The newly obtained features are added to the original input features to form a residual connection, and then passed through another normalization layer, and then sent into the multi-layer perception module (MLP) for further transformation; the multi-layer perception module usually contains two linear layers and a GELU activation function to enhance the non-linear expression ability. The output of the multi-layer perception module is also added with another residual connection, that is, added to the input features after normalization, and finally used as the output signal of this module to be input into the next module.

[0061] When the multi-head self-attention module calculates self-attention, the input feature sequence is mapped into three different feature representations, namely Q, K, and V. The Q, K, and V respectively represent the query matrix, the key matrix, and the value matrix. Specifically: , , ,where represents the input sequence, represents the learnable parameter; the calculation formula of self-attention is as follows: ,where Z represents the learnable relative position encoding, represents the dimension of K, represents the transpose of K.

[0062] The role of the image fusion network discriminator is to evaluate the quality of image fusion, provide a feedback signal to help the generator improve its output, and generate a fused image that fully integrates the information of the two source images. Here, the Gan adversarial network is introduced to construct a dual discriminator module. Through adversarial training, the fused image generated by the improved generator module is forced to retain more source image information, enabling the generated fused image to possess the complementary information of both source images and achieving the fusion of complementary information.

[0063] As Figure 2 shown, the discriminator 1 sequentially passes through the first convolutional module, the second convolutional module, the third convolutional module, and the fully connected layer from input to output; the second convolutional module and the third convolutional module both use batch normalization layers; the first convolutional module, the second convolutional module, and the third convolutional module all use LeakyReLU activation functions; the fully connected layer uses tanh activation functions; the structure of discriminator 2 is the same as that of discriminator 1.

[0064] LeakyReLU The calculation formula of the activation function is:

[0065] , where, represents that the input can be any real number, represents the leakage coefficient, with a value between 0 and 1. In this embodiment, takes the value of 0.01.

[0066] Discriminator 1 is used to distinguish between the fused image and the first source image, and discriminator 2 is used to distinguish between the fused image and the second source image; taking discriminator 1 as an example: the input of discriminator 1 passes through the first convolutional module and outputs an output tensor with 16 channels. The second convolutional module receives the output of the first convolutional module and outputs an output tensor with 32 channels. The third convolutional module receives the output of the second convolutional module and outputs an output tensor with 64 channels. The fully connected layer uses tanh activation function to generate a scalar to estimate the probability that the input image comes from the first type of image rather than the fused image. Specifically: , where, t represents that the input can be any real number, and e represents the natural number;

[0067] The role of the image fusion network loss function is to define the difference between the model's predicted output and the expected output, and guide the model on how to adjust its parameters to minimize this difference. Here, by combining the total loss of the improved generator module and the dual discriminator loss to improve the PFNET image fusion network loss function, it helps to improve the quality of the fused image, stabilize the training process, and achieve the best fusion effect.

[0068] The total loss of the improved generator module includes adversarial loss and content loss; the calculation formula for the total loss of the improved generator module is: , where represents the adversarial loss, represents the content loss, represents the training parameter, which is set to 2 in this embodiment .

[0069] The calculation formula is: , where represents the mathematical expectation, represents the fused image containing only luminance information generated by the improved generator module, represents discriminator 1, represents discriminator 2.

[0070] The calculation formula is: , where represents the structural similarity loss, , represents the harmonic coefficient, which is set to 0.01 in this embodiment .

[0071] Structural similarity is part of the most widely used content loss function in the field of image fusion. It reflects the similarity between images from three aspects: light intensity, contrast, and structure; the calculation formula is: , where represents calculating the similarity between two images, and respectively represent the luminance information images of source image 1 and source image 2, and represent the weight parameters used to balance the contributions of the two structural similarity terms (Structural Similarity Index, abbreviated as SSIM) in the total loss function Lssim. In this embodiment and are both set to 0.5;

[0072] represents the mean absolute error loss, which can directly reflect the average difference between the predicted value and the true value. The calculation formula is: , where M and N respectively represent the height and width of the image, represents the intensity image and the average value of the degree of linear polarization image DOLP at the pixel level. Since the intensity image The intensity image and the degree of linear polarization image DOLP are complementary information of the same scene. Therefore, taking the average value of the intensity image and the degree of linear polarization image DOLP will be beneficial to the fusion result.

[0073] The fused image generated by the improved generator module includes information from two different source images. Therefore, using only one discriminator may be biased towards images from a certain source, resulting in the loss of some information. So we constructed a dual discriminator module through the Gan adversarial network. Discriminator 1 and discriminator 2 calculate the discriminator loss respectively to better preserve the information of the source images. Discriminator 1 and discriminator 2 are independent of each other but use the same loss function. The calculation formula for the loss function of a single discriminator is:

[0074] .

[0075] Among them, represents the discriminator , represents the brightness information of the source image , , and represent the source image and the source image .

[0076] S3. Input the training set into the improved PFNET model for iterative training. When the number of training times is greater than 100 and the total loss of the generator module and the losses of discriminator 1 and discriminator 2 no longer decrease significantly, the target model is obtained.

[0077] The relevant operations in step S3 are introduced with a specific example:

[0078] During iterative training, set the learning rate of the generator network of the improved PFNET model to 0.0001, the optimizer to AdamW, the learning rate of the dual discriminator network to 0.00001, the optimizer to Adam, the random rotation angle of the training set images to 5 degrees, the probability of random horizontal flipping to 5, and the probability of random vertical flipping to 5.

[0079] S4. Input the test set into the target model for testing to obtain performance test metrics: information entropy, standard deviation, visual information fidelity, structural similarity index, and mutual information. When the performance test metrics reach a fixed threshold, an executable improved PFNET model is obtained.

[0080] The relevant operations in step S4 are introduced with a specific example:

[0081] When the performance test metrics reach the fixed threshold specifically: information entropy: 7.0, standard deviation: , Visual information fidelity: , Structural similarity index: and Mutual information: .

[0082] Information entropy is a no-reference evaluation metric based on information theory, used to measure the amount of information contained in the fused image. The calculation formula of EN is as follows: , where L represents the number of gray levels of the image, represents the normalized histogram of the gray levels in the

[0083] Standard deviation represents the distribution and contrast of pixel values in the image. The larger the value of , the higher the contrast. The calculation formula of the standard deviation is as follows: , where M and N represent the height and width of the image respectively, represents the mean value of the image pixel values, represents the pixel value at the position

[0084] Visual information fidelity is consistent with the human visual system and can measure information fidelity. It can be calculated in four steps: (1) Filter and divide the source image and the fused image into different blocks; (2) Evaluate the visual information of each block; (3) Calculate the VIF of each sub-band; (4) Calculate the overall metric. The visual information fidelity is directly proportional to the performance of the fusion method.

[0085] Structural similarity index is used to measure the similarity between two images in terms of correlation, brightness, and contrast distortion, specifically: , where, represents the local similarity between images x and y with a window of size , represents the partial image of image x within the window , and respectively represent the mean and variance of represents and the covariance of represents the partial image of image y within the window , and respectively represent the mean and variance of and For maintaining stability, in this embodiment, and are respectively set to and is .

[0086] In polarization image fusion, SSIM consists of the structural similarity between two source images and the fused image.

[0087] Structural similarity index The calculation formula is as follows: .

[0088] Among them, and represent weight coefficients, which are used to balance image and image's contribution to the total structural similarity index. In this embodiment, and are both set to 0.5.

[0089] Mutual information is a way to measure the information similarity between the fused image and the source image. The larger the mutual information value, the more source image information is contained in the fused image, and the better the fusion effect. Mutual information The calculation formula is as follows: , among which, represents the mutual information function formula, W includes source image A and source image B, represents the fused image, represents the source image and the fused image 's joint histogram, and are the marginal histogram statistical probabilities of the source image and the fused image , represents the pixel intensity value (or gray value) of the source image W, represents the pixel intensity value (or gray value) of the fused image F.

[0090] S5. Input the validation dataset into the executable improved PFNET model to obtain the RGB fused image.

[0091] Introduce the relevant operations in step S5 with a specific example:

[0092] The implementation process of inputting the validation dataset into the executable improved PFNET model to obtain the RGB fused image includes the following steps:

[0093] S91. The degree of linear polarization image DOLP and the intensity image in the training set Convert from the RGB color space to the YCrCb color space to obtain the luminance information and the Cr and Cb channel information;

[0094] S92. The upper encoder module and the lower encoder module respectively perform feature extraction on the luminance information. In the upper encoder module and the lower encoder module, the image block module divides the input image into image blocks of size 4×4, converts the minimum operation unit of the image from pixels to image blocks, and the image processed by the image block module then passes through the first-stage layer to embed the 4×4 image blocks into vectors. Subsequently, the image processed by the first-stage layer then successively passes through the second, third, and fourth-stage layers. In the second, third, and fourth-stage layers, every time passing through a stage layer, the size of the input image becomes half of the original, and at the same time the feature dimension becomes twice the original. The feature information extracted by the upper encoder module and the lower encoder module is concatenated to obtain the luminance feature. The role of the image block module is to reduce the size of the input image while realizing the concatenation of image blocks at the feature dimension level;

[0095] S93. The luminance feature successively passes through the fusion module and the decoder module to obtain the reconstructed image of the fused image containing only luminance information;

[0096] S94. Input the reconstructed image and the luminance information into discriminator 1 and discriminator 2 for adversarial training to obtain a feedback signal;

[0097] S95. Concatenate the reconstructed image and the Cr and Cb channel information, and convert the concatenated image to the RGB color space to obtain the RGB fused image.

[0098] As Figure 2 shown, the fusion module successively passes through two convolutional and ReLU activation function layers from input to output, and each convolutional layer uses the ReLU activation function; the decoder module successively passes through five convolutional layers from input to output, and the first four convolutional layers all use the ReLU activation function.

[0099] The feedback signal is used to optimize the improved generator module: guide the generator to improve the generated image by adjusting parameters (adversarial loss).

Claims

1. A color polarization image fusion method improved based on PFNET, characterized in that, The method includes the following steps: S1. Obtain a color polarization image dataset and perform preprocessing. Divide the preprocessed color polarization image dataset into a training set, a test set, and a validation dataset according to a ratio of 6:2:2; S2. Improve the PFNET model. Specifically: S21. Replace the original generator module with an improved generator module. The improved generator module is obtained by introducing the encoder module in the fusion image generator constructed by the shifted window transform feature extraction network into the original generator module of the PFNET model. The improved generator module includes an upper encoder module and a lower encoder module; From input to output, the upper encoder module sequentially passes through an image block module, a first-stage layer module, a second-stage layer module, a third-stage layer module, and a fourth-stage layer module; The structure of the lower encoder module is the same as that of the upper encoder module; The first-stage layer module includes: a linear embedding module and two shifted window transform modules; both the second-stage layer module and the fourth-stage layer module include: a block merging module and two shifted window transform modules; the third-stage layer module includes: a block merging module and six shifted window transform modules; S22. Introduce a dual discriminator module constructed by the Gan adversarial network into the PFNET model. The dual discriminator module includes discriminator 1 and discriminator 2; S23. Introduce the total loss of the Gan adversarial network into the loss function of the PFNET model to obtain an improved loss function, and the improved loss function is specifically: , where represents the total loss of the improved generator module, and respectively represent the losses of discriminator 1 and discriminator 2; S3. Input the training set into the improved PFNET model for iterative training. When the number of training times is greater than 100, and the total loss of the generator module and the losses of discriminator 1 and discriminator 2 no longer decrease significantly, obtain the target model; S4. Input the test set into the target model for testing, and obtain performance test metrics: information entropy, standard deviation, visual information fidelity, structural similarity index, and mutual information. When the performance test metrics reach a fixed threshold, obtain an executable improved PFNET model; S5. Input the validation dataset into the executable improved PFNET model to obtain an RGB fusion image.

2. The improved color polarization image fusion method based on PFNET according to claim 1, wherein The specific preprocessing is as follows: the Stokes vector is used to describe the polarization characteristics of the color polarization image dataset; the preprocessed color polarization image dataset includes: an intensity image and a degree of linear polarization (DOLP) image.

3. The color polarization image fusion method improved based on PFNET according to claim 2, characterized in that, From input to output, the shifted window transform module sequentially passes through a multi-head self-attention module and a multi-layer perceptron module. A normalization layer is set before each multi-head self-attention module and multi-layer perceptron module; residual connections are made between each normalization layer; The setting method of the multi-head self-attention module is as follows: The multi-head self-attention module includes a window multi-head self-attention module and a shifted window multi-head self-attention module; in several consecutive shifted window transform modules, from input to output, the multi-head self-attention module in the odd-numbered shifted window transform modules is: the window multi-head self-attention module, and the multi-head self-attention module in the even-numbered shifted window transform modules is: the shifted window multi-head self-attention module.

4. The improved color polarization image fusion method based on PFNET according to claim 1, characterized in that The discriminator 1 sequentially passes through a first convolutional module, a second convolutional module, a third convolutional module, and a fully connected layer from input to output; the second convolutional module and the third convolutional module both use batch normalization layers; the first convolutional module, the second convolutional module, and the third convolutional module all use the LeakyReLU activation function; the fully connected layer uses the tanh activation function; the structure of discriminator 2 is the same as that of discriminator 1.

5. The color polarization image fusion method improved based on PFNET according to claim 1, characterized in that The total loss of the improved generator module includes adversarial loss and content loss; the calculation formula for the total loss of the improved generator module is: , where represents the adversarial loss, represents the training parameters, represents the content loss.

6. The improved color polarization image fusion method based on PFNET according to claim 1, wherein During iterative training, set the learning rate of the generator network of the improved PFNET model to 0.0001, the optimizer to AdamW, the learning rate of the dual discriminator network to 0.00001, the optimizer to Adam, the random rotation angle of the training set images to 5 degrees, the probability of random horizontal flipping to 5, and the probability of random vertical flipping to 5.

7. The improved color polarization image fusion method based on PFNET according to claim 1, wherein When the performance test metrics reach a fixed threshold, specifically: information entropy: 7.0, standard deviation: , visual information fidelity: , structural similarity index: and mutual information: .

8. The color polarization image fusion method improved based on PFNET according to claim 1, characterized in that, The implementation process of inputting the validation data set into the executable improved PFNET model to obtain the RGB fusion image includes the following steps: S81. Convert the degree of linear polarization image DOLP and the intensity image in the training set from the RGB color space to the YCrCb color space to obtain the luminance information and the Cr and Cb channel information; S82. The upper encoder module and the lower encoder module respectively extract features from the luminance information, and splice the feature information extracted by the upper encoder module and the lower encoder module to obtain luminance features; S83. The luminance features sequentially pass through the fusion module and the decoder module to obtain the reconstructed image; S84. Input the reconstructed image and the luminance information into discriminator 1 and discriminator 2 for adversarial training to obtain a feedback signal; S85. Splice the reconstructed image and the Cr and Cb channel information, and convert the spliced image to the RGB color space to obtain the RGB fusion image.

9. The color polarization image fusion method improved based on PFNET according to claim 8, wherein The feedback signal is used to optimize the improved generator module.