Polarization image fusion method and system with global perception and multi-branch heterogeneous attention

By employing a polarization image fusion method based on global perception and multi-branch heterogeneous attention, the problems of information loss and insufficient feature extraction in polarization image fusion are solved, achieving efficient polarization feature preservation and detail enhancement, thereby improving image quality.

CN121120426BActive Publication Date: 2026-03-24HUAQIAO UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies for polarization image fusion suffer from loss of polarization information, limited feature extraction capabilities, and insufficient attention focus areas, making it difficult to effectively preserve and enhance detailed information.

Method used

A polarization image fusion method based on global perception and multi-branch heterogeneous attention is adopted. By constructing a dual-branch structure and a multi-branch heterogeneous polarization cross attention module (MPCA), a polarization channel attention module (CA), and a collaborative cross-dimensional module (CCDM), combined with unsupervised training and a designed loss function, multi-dimensional feature extraction and detail preservation of polarization images are achieved.

Benefits of technology

It improves the ability to capture details and the accuracy of feature recognition in polarization images, reduces computational complexity, and maintains and enhances the information flow and fusion effect of polarization features in multiple dimensions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120426B_ABST
    Figure CN121120426B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of optical image processing, and discloses a polarization image fusion method and system with global perception and multi-branch heterogeneous attention, which comprises the following steps: acquiring polarization images at four angles, obtaining an intensity image and a linear polarization degree image dataset according to a Stokes vector synthesis method; constructing a polarization image fusion model with global perception and multi-branch heterogeneity, and training the polarization image fusion model by using the corresponding intensity image and linear polarization degree image dataset; and realizing polarization image fusion by using the polarization image fusion model; while realizing polarization image fusion, the application significantly enhances the perception of polarization features and effectively maintains image details; meanwhile, by improving the contrast and brightness of the image, the fused image is clearer and more detailed; in complex scene processing, the application can better retain object edge features and subtle light reflection information, and effectively improves the detail loss problem in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical image processing technology, and in particular to a polarization image fusion method and system based on global perception and multi-branch heterogeneous attention. Background Technology

[0002] Polarization detection imaging, as an advanced imaging method based on the polarization characteristics of light waves, can further extract parameters such as linear polarization degree (DoLP) and angle of polarization (AoP) on top of the intensity information of traditional imaging, expanding the perception dimension of the imaging system. The core of polarization image fusion lies in making full use of the complementarity of different polarization information to improve the recognizability of targets and the ability to express scene details, especially showing unique advantages in the analysis of object surface characteristics. Intensity images contain basic brightness and color information of the scene, usually used to describe the lighting conditions and surface reflection properties of the scene, providing an overall understanding of the scene. However, intensity images are relatively limited in their representation of details and object surface structures. Linear polarization degree (DoLP) can capture the differences in polarization characteristics caused by changes in the surface structure of an object, thus providing rich detail information and having a better ability to express surface details. However, the unique polarization characteristics in linear polarization degree images are often difficult to perceive directly with the human eye. Therefore, how to effectively fuse this image information becomes the key to improving image quality. Through fusion methods, the texture details, contrast, and sharpness of the image can be significantly improved.

[0003] The dimension of polarization information holds significant value across multiple application areas. In the military field, polarization imaging effectively enhances target recognition capabilities, particularly in low-light, complex background, or camouflage conditions, enabling accurate identification and tracking of objects. In industrial applications, polarization images can be used to detect surface defects, cracks, and corrosion, contributing to improved quality control in production processes and enhanced equipment maintenance efficiency. In target detection, polarization information enhances the identification of complex targets, providing additional detail and improving accuracy, especially in the face of environmental interference or background noise. In the biomedical field, polarization imaging is used for tissue imaging and lesion detection, particularly in cancer screening and wound healing monitoring. By analyzing polarization information, structural changes in tissues are obtained, providing more precise and detailed information than traditional imaging methods. Combining polarization information with other imaging techniques can significantly improve imaging quality and analytical capabilities across various fields.

[0004] In existing technologies, both convolutional neural networks and autoencoder-based methods utilize convolutional layers to extract image features, possessing hierarchical feature abstraction capabilities, and optimize parameters through end-to-end training. Zhang et al. proposed a network structure consisting of three modules: feature extraction, connection, and reconstruction, and improved the loss function, achieving good results (Zhang, Junchao, et al. "Polarization image fusion with self-learned fusion strategy." PatternRecognition 118 (2021): 108045.). Liu et al. proposed a pixel-level image fusion method and combined it with saliency detection-guided image encoding, using its features as fusion weights (Liu, Jinyang, et al. "SGFusion: Asaliency guided deep-learning framework for pixel-level image fusion." Information Fusion 91 (2023): 205-214.). Han et al. introduced structural loss and designed adaptive intensity loss, using weight coefficients to constrain image similarity (Xu, Han, et al. "Attention-guided polarization image fusion using salient information distribution." IEEE Transactions on Computational Imaging 8 (2022): 1117-1130.). Wang et al. used wavelet transform to expand the receptive field of convolutional layers, extracting shared shallow features to encode low-frequency structural contours and high-frequency textures (Wang Y, Liu J, Wang J, et al. HaarFuse: A Dual-Branch Infrared and Visible LightImage Fusion Network based on Haar Wavelet Transform[J]. Pattern Recognition,2025: 111594.). These image fusion methods based on convolutional and autoencoder structures show significant advantages in terms of deep representation of features, flexibility of fusion mechanisms, and optimized design of loss functions, especially in achieving good results in image sharpness enhancement and salient region enhancement; however, the inability to effectively perceive polarization information leads to loss of detail, which affects subsequent tasks.

[0005] Although the Chinese invention application CN120355594A, entitled "A Deep Learning-Based Polarization Image Fusion Method for Sparse Aperture Optical Systems," can achieve polarization image fusion, effectively suppress noise, improve the contrast of the fused image, and integrate polarization information, it still has problems such as loss of polarization information, limited feature extraction capabilities, insufficient attention area, and easy neglect of effective information. Summary of the Invention

[0006] The purpose of this invention is to solve the problems in the prior art.

[0007] The technical solution adopted by this invention to solve its technical problem is: to provide a polarization image fusion method based on global perception and multi-branch heterogeneous attention, comprising the following steps:

[0008] The polarization images at four angles were obtained, and a dataset containing intensity images and linear polarization degree images was obtained by using the Stokes vector synthesis method.

[0009] A global perception and multi-branch heterogeneous polarization image fusion model is constructed and trained using a dataset containing intensity images and linear polarization degree images to obtain the training weights of the polarization image fusion model.

[0010] Load the training weights and use the fusion model to complete the polarization image fusion task;

[0011] The polarization image fusion model includes a dual-branch structure and an image reconstruction structure. The dual-branch structure comprises a first branch and a second branch, which respectively receive and process the intensity image and the linear polarization degree image. Both the first and second branches include several sequentially connected multi-branch heterogeneous polarization cross-attention modules (MPCA), polarization channel attention modules (CA), and collaborative cross-dimensional modules (CCDM). Each MPCA performs multi-feature extraction and cross-fusion on the input features, with skip connections in the middle of the MPCA to achieve comprehensive perception. The features output by the last MPCA serve as the input features of the CA, and the output features of the CA serve as the input features of the collaborative cross-dimensional module (CCDM). The collaborative cross-dimensional module (CCDM) performs overall spatial dimension transformation to separate coordinate dimensions on the input features. The features output by the first branch CCDM and the features output by the second branch CCDM are input into the image reconstruction structure, merged in channels, and reconstructed through three convolutional layers to obtain the fused polarization image.

[0012] Preferably, the first branch includes an RGB2YCrCb module, which converts the intensity image from RGB three-channel to YCrCb three-channel space, and separates the luminance channel and chrominance channel Y of the intensity image, and inputs the image of the luminance channel into the plurality of multi-branch heterogeneous polarization cross-attention modules;

[0013] The features output by the upper branch's collaborative cross-dimensional module CCDM and the lower branch's collaborative cross-dimensional module CCDM are merged in the channel, then reconstructed through 3 convolutional layers, and finally converted from YCrCb space to RGB space by the YCrCb2RGB module to obtain the fused polarization image.

[0014] Preferably, the acquisition of polarization images at four angles, and the generation of a dataset containing intensity images and linear polarization degree images using the Stokes vector synthesis method, are represented as follows:

[0015] ;

[0016] ;

[0017] ;

[0018] ;

[0019] in, Represents an intensity image. This represents the polarization difference between the 0° and 90° directions; This represents the polarization difference between the 45° and 135° directions; Represents the linear polarization degree image; , , and These represent polarization images at 0°, 45°, 90°, and 135°, respectively.

[0020] Preferably, the multi-branch heterogeneous polarization cross-attention module includes a preprocessing layer PA, an upper branch UP, and a lower branch DOWN;

[0021] The preprocessing layer PA consists of three layers connected in sequence. The first layer includes a 3×3 convolution, batch normalization (BN), and LReLU activation function. The second layer includes a 3×3 convolution and LReLU activation function. The structure of the third layer is the same as that of the first layer. Finally, the output of the first layer is concatenated with the output of the third layer by adding the elements of the first layer and the output of the third layer to obtain the output of the preprocessing layer.

[0022] After receiving the output of the preprocessing layer, the upper branch UP sequentially processes the output through 1×5 convolution, batch normalization (BN), LReLU activation function, 5×1 convolution, batch normalization (BN), and Sigmoid activation function. Then, it performs element-wise multiplication with the output of the preprocessing layer PA to obtain the output of the upper branch.

[0023] After receiving the output of the preprocessing layer PA, the lower branch DOWN sequentially processes the output through 1×1 convolution, batch normalization (BN), LReLU activation function, 1×1 convolution, batch normalization (BN), and Sigmoid activation function. Then, it performs element-wise multiplication with the output of the preprocessing layer to obtain the output of the lower branch.

[0024] Finally, output the UP branch. The output of the lower branch DOWN Output of preprocessing layer PA Adding them together yields the output of the first multi-branch heterogeneous polarization cross-attention module. , represented as:

[0025] ;

[0026] ;

[0027] ;

[0028] Where F represents the input of the multi-branch heterogeneous polarization cross attention module.

[0029] Preferably, the multi-branch heterogeneous polarization cross-attention module introduces skip connections to achieve global perception in order to compensate for possible loss of detail;

[0030] Both the first branch and the second branch include five multi-branch heterogeneous polarization cross-attention modules; skip connections are set between the output of the preprocessing layer of the first multi-branch heterogeneous polarization cross-attention module and the output of the fifth multi-branch heterogeneous polarization cross-attention module, between the output of the first multi-branch heterogeneous polarization cross-attention module and the output of the fourth multi-branch heterogeneous polarization cross-attention module, and between the output of the second multi-branch heterogeneous polarization cross-attention module and the output of the third multi-branch heterogeneous polarization cross-attention module.

[0031] Preferably, the channel attention module CA is represented as follows:

[0032] ;

[0033] in, This represents the output after passing through the channel attention module. This represents the input to the channel attention module (CA). This represents the Sigmoid activation function. Indicates a fully connected layer. This represents the LRelu activation function. This indicates the average pooling operation. This indicates a max pooling operation.

[0034] Preferably, the collaborative cross-dimensional module is represented as follows:

[0035] ;

[0036] ;

[0037] ;

[0038] ;

[0039] ;

[0040] in, This represents the output after passing through the channel attention module. This represents the Sigmoid activation function. This indicates that the convolution has been performed using a 3×3 method. This represents the LRelu activation function. Indicates attentional characteristics; This indicates average pooling along the x-axis. This indicates average pooling along the y-axis, and Concat indicates merging across channels. Indicates will Features obtained by merging two types of average pooling on the channel; This indicates that after a 1×1 convolution, This represents the activation function. Indicates batch normalization. Indicates a separation operation. and Indicates intermediate features during the operation process; This indicates the final output obtained after matrix multiplication, which is the output of the collaborative cross-dimensional module.

[0041] Preferably, the training adopts an unsupervised training method, and the designed loss function includes multi-scale weighted structural loss, intensity loss, and significant gradient loss;

[0042] The multi-scale weighted structural loss , represented as:

[0043] ;

[0044] ;

[0045] in, It is the size of the sliding window. and They represent in × Sliding window size Local mean of the image Local mean of the image and These represent the corresponding local variances. This represents the local covariance of the two input images. and It is the stability coefficient. It is a similarity measurement function. It is the merged image; It is a constant. Pick and The maximum value;

[0046] The strength loss , represented as:

[0047] ;

[0048] ;

[0049] in, This means that a pixel value less than 0.5 is equal to 0, and a pixel value greater than 0.5 is equal to 1. This means that pixels less than 0 are set to 0, and pixels greater than 0 are set to 1. Indicates upsampling, It is average pooling. express Norm; * indicates element-wise multiplication. Represents intermediate variables in the calculation process. It is a fused image;

[0050] The significant gradient loss , represented as:

[0051] ;

[0052] ;

[0053] ;

[0054] in, Represents the gradient operator. , and This represents the computation of S0 image, DoLP image, and fused image. gradient, This indicates taking the maximum value of the gradient magnitudes of the S0 image and the DoLP image. This indicates a mean operation on the image domain; * represents element-wise multiplication. It is used for weight adjustment and filtering out meaningless, non-salient texture regions.

[0055] This invention also provides a polarization image fusion system based on global perception and multi-branch heterogeneous attention, comprising:

[0056] The dataset acquisition module acquires polarization images from four angles and uses the Stokes vector synthesis method to obtain a dataset containing intensity images and linear polarization images.

[0057] The fusion model acquisition module constructs a polarization image fusion model with nonlocal skip connections and multi-branch intersections, and trains it using a dataset containing intensity images and linear polarization degree images to obtain the polarization image fusion model.

[0058] The polarization image fusion model includes a dual-branch structure and an image reconstruction structure. The dual-branch structure comprises a first branch and a second branch, which respectively receive and process the intensity image and the linear polarization degree image. Both the first and second branches include several sequentially connected multi-branch heterogeneous polarization cross-attention modules (MPCA), polarization channel attention modules (CA), and collaborative cross-dimensional modules (CCDM). Each MPCA performs multi-feature extraction and cross-fusion on the input features, with skip connections in the middle of the MPCA to achieve comprehensive perception. The features output by the last MPCA serve as the input features of the CA, and the output features of the CA serve as the input features of the collaborative cross-dimensional module (CCDM). The collaborative cross-dimensional module (CCDM) performs overall spatial dimension transformation to separate coordinate dimensions on the input features. The features output by the first branch CCDM and the features output by the second branch CCDM are input into the image reconstruction structure, merged in channels, and reconstructed through three convolutional layers to obtain the fused polarization image.

[0059] The present invention has the following beneficial effects:

[0060] (1) This invention designs a multi-branch heterogeneous polarization cross-attention module (MPCA), which extends the model's ability to capture polarization image features through a multi-branch structure, enabling it to understand data from different angles and levels. The cross-attention mechanism allows the model to automatically focus on the most critical features of the current task, thereby improving the perception ability of polarization features and the accuracy of the model. While enhancing detail capture, this module adopts a lightweight design to reduce computational complexity, ensuring that the model performance is improved without causing excessive computational burden. The multi-branch polarization cross-attention skip connection mechanism compensates for the possible loss of details during cross-module transmission, strengthening the preservation of polarization detail information.

[0061] (2) The present invention constructs a collaborative cross-dimensional module (CCDM) to realize the transformation of the overall spatial dimension to the processing of the separate coordinate dimension, so as to further realize the flow of detailed information of polarization features in different dimensions, thereby enriching the model's perception and understanding of polarization features and improving the accuracy and efficiency of feature recognition.

[0062] (3) In order to enhance the network’s ability to perceive polarization features, the present invention constructs a polarization enhancement perception loss function. By designing and utilizing the polarization mask to enhance the model, the polarization perception ability is improved while ensuring the consistency of details between the fused image and the source image, thereby achieving the ability to further preserve the details of polarization features.

[0063] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. Attached Figure Description

[0064] Figure 1 This is a diagram illustrating the method steps of an embodiment of the present invention;

[0065] Figure 2 This is a schematic diagram of the structure of the polarization image fusion model according to an embodiment of the present invention;

[0066] Figure 3 This is a schematic diagram of the Multi-Branch Polarization Cross-Attention (MPCA) fusion module according to an embodiment of the present invention;

[0067] Figure 4 This is a schematic diagram of the Collaborative Cross-Dimensional Module (CCDM) according to an embodiment of the present invention;

[0068] Figure 5 The present invention provides the original image, the fused image, and the qualitative results obtained from the polarization image in an embodiment of the invention.

[0069] Figure 6 This is a structural diagram of the device according to an embodiment of the present invention. Detailed Implementation

[0070] See Figure 1 The diagram shows the method steps of an embodiment of the present invention, including the following steps:

[0071] S101, acquire polarization images from four angles, and obtain a dataset containing intensity images and linear polarization images using the Stokes vector synthesis method;

[0072] S102, Construct a polarization image fusion model with nonlocal skip connections and multi-branch intersections, and train it using the dataset to obtain the polarization image fusion model;

[0073] S103, polarization image fusion is achieved using a polarization image fusion model;

[0074] Among them, see Figure 2As shown, the polarization image fusion model includes a dual-branch structure and an image reconstruction structure. The dual-branch structure includes a first branch and a second branch, which respectively receive and process the intensity image and the linear polarization degree image. Both the first and second branches include several sequentially connected multi-branch heterogeneous polarization cross-attention modules (MPCA), polarization channel attention modules (CA), and collaborative cross-dimensional modules (CCDM). Each MPCA performs multi-feature extraction and cross-fusion on the input features, with skip connections in the middle of the MPCA to achieve comprehensive perception. The features output by the last MPCA serve as the input features of the CA, and the output features of the CA serve as the input features of the collaborative cross-dimensional module (CCDM). The collaborative cross-dimensional module (CCDM) performs overall spatial dimension transformation to separate coordinate dimensions on the input features. The features output by the first branch CCDM and the features output by the second branch CCDM are input into the image reconstruction structure, merged in channels, and reconstructed through three convolutional layers to obtain the fused polarization image. The multi-branch heterogeneous polarization cross-attention module significantly enhances the perception capability of polarization features through multi-scale feature extraction and cross-fusion strategies, extracting polarization information at different scales and viewpoints, allowing for full expression of details in the image and avoiding the problem of insufficient low-dimensional feature representation in traditional methods. Through a cross-attention mechanism, features at different scales are complemented and fused, further enhancing the overall feature expressiveness. The nonlocal skip group connection mechanism introduces nonlocal information integration and skip connections, effectively mitigating the potential detail loss problem in deep networks. Nonlocal connections enhance the transmission of global information, while skip connections ensure the preservation of important details between different layers, enabling a smoother transition between features during fusion and further improving detail preservation. The collaborative cross-dimensional module achieves efficient information flow between different dimensions through flexible transformation between spatial and coordinate dimensions. This module optimizes the complementarity and information interaction of polarization features across multiple dimensions, resulting in more refined and comprehensive fused features. In this process, polarization features are not only preserved in the spatial dimension but also achieve detail flow between different coordinate dimensions, further enhancing the global information integration capability of polarization images. Overall, this model, through innovative multi-dimensional information fusion and cross-scale feature transmission strategies, addresses the limitations of traditional polarization image fusion methods in detail preservation and feature enhancement. The modules work together to ensure that the fusion results are accurately preserved at the level of detail and that features are effectively transferred in multiple dimensions, which significantly improves the fusion effect of polarization images.

[0075] Specifically, the first branch includes an RGB2YCrCb module, which converts the intensity image from RGB three-channel to YCrCb three-channel space, and separates the luminance channel and chrominance channel Y of the intensity image, and inputs the image of the luminance channel into the plurality of multi-branch heterogeneous polarization cross-attention modules;

[0076] The features output by the upper branch's collaborative cross-dimensional module CCDM and the lower branch's collaborative cross-dimensional module CCDM are merged in the channel, then reconstructed through 3 convolutional layers, and finally converted from YCrCb space to RGB space by the YCrCb2RGB module to obtain the fused polarization image.

[0077] Specifically, the process involves acquiring polarization images from four angles, and then using the Stokes vector synthesis method to obtain a dataset containing intensity images and linear polarization degree images, represented as follows:

[0078] ;

[0079] ;

[0080] ;

[0081] ;

[0082] in, Represents an intensity image. This represents the polarization difference between the 0° and 90° directions; This represents the polarization difference between the 45° and 135° directions; Represents the linear polarization degree image; , , and These represent polarization images at 0°, 45°, 90°, and 135°, respectively.

[0083] For details, see Figure 3 As shown, the multi-branch heterogeneous polarization cross-attention module includes a preprocessing layer PA, an upper branch UP, and a lower branch DOWN;

[0084] The preprocessing layer PA consists of three layers connected in sequence. The first layer includes a 3×3 convolution, batch normalization (BN), and LReLU activation function. The second layer includes a 3×3 convolution and LReLU activation function. The structure of the third layer is the same as that of the first layer. Finally, the output of the first layer is concatenated with the output of the third layer by adding the elements of the first layer and the output of the third layer to obtain the output of the preprocessing layer.

[0085] After receiving the output of the preprocessing layer, the upper branch UP sequentially processes the output through 1×5 convolution, batch normalization (BN), LReLU activation function, 5×1 convolution, batch normalization (BN), and Sigmoid activation function. Then, it performs element-wise multiplication with the output of the preprocessing layer PA to obtain the output of the upper branch.

[0086] After receiving the output of the preprocessing layer PA, the lower branch DOWN sequentially processes the output through 1×1 convolution, batch normalization (BN), LReLU activation function, 1×1 convolution, batch normalization (BN), and Sigmoid activation function. Then, it performs element-wise multiplication with the output of the preprocessing layer to obtain the output of the lower branch.

[0087] Finally, output the UP branch. The output of the lower branch DOWN Output of preprocessing layer PA Adding them together yields the output of the first multi-branch heterogeneous polarization cross-attention module. , represented as:

[0088] ;

[0089] ;

[0090] ;

[0091] Here, F represents the input of the multi-branch heterogeneous polarization cross-attention module. The upper and lower branches can adapt to different input image sizes by using different convolution kernels for upsampling and downsampling, thus exhibiting higher robustness.

[0092] Specifically, the multi-branch heterogeneous polarization cross-attention module introduces skip connections to achieve global perception, thereby compensating for potential loss of detail. In this embodiment of the invention, both the first branch and the second branch include five multi-branch heterogeneous polarization cross-attention modules. Skip connections are set between the output of the preprocessing layer of the first multi-branch heterogeneous polarization cross-attention module and the output of the fifth multi-branch heterogeneous polarization cross-attention module, between the output of the first multi-branch heterogeneous polarization cross-attention module and the output of the fourth multi-branch heterogeneous polarization cross-attention module, and between the output of the second multi-branch heterogeneous polarization cross-attention module and the output of the third multi-branch heterogeneous polarization cross-attention module.

[0093] Specifically, the channel attention module CA is represented as:

[0094] ;

[0095] in, This represents the output after passing through the channel attention module. This represents the input to the channel attention module (CA). This represents the Sigmoid activation function. Indicates a fully connected layer. This represents the LRelu activation function. This indicates the average pooling operation. This indicates a max pooling operation.

[0096] For details, see Figure 4 As shown, the collaborative cross-dimensional module is represented as follows:

[0097] ;

[0098] ;

[0099] ;

[0100] ;

[0101] ;

[0102] in, This represents the output after passing through the channel attention module. This represents the Sigmoid activation function. This indicates that the convolution has been performed using a 3×3 method. This represents the LRelu activation function. Indicates attentional characteristics; This indicates average pooling along the x-axis. This indicates average pooling along the y-axis, and Concat indicates merging across channels. Indicates will Features obtained by merging two types of average pooling on the channel; This indicates that after a 1×1 convolution, This represents the activation function. Indicates batch normalization. Indicates a separation operation. and Indicates intermediate features during the operation process; This indicates the final output obtained after matrix multiplication, which is the output of the collaborative cross-dimensional module.

[0103] Specifically, in order to achieve effective fusion of polarization images, make the fusion effect conform to human subjective visual perception, and enhance polarization perception ability to effectively preserve detailed information, the model adopts an unsupervised training method. The designed loss function consists of three parts: a multi-scale weighted structural loss, an intensity loss, and a significant gradient loss.

[0104] The multi-scale weighted structural loss To ensure a high degree of similarity between the fused image and the source image, it is represented as:

[0105] ;

[0106] ;

[0107] in, It is the size of the sliding window. and They represent in × Sliding window size Local mean of the image Local mean of the image and These represent the corresponding local variances. This represents the local covariance of the two input images. and It is the stability coefficient. It is a similarity measurement function. It is the merged image; It is a very small constant. Pick and The maximum value;

[0108] The strength loss A salient target that retains sufficient polarization information is represented as:

[0109] ;

[0110] ;

[0111] in, This means that a pixel value less than 0.5 is equal to 0, and a pixel value greater than 0.5 is equal to 1. This means that pixels less than 0 are set to 0, and pixels greater than 0 are set to 1. Indicates upsampling, It is average pooling. express Norm; * indicates element-wise multiplication. Represents intermediate variables in the calculation process. It is a fused image;

[0112] The significant gradient loss This ensures that the fused image retains significant edge and texture features from the input image, filters out textures in non-salient regions, and avoids interference from weak gradients during the fusion process. This can be represented as:

[0113] ;

[0114] ;

[0115] ;

[0116] in, Represents the gradient operator. , and This represents the computation of S0 image, DoLP image, and fused image. gradient, This indicates taking the maximum value of the gradient magnitudes of the S0 image and the DoLP image. This indicates a mean operation on the image domain; * represents element-wise multiplication. Used for weight adjustment and filtering out meaningless, non-salient texture regions. It is a small constant used to avoid division by zero errors.

[0117] See Figure 5 As shown in Table 1, this invention is applied to the original and fused images of polarization images. As can be seen from the figures, this invention achieves excellent polarization image fusion results, enriching the detailed features.

[0118] Table 1:

[0119]

[0120] See Figure 6 The diagram shown is a structural diagram of a device according to an embodiment of the present invention, comprising:

[0121] The dataset acquisition module 601 acquires polarization images from four angles and obtains a dataset containing intensity images and linear polarization images using the Stokes vector synthesis method.

[0122] The fusion model acquisition module 602 constructs a polarization image fusion model with nonlocal skip connections and multi-branch intersections, and trains it using a dataset containing intensity images and linear polarization degree images to obtain the polarization image fusion model.

[0123] The polarization image fusion module 603 uses a polarization image fusion model to achieve polarization image fusion.

[0124] As can be seen, the loss function of this invention employs intensity loss and a mask design for DOLP images, taking into account the physical properties of DOLP images. The mask can more effectively enhance polarization features, effectively guiding the network to strengthen feature extraction from DOLP images. By comparing with more advanced methods in recent years, and selecting a wider range of evaluation metrics, this invention achieves advantages in multiple metrics, demonstrating its superiority. In terms of feature extraction, various convolutional kernels are introduced, such as 1×5 and 5×1 convolutions, which better focus on vertical and horizontal directional features and can adapt to different polarization feature distributions. The cross-dimensional collaborative module compensates for the shortcomings of traditional attention mechanisms in feature extraction in the x-axis and y-axis directions.

[0125] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A polarization image fusion method combining global perception and multi-branch heterogeneous attention, characterized in that, Includes the following steps: The polarization images at four angles were obtained, and a dataset containing intensity images and linear polarization degree images was obtained by using the Stokes vector synthesis method. A global perception and multi-branch heterogeneous polarization image fusion model is constructed and trained using a dataset containing intensity images and linear polarization degree images to obtain the training weights of the polarization image fusion model. Load the training weights and use the fusion model to complete the polarization image fusion task; The polarization image fusion model includes a dual-branch structure and an image reconstruction structure. The dual-branch structure includes a first branch and a second branch, which respectively receive and process the intensity image and the linear polarization degree image. Both the first branch and the second branch include several multi-branch heterogeneous polarization cross-attention modules (MPCA), polarization channel attention modules (CA), and collaborative cross-dimensional modules (CCDM) connected in sequence. Each MPCA performs multi-feature extraction and cross-fusion on the input features, with skip connections in the middle of the MPCA to achieve comprehensive perception. The features output by the last MPCA serve as the input features of the CA, and the output features of the CA serve as the input features of the collaborative cross-dimensional module (CCDM). The collaborative cross-dimensional module (CCDM) performs overall spatial dimension transformation to separate coordinate dimension processing on the input features. The features output by the first branch CCDM and the features output by the second branch CCDM are input into the image reconstruction structure, merged in channels, and reconstructed through three convolutional layers to obtain the fused polarization image. The multi-branch heterogeneous polarization cross-attention module introduces skip connections to achieve global perception in order to compensate for the loss of detail; Both the first branch and the second branch include five multi-branch heterogeneous polarization cross-attention modules; skip connections are set between the output of the preprocessing layer of the first multi-branch heterogeneous polarization cross-attention module and the output of the fifth multi-branch heterogeneous polarization cross-attention module, between the output of the first multi-branch heterogeneous polarization cross-attention module and the output of the fourth multi-branch heterogeneous polarization cross-attention module, and between the output of the second multi-branch heterogeneous polarization cross-attention module and the output of the third multi-branch heterogeneous polarization cross-attention module.

2. The polarization image fusion method based on global perception and multi-branch heterogeneous attention according to claim 1, characterized in that, The first branch includes an RGB2YCrCb module, which converts the intensity image from RGB three-channel to YCrCb three-channel space, and separates the luminance channel and chrominance channel Y of the intensity image, and inputs the image of the luminance channel into the plurality of multi-branch heterogeneous polarization cross-attention modules; The features output by the upper branch's collaborative cross-dimensional module CCDM and the lower branch's collaborative cross-dimensional module CCDM are merged in the channel, then reconstructed through 3 convolutional layers, and finally converted from YCrCb space to RGB space by the YCrCb2RGB module to obtain the fused polarization image.

3. The polarization image fusion method based on global perception and multi-branch heterogeneous attention according to claim 1, characterized in that, The process involves acquiring polarization images at four angles, and then using the Stokes vector synthesis method to obtain a dataset containing intensity and linear polarization images, represented as follows: ; ; ; ; in, Represents an intensity image. This represents the polarization difference between the 0° and 90° directions; This represents the polarization difference between the 45° and 135° directions; Represents the linear polarization degree image; , , and These represent polarization images at 0°, 45°, 90°, and 135°, respectively.

4. The polarization image fusion method based on global perception and multi-branch heterogeneous attention according to claim 1, characterized in that, The multi-branch heterogeneous polarization cross-attention module includes a preprocessing layer PA, an upper branch UP, and a lower branch DOWN; The preprocessing layer PA consists of three layers connected in sequence. The first layer includes a 3×3 convolution, batch normalization (BN), and LReLU activation function. The second layer includes a 3×3 convolution and LReLU activation function. The structure of the third layer is the same as that of the first layer. Finally, the output of the first layer is concatenated with the output of the third layer by adding the elements of the first layer and the output of the third layer to obtain the output of the preprocessing layer. After receiving the output of the preprocessing layer, the upper branch UP sequentially processes the output through 1×5 convolution, batch normalization (BN), LReLU activation function, 5×1 convolution, batch normalization (BN), and Sigmoid activation function. Then, it performs element-wise multiplication with the output of the preprocessing layer PA to obtain the output of the upper branch. After receiving the output of the preprocessing layer PA, the lower branch DOWN sequentially processes the output through 1×1 convolution, batch normalization (BN), LReLU activation function, 1×1 convolution, batch normalization (BN), and Sigmoid activation function. Then, it performs element-wise multiplication with the output of the preprocessing layer to obtain the output of the lower branch. Finally, output the UP branch. The output of the lower branch DOWN Output of preprocessing layer PA Adding them together yields the output of the first multi-branch heterogeneous polarization cross-attention module. , represented as: ; ; ; Where F represents the input of the multi-branch heterogeneous polarization cross attention module.

5. The polarization image fusion method based on global perception and multi-branch heterogeneous attention according to claim 1, characterized in that, The channel attention module CA is represented as follows: ; in, This represents the output after passing through the channel attention module. This represents the input to the channel attention module (CA). This represents the Sigmoid activation function. Indicates a fully connected layer. This represents the LRelu activation function. This indicates the average pooling operation. This indicates a max pooling operation.

6. The polarization image fusion method based on global perception and multi-branch heterogeneous attention according to claim 1, characterized in that, The collaborative cross-dimensional module is represented as follows: ; ; ; ; ; in, This represents the output after passing through the channel attention module. This represents the Sigmoid activation function. This indicates that the convolution has been performed using a 3×3 method. This represents the LRelu activation function. Indicates attentional characteristics; This indicates average pooling along the x-axis. This indicates average pooling along the y-axis, and Concat indicates merging across channels. Indicates will Features obtained by merging two types of average pooling on the channel; This indicates that after a 1×1 convolution, This represents the activation function. Indicates batch normalization. Indicates a separation operation. and Indicates intermediate features during the operation process; This indicates the final output obtained after matrix multiplication, which is the output of the collaborative cross-dimensional module.

7. The polarization image fusion method based on global perception and multi-branch heterogeneous attention according to claim 1, characterized in that, The training adopts an unsupervised training method, and the designed loss function includes multi-scale weighted structural loss, intensity loss and significant gradient loss; The multi-scale weighted structural loss , represented as: ; ; in, It is the size of the sliding window. and They represent in × Sliding window size Local mean of the image Local mean of the image and These represent the corresponding local variances. This represents the local covariance of the two input images. and It is the stability coefficient. It is a similarity measurement function. It is the merged image; It is a constant. Pick and The maximum value; The strength loss , represented as: ; ; in, This means that a pixel value less than 0.5 is equal to 0, and a pixel value greater than 0.5 is equal to 1. This means that pixels less than 0 are set to 0, and pixels greater than 0 are set to 1. Indicates upsampling, It is average pooling. express Norm; * indicates element-wise multiplication. Represents intermediate variables in the calculation process. It is a fused image; The significant gradient loss , represented as: ; ; ; in, Represents the gradient operator. , and This represents the computation of S0 image, DoLP image, and fused image. gradient, This indicates taking the maximum value of the gradient magnitudes of the S0 image and the DoLP image. This indicates a mean operation on the image domain; * represents element-wise multiplication. It is used for weight adjustment and filtering out meaningless, non-salient texture regions.

8. A polarization image fusion system with global perception and multi-branch heterogeneous attention, characterized in that, include: The dataset acquisition module acquires polarization images from four angles and uses the Stokes vector synthesis method to obtain a dataset containing intensity images and linear polarization images. The fusion model acquisition module constructs a polarization image fusion model with nonlocal skip connections and multi-branch intersections, and trains it using a dataset containing intensity images and linear polarization degree images to obtain the polarization image fusion model. The polarization image fusion model includes a dual-branch structure and an image reconstruction structure. The dual-branch structure includes a first branch and a second branch, which respectively receive and process the intensity image and the linear polarization degree image. Both the first branch and the second branch include several multi-branch heterogeneous polarization cross-attention modules (MPCA), polarization channel attention modules (CA), and collaborative cross-dimensional modules (CCDM) connected in sequence. Each MPCA performs multi-feature extraction and cross-fusion on the input features, with skip connections in the middle of the MPCA to achieve comprehensive perception. The features output by the last MPCA serve as the input features of the CA, and the output features of the CA serve as the input features of the collaborative cross-dimensional module (CCDM). The collaborative cross-dimensional module (CCDM) performs overall spatial dimension transformation to separate coordinate dimension processing on the input features. The features output by the first branch CCDM and the features output by the second branch CCDM are input into the image reconstruction structure, merged in channels, and reconstructed through three convolutional layers to obtain the fused polarization image. The multi-branch heterogeneous polarization cross-attention module introduces skip connections to achieve global perception in order to compensate for the loss of detail; Both the first branch and the second branch include five multi-branch heterogeneous polarization cross-attention modules; skip connections are set between the output of the preprocessing layer of the first multi-branch heterogeneous polarization cross-attention module and the output of the fifth multi-branch heterogeneous polarization cross-attention module, between the output of the first multi-branch heterogeneous polarization cross-attention module and the output of the fourth multi-branch heterogeneous polarization cross-attention module, and between the output of the second multi-branch heterogeneous polarization cross-attention module and the output of the third multi-branch heterogeneous polarization cross-attention module.

Citation Information

Patent Citations

  • Sparse aperture optical system polarization image fusion method based on deep learning

    CN120355594A

  • Polarized SAR image classification method based on uniform graph guide fusion

    CN118097294A

  • Infrared polarization image super-resolution method based on cross attention double-branch network

    CN120876234A