Sparse aperture optical system polarization image fusion method based on deep learning
The polarized image fusion model of sparse aperture optical system is constructed through deep learning methods, which solves the problem of insufficient imaging quality in sparse aperture optical systems, and effectively fusion of polarization imaging and sparse aperture imaging is achieved, thereby improving imaging contrast and resolution.
Patent Information
- Application Number
- CN202510852339.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art cannot effectively solve the problems of loss of intermediate frequency information of modulation transfer function, decreased imaging contrast and enhanced noise sensitivity in sparse aperture optical systems, especially when polarized images are fused.
The polarization image fusion method of sparse aperture optical system based on deep learning is adopted. By obtaining the multi-angle polarization angle original image of the sparse aperture optical system, linear polarization degree map, polarization angle map, and polarization intensity map are generated, and an encoder, a multi-modal fusion module, and a decoder are built. The edge gradient compensation module, a polarization attention mechanism and a residual aggregation module are used to train the model together with the total loss function to achieve effective fusion of polarized images.
Effectively suppress noise, improve imaging contrast and resolution, realize the effective fusion of polarization imaging and sparse aperture imaging, improve image smoothing and texture weakening problems, and enhance imaging quality in complex scenarios.
Smart Images

Figure CN120355594A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a polarization image fusion method for a sparse aperture optical system based on deep learning. Background Art
[0002] The sparse aperture optical system realizes the interference synthesis of an equivalent large aperture by arranging multiple sub - mirrors in a non - redundant manner, breaks through the physical constraints in the way of coherent superposition of sub - apertures, significantly reduces the volume and weight of the system, and at the same time maintains the resolution ability close to that of a complete aperture, solving the volume and cost problems of traditional large - aperture optical systems while achieving high - resolution imaging.
[0003] However, the sparse structure of the sparse aperture optical system causes the loss of intermediate - frequency information of the modulation transfer function and a decrease in imaging contrast; moreover, the sparse aperture optical system has an enhanced sensitivity to noise, specifically manifested as defects such as blurred image details and insufficient edge sharpness. In order to improve these defects, existing technologies have tried to improve the imaging contrast of sparse apertures and process noise by starting from aspects such as structural optimization and image restoration, but these methods do not consider the influence of the polarization factors of the target object on imaging and image restoration.
[0004] Polarization imaging technology provides a physical property dimension other than light intensity for target recognition by analyzing the polarization state information of light, showing unique advantages in complex scenes. However, most current polarization image fusions are for highlighting the target to obtain a fusion image with prominent target features, rather than for improving imaging quality. And when performing polarization image fusion, existing technologies only aim to solve the noise existing in polarization images and cannot effectively solve the noise caused by sparse aperture imaging. For example, a full - polarization image fusion method based on an auto - encoder in the prior art (patent publication number CN115393233B) can reduce polarization blur caused by material properties and illumination environments and retain and enhance the polarization information of the target, but it also cannot effectively solve problems such as image smoothing and texture weakening caused by the sparse aperture structure. Therefore, the prior art cannot achieve the fusion of polarization imaging and sparse aperture imaging, and the imaging quality is not high. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide a polarization image fusion method for a sparse aperture optical system based on deep learning, which can effectively fuse polarization imaging and sparse aperture imaging, effectively suppress noise, and improve the contrast and resolution of imaging.
[0006] To solve the above - mentioned technical problem, the present invention provides a polarization image fusion method for a sparse aperture optical system based on deep learning, including: Obtain the original multi - angle polarization angle images collected by the sparse aperture optical system, and generate the degree of linear polarization map, polarization angle map, and polarization intensity map; Construct a polarization image fusion model, which includes an encoder, a multi - modal fusion module, and a decoder; the encoder includes a first - branch encoder and a second - branch encoder. The first - branch encoder extracts the global semantic features of the polarization angle map and the polarization intensity map, and the second - branch encoder extracts the local detailed features of the degree of linear polarization map and the polarization intensity map; the multi - modal fusion module includes an edge gradient compensation module, a polarization attention mechanism, and a residual aggregation module. The edge gradient compensation module extracts multi - level edge features, the polarization attention mechanism adaptively weights and fuses the features, and the residual aggregation module obtains the aggregated features by retaining the original features through skip connections. The decoder uses multi - kernel deconvolution to decode the aggregated features; Construct a total loss function in combination with the characteristics of the degree of linear polarization map and the polarization intensity map, and train the polarization image fusion model. Input the degree of linear polarization map, polarization angle map, and polarization intensity map to be fused into the trained polarization image fusion model to obtain a polarization fusion image.
[0007] Further, when the second - branch encoder extracts the local detailed features of the polarization intensity map, specifically: Perform a 3×3 convolution on the polarization intensity map, and denote the obtained feature as F S0_init ; Perform a 3×3 convolution on F S0_init , and denote the obtained feature as F S0_16 ; Concatenate F S0_init and F S0_16 channel - wise, and denote the obtained feature as F S0_32_ ; Perform a 3×3 convolution on F S0_32_ , and denote the obtained feature as F S0_32 ; Concatenate F S0_32 , F S0_16 and F S0_init channel - wise, and denote the obtained feature as F S0_48_ ; Perform a 3×3 convolution on F S0_48_ , and denote the obtained feature as F S0_48 ; Concatenate F S0_48 , F S0_32 , F S0_16 and F S0_init channel - wise to obtain the local detailed features of the polarization intensity map.
[0008] Further, when the edge gradient compensation module extracts multi - level edge features, specifically: Use the Sobel operator to extract the horizontal gradient features and vertical gradient features of the features extracted by the first-branch encoder and the second-branch encoder respectively. Use the Laplace operator to extract the high-frequency detail features in all directions of the features extracted by the first-branch encoder and the second-branch encoder respectively. Concatenate the features extracted by the Sobel operator and the features extracted by the Laplace operator along the channel dimension, and obtain multi-level edge features through two 3×3 convolutions.
[0009] Furthermore, the encoder further includes a cross-branch feature interaction and fusion module, which fuses the features extracted by the first-branch encoder to generate the fused features of the first-branch encoder, and fuses the features extracted by the second-branch encoder to generate the fused features of the second-branch encoder. When the polarization attention mechanism adaptively weights and fuses features, it combines channel attention and spatial attention. Specifically: Perform channel attention operation on the fused features of the first-branch encoder using channel attention, and obtain features through 1×1 convolution and 3×3 convolution, denoted as F enc1 ’; Perform spatial attention operation on the fused features of the second-branch encoder using spatial attention, and obtain features through 1×1 convolution and 3×3 convolution, denoted as F enc2 ’; Fuse the features extracted by channel attention and spatial attention to obtain polarization attention features, denoted as F, F = concat(F enc1 ’ + F enc2 ’), where concat( ) is the concatenation operation.
[0010] Furthermore, the residual aggregation module obtains the aggregated features by retaining the original features through skip connections. Specifically: Denote the features obtained by passing the polarization attention features through 1×1 convolution as F’, and the features obtained by passing through two 3×3 convolutions as ; Denote the aggregated features as F Agg , F Agg = F edge + F’ + , where F edge is the multi-level edge features extracted by the edge gradient compensation module.
[0011] Furthermore, the total loss function is: L = L MSWSSIM + α L mse + β L mae +γ L edge +L penalty , wherein, L is the total loss function, L MSWSSIM is the multi-size window loss function, L mse is the pixel loss function combining the degree of linear polarization map and the polarization intensity map, L mae is the pixel loss function combining the polarization intensity map, L edge is the edge loss function, L penalty is the adaptive penalty term, α 、 β 、 γ are the balance parameters.
[0012] Furthermore, the multi-size window loss function is: , wherein, N is the number of window sizes, ni is the window corresponding to the i th window size, is the weight coefficient, w is { n 1,…, ni, …, nN}, any window in ( ) is the structural loss function, S 0 is the polarization intensity map, DoLP is the degree of linear polarization map, is the predicted image of the polarization image fusion model.
[0013] Furthermore, is calculated as: , wherein, is the mean value of , is the region of the polarization intensity map within the window w ; is the mean value of , is the region of the predicted image of the polarization image fusion model within the window w ; is a preset constant, is the covariance of and , is the variance of , is the variance of ; The calculation method of the weight coefficient is: , Among them, is the variance of is the region of the linear polarization degree map within the window w , is a correction function.
[0014] Furthermore, the pixel loss function combining the linear polarization degree map and the polarization intensity map is: , where is the i th pixel in the polarization intensity map, is the i th pixel in the linear polarization degree map, is the i th pixel in the predicted image of the polarization image fusion model, m is the number of pixels.
[0015] Furthermore, the pixel loss function combining the polarization intensity map is: , where is the i th pixel in the polarization intensity map, is the i th pixel in the predicted image of the polarization image fusion model, m is the number of pixels.
[0016] The above technical solutions of the present invention have the following beneficial effects compared with the prior art: The present invention obtains the linear polarization degree map, the polarization angle map, and the polarization intensity map through a sparse aperture optical system, and respectively extracts the global semantic features of the polarization angle map and the polarization intensity map, and the local detail features of the linear polarization degree map and the polarization intensity map through a polarization image fusion model with a double-branch structure; at the same time, the edge gradient compensation layer is used to strengthen the target contour, the attention mechanism is combined to dynamically allocate feature weights, and the original information is retained through residual connection, improving problems such as image smoothing and texture weakening caused by the sparse aperture structure, and effectively suppressing noise; on this basis, the image resolution is gradually restored through the deconvolution decoding structure, and the polarization image fusion model is trained in combination with the characteristics of the polarization image, and finally polarization fusion imaging with richer information in complex scenes is realized, realizing the effective fusion of polarization imaging and sparse aperture imaging, effectively suppressing noise, and improving the contrast and resolution of imaging. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to make the content of the present invention easier to be clearly understood, the following further describes the present invention in detail according to the specific embodiments of the present invention in conjunction with the drawings, where: Figure 1 It is a flowchart of the method in the preferred embodiment of the present invention.
[0018] Figure 2 It is a structural diagram of the method in the preferred embodiment of the present invention.
[0019] Figure 3 It is a structural diagram of the first branch encoder and the second branch encoder in the preferred embodiment of the present invention.
[0020] Figure 4 It is a structural diagram of the multi-modal fusion module in the preferred embodiment of the present invention.
[0021] Figure 5 It is a structural diagram of each level of transposed convolution of the decoder in the preferred embodiment of the present invention.
[0022] Figure 6 It is a result diagram of fusing indoor scene images by different methods in the simulation experiment of the preferred embodiment of the present invention.
[0023] Figure 7 It is a result diagram of fusing outdoor scene images by different methods in the simulation experiment of the preferred embodiment of the present invention. Detailed implementation manners
[0024] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments given are not intended to limit the present invention.
[0025] Refer to Figure 1 、 Figure 2 As shown, the present invention discloses a polarization image fusion method for a sparse aperture optical system based on deep learning, including the following steps: S1: Obtain the original multi-angle polarization angle images collected by the sparse aperture optical system, and generate a degree of linear polarization map (DoLP), an angle of polarization map (AoP), and a polarization intensity map ( S 0 ).
[0026] The polarization imaging system can synchronously obtain the light intensity of the scene ( S 0), polarization angle (AoP), and degree of linear polarization (DoLP) information. These three types of data are significantly complementary in target characterization. The intensity image reflects the macroscopic details of the scene through pixel intensity and spatial distribution, but it is vulnerable to environmental light interference and difficult to distinguish material properties. The polarization angle and degree of linear polarization, as physical properties of the interaction between light and objects, are directly related to characteristics such as the shape of the target surface and material roughness. The degree-of-polarization image can enhance the contrast between the target and the background through brightness differences, but it does not contain the spatial distribution information of the scene itself. Therefore, fusing the intensity and degree-of-polarization images can synergistically utilize the advantages of both: the intensity data provides the spatial details and texture information of the scene, while the degree-of-polarization data enhances the target edges and material features through physical property differences, ultimately generating a fused image with both high information integrity and target saliency, providing reliable support for accurate detection and recognition in complex scenes.
[0027] S1-1: In this embodiment, four groups of original polarization angle images at 0°, 45°, 90°, and 135° collected by a sparse aperture optical system are obtained. The resolution of a single-frame image is set to 40×40 pixels, and the bit depth is 8 bits.
[0028] S1-2: Generate the degree-of-linear-polarization map (DoLP), polarization-angle map (AoP), and polarization-intensity map ( S 0 ): S1-2-1: Calculate the Stokes vector based on the four groups of original polarization angle images: ; Among them, S 0 is the polarization-intensity map, I 0 is the original 0° polarization angle image, I 90 is the original 90° polarization angle image, I 45 is the original 45° polarization angle image, I 135 is the original 135° polarization angle image.
[0029] S1-2-2: Calculate the degree-of-linear-polarization map and the polarization-angle map as: ; Among them, DoLP is the degree-of-linear-polarization map, and AoP is the polarization-angle map.
[0030] S1-3: Perform data normalization and three-channel stitching to generate the input data for the polarization image fusion model. In this embodiment, S 0, The DoLP, AoP three-channel data is normalized to the range of [0, 1] and concatenated along the channel dimension into an input tensor of 1024×1024×3.
[0031] S2: Construct a polarization image fusion model, which includes an encoder, a multi-modal fusion module, and a decoder.
[0032] S2-1: The encoder includes a first-branch encoder, a second-branch encoder, and a cross-branch feature interaction and fusion module. The first-branch encoder performs multi-level convolution and downsampling operations on the polarization angle map and the polarization intensity map to extract the global semantic features of the scene. The second-branch encoder performs multi-scale extended convolution on the degree of linear polarization map and the polarization intensity map respectively to extract the local detailed features of the target. The cross-branch feature interaction and fusion module fuses the features extracted by the first-branch encoder and the second-branch encoder respectively to generate the fusion features of the first-branch encoder with 128 channels and the fusion features of the second-branch encoder.
[0033] S2-1-1: In this embodiment, as Figure 3 shown, the feature extraction process of the first-branch encoder is specifically as follows: S2-1-1-1: Perform 3×3 convolution operations with four levels on the polarization angle map and the polarization intensity map respectively, with step sizes of 1, 2, 2, 2 respectively, and the number of channels changes successively as 32, 64, 128, 64; after each level of convolution operation, connect the LeakyReLU activation function with a negative axis slope of 0.2 to obtain the 64-channel features of the polarization angle map and the 64-channel features of the polarization intensity map.
[0034] S2-1-1-2: Perform bilinear upsampling by 8 times on the 64-channel features of the polarization angle map to restore to the original resolution, denoted as F Aop ; perform bilinear upsampling by 8 times on the 64-channel features of the polarization intensity map to restore to the original resolution, denoted as F S0 .
[0035] S2-1-2: In this embodiment, as Figure 3 shown, the feature extraction process of the second-branch encoder is specifically as follows: S2-1-2-1: Extract the features of the polarization intensity map, specifically: (1) The input S 0 image undergoes 3×3 convolution (step size = 1), and the output 16-channel features are denoted as F S0_init ; (2) Perform 3×3 convolution (step size = 1) on F S0_init , and the output 16-channel features are denoted as F S0_16 ; (3) Combine F S0_init with FS0_16 Concatenate by channel to generate 32-channel features, denoted as F S0_32_ ; (4)For F S0_32_ Perform 3×3 convolution (stride = 1), and output 16-channel features, denoted as F S0_32 ; (5)Concatenate F S0_32 , F S0_16 and F S0_init by channel to generate 48-channel features, denoted as F S0_48_ ; (6)For F S0_48_ Perform 3×3 convolution (stride = 1), and output 16-channel features, denoted as F S0_48 ; (7)Concatenate F S0_48 , F S0_32 , F S0_16 and F S0_init by channel to generate 64-channel features, and obtain the local detail features of the polarization intensity map, denoted as F S0 ’.
[0036] S2-1-2-2: Extract the features of the linear polarization degree map, which is symmetric to the feature extraction path design of the polarization intensity map, and generate the final 64-channel features through the same-level convolution and concatenation operations, and obtain the local detail features of the linear polarization degree map, denoted as F DoLP .
[0037] S2-1-3: The cross-branch feature interaction and fusion module concatenates the 64-channel features of the polarization angle map and the 64-channel features of the polarization intensity map extracted by the first-branch encoder in the channel dimension to generate the fusion features of the first-branch encoder with 128 channels, denoted as F enc1= concat(F Aop , F S0 ).
[0038] The cross-branch feature interaction and fusion module concatenates the 64-channel features of the linear polarization degree map and the 64-channel local features of the polarization intensity map extracted by the second-branch encoder in the channel dimension to generate the fusion features of the second-branch encoder with 128 channels, denoted as F enc2= concat(F DoLP , F S0 ’).
[0039] S2-2: The multi-modal fusion module includes an edge gradient compensation module, a polarization attention mechanism, and a residual aggregation module, as Figure 4As shown, the edge gradient compensation module extracts multi-level edge features based on the Sobel operator and the Laplace operator. The polarization attention mechanism adaptively weights and fuses features by combining channel attention and spatial attention. The residual aggregation module retains the original features through skip connections and outputs high-resolution aggregated features with 64 channels. Figure 4 In " Figure 4 ", "C" represents the concatenation operation along the channel dimension, and "+" represents the feature addition operation.
[0040] S2-2-1: The edge gradient compensation module extracts multi-level edge features to enhance the target contour and material boundary. Specifically: S2-2-1-1: Use the Sobel operator to extract the horizontal gradient features and vertical gradient features of F Aop , F S0 , F DoLP , F S0 ', respectively. The convolution kernels are: ; Among them, G x is the convolution kernel in the horizontal direction, G y is the convolution kernel in the vertical direction.
[0041] S2-2-1-2: Use the Laplace operator to extract the high-frequency detail features in all directions of F Aop , F S0 , F DoLP , F S0 ', respectively. The convolution kernel is .
[0042] S2-2-1-3: Concatenate the features extracted by the Sobel operator and the Laplace operator along the channel dimension to obtain 256-channel features, and then compress them into 64-channel features through two 3×3 convolutions to obtain the multi-level edge features extracted by the edge gradient compensation module, denoted as F edge .
[0043] S2-2-2: The polarization attention mechanism adaptively weights and fuses features by combining channel attention (focusing on material differences) and spatial attention (focusing on the target area). Specifically: S2-2-2-1: Perform channel attention operation on the fused feature F enc1 of the first branch encoder using channel attention. The feature obtained after passing through 1×1 convolution and 3×3 convolution is denoted as F enc1 '.
[0044] S2-2-2-2: Perform spatial attention on the fused feature F enc2Perform spatial attention operation, and denote the feature obtained through 1×1 convolution and 3×3 convolution as F enc2 ’.
[0045] S2-2-2-3: Fuse the features extracted by channel attention and spatial attention to obtain polarization attention features: F = concat(F enc1 ’ + F enc2 ’), where F is the polarization attention feature, and concat( ) is the concatenation operation.
[0046] S2-2-2-4: Denote the feature obtained by 1×1 convolution of the polarization attention feature F as F’, and then denote the feature obtained by two 3×3 convolutions as .
[0047] S2-2-3: The high-resolution aggregation feature output by the residual aggregation module is: F Agg = F edge + F’ + .
[0048] S2-3: The decoder uses multi-core transposed convolution to decode the aggregation feature and gradually restore the image resolution; as Figure 5 shown, in this embodiment, the specific structure of the multi-core transposed convolution is: S2-3-1: The first-level transposed convolution: The input is a 64-channel feature, and 3×3, 1×3, and 3×1 transposed convolution operations are respectively performed, with a stride = 1 and padding = 1, and a 48-channel feature is output.
[0049] S2-3-2: The second-level transposed convolution: The structure is the same as that of the first-level transposed convolution. The input is a 48-channel feature, and the output is a 32-channel feature, retaining multi-scale details.
[0050] S2-3-3: The third-level transposed convolution: The structure is the same as that of the first-level transposed convolution. The input is a 32-channel feature, and the output is a 16-channel feature, suppressing noise interference.
[0051] S2-3-4: Output layer: 1×1 transposed convolution compresses the 16-channel feature to a single channel, and generates the final single-channel polarization fusion image through the Sigmoid activation function.
[0052] S3: Construct a total loss function in combination with the characteristics of the linear polarization degree map and the polarization intensity map, and train the polarization image fusion model.
[0053] When training the polarization image fusion model, the constructed total loss function is: L = L MSWSSIM + α Lmse + β L mae + γ L edge +L penalty , where L is the total loss function, L MSWSSIM is the multi-scale window loss function, L mse is the pixel loss function that combines the degree of linear polarization map and the polarization intensity map, L mae is the pixel loss function that combines the polarization intensity map, L edge is the edge loss function, L penalty is the adaptive penalty term, α 、 β 、 γ are the balance parameters. In this embodiment α = 1, β = 0.1, γ = 0.1.
[0054] The multi-scale window loss function considers the optimized structural similarity under multiple window sizes, specifically: , where N is the number of window sizes, ni is the window corresponding to the i -th window size, is the weight coefficient, w is { n 1,…, ni, …, nN} in any window, ( ) is the structural loss function, S 0 is the polarization intensity map, DoLP is the degree of linear polarization map, is the predicted image of the polarization image fusion model. In this embodiment, the optimized structural similarity is carried out at 5 scales, that is N = 5, and the 5 window sizes are 3×3 (i.e., n 1 window size), 5×5 (i.e., n 2 window size), 7×7 (i.e., n 3 window size), 9×9 (i.e., n 4 window size), 11×11 (i.e., n 5 window size), and the standard deviation of the Gaussian kernel
[0055] The calculation method of , where is The mean value of is the region of the polarization intensity map within the window w ; is The mean value of is the region of the predicted image of the polarization image fusion model within the window w ; is a preset constant. In this embodiment, are respectively 1×10 -4 and 9×10 -4 ; is and The covariance of is The variance of is The variance of
[0056] The calculation method of is the same as the principle of the calculation method of
[0057] The calculation method of the weight coefficient is: where where is The variance of is the region of the linear polarization degree map within the window w ; is a correction function. In this embodiment, =max(x, 0.0001), which is used to enhance the robustness.
[0058] The pixel loss function that combines the linear polarization degree map and the polarization intensity map is: where where is the i th pixel in the polarization intensity map, is the i th pixel in the linear polarization degree map, is the i th pixel in the predicted image of the polarization image fusion model, m is the number of pixels.
[0059] ;
[0060] The present invention can obtain better robustness and better fuse different types of features by combining L mse and L mae .
[0061] The edge loss function detects the edges of an image through the Sobel operator and calculates the MAE loss between the predicted image and the edge feature map of the real image, specifically as follows: , where is the i th pixel in the real image, is the axis coordinate of x after passing through the Sobel operator, is the axis coordinate of y after passing through the Sobel operator.
[0062] By constraining the Sobel gradient difference, the target edges are strengthened, enabling the model to retain the edge features of the original image during prediction and improving the quality of the generated image.
[0063] When the single-scale SSIM is lower than 0.7, a linear penalty is imposed, and the adaptive penalty term , where SSIM is the structural similarity between the real image and the predicted image.
[0064] During the training process, Adam is used for backpropagation optimization with a learning rate of 1e-4, β1 = 0.9, β2 = 0.999, the number of training epochs is set to 30, and the batch size is set to 128.
[0065] S4: Input the linear polarization degree map, polarization angle map, and polarization intensity map to be fused into the trained polarization image fusion model to obtain the predicted polarization fusion image.
[0066] The present invention also discloses a polarization image fusion system for a sparse aperture optical system based on deep learning, including a data acquisition module, a model construction module, a training module, and a fusion module.
[0067] The data acquisition module acquires the multi-angle polarization angle original images collected by the sparse aperture optical system and generates a linear polarization degree map, a polarization angle map, and a polarization intensity map.
[0068] The model construction module constructs a polarization image fusion model, which includes an encoder, a multimodal fusion module, and a decoder; the encoder includes a first-branch encoder and a second-branch encoder. The first-branch encoder extracts the global semantic features of the polarization angle map and the polarization intensity map, and the second-branch encoder extracts the local detail features of the degree of linear polarization map and the polarization intensity map; the multimodal fusion module includes an edge gradient compensation module, a polarization attention mechanism, and a residual aggregation module. The edge gradient compensation module extracts multi-level edge features, the polarization attention mechanism adaptively weights and fuses features, and the residual aggregation module obtains aggregated features by retaining the original features through skip connections. The decoder decodes the aggregated features using multi-core transposed convolution.
[0069] The training module constructs a total loss function in combination with the characteristics of the degree of linear polarization map and the polarization intensity map, and trains the polarization image fusion model to obtain a trained polarization image fusion model.
[0070] The fusion module inputs the degree of linear polarization map, the polarization angle map, and the polarization intensity map to be fused into the trained polarization image fusion model to obtain a polarization fusion image.
[0071] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a polarization image fusion method for a sparse aperture optical system based on deep learning.
[0072] The present invention also discloses a device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a polarization image fusion method for a sparse aperture optical system based on deep learning.
[0073] The present invention combines the high-resolution advantage of the sparse aperture system with the multi-dimensional information analysis ability of polarization imaging, and realizes the effective fusion of polarization imaging and sparse aperture imaging through deep learning methods. Compared with the prior art, the advantages of the present invention are: 1. The present invention obtains the degree of linear polarization map, the polarization angle map, and the polarization intensity map through a sparse aperture optical system, and respectively extracts the global semantic features of the polarization angle map and the polarization intensity map, and the local detail features of the degree of linear polarization map and the polarization intensity map through a polarization image fusion model with a dual-branch structure, overcoming the defect of ignoring physical relevance in traditional methods.
[0074] 2. The edge gradient compensation layer is used to strengthen the target contour, the attention mechanism is combined to dynamically allocate feature weights, and the original information is retained through residual connections, improving problems such as image smoothing and texture weakening caused by the sparse aperture structure, and effectively suppressing noise.
[0075] 3. Gradually restore the image resolution through a hybrid deconvolution structure combined with multi-core operations, effectively suppressing motion blur and noise interference, and finally realizing polarization fusion imaging with richer information in complex scenes, achieving an effective fusion of polarization imaging and sparse aperture imaging, and improving the contrast and resolution of imaging.
[0076] 4. Train a polarization image fusion model in combination with the characteristics of polarization images, and optimize network parameters by jointly using multi-scale structural similarity loss and edge-preserving loss, etc., to further improve the contrast and resolution of imaging.
[0077] To further prove the advantages of the present invention, in this embodiment, on a hardware platform with NVIDIA 4060Ti GPU and 16GB video memory, the present invention and existing CVT (Convolutions to Vision Transformers), RP (Rapid Prototyping Model), PCNN (PCNN - Pulse Coupled Neural Network), PFNet (see the paper "Zhang J C, Shao J B, Chen J L, et al. PFNet: an unsupervised deep network for polarization image fusion[J]. Optics Letters, 2020, 45(6): 1507 - 1510.") models are respectively used to conduct fusion simulation experiments on indoor scene images and outdoor scene images.
[0078] Three indicators of information entropy, standard deviation, and multi-scale structural similarity are used to evaluate the effect of image fusion. The results of fusing indoor scene images by different methods are shown in Table 1, and the results of fusing outdoor scene images by different methods are shown in Table 2.
[0079] Table 1 Results table of fusing indoor scene images by different methods
[0080] Table 2 Results table of fusing outdoor scene images by different methods
[0082] As can be seen from Table 1 and Table 2, in both indoor and outdoor scenes, all indicators of the present invention are better than those of other existing models.
[0083] The result diagram of fusing indoor scene images by different methods is as Figure 6 shown, Figure 6 in which (a) is the imaging result diagram of CVT;Figure 6 In (b) is the RP imaging result diagram, Figure 6 In (c) is the PCNN imaging result diagram, Figure 6 In (d) is the PFNet imaging result diagram, Figure 6 In (e) is the imaging result diagram of the present invention. The result diagrams of fusing outdoor scene images by different methods are as Figure 7 shown, Figure 7 In (a) is the CVT imaging result diagram, Figure 7 In (b) is the RP imaging result diagram, Figure 7 In (c) is the PCNN imaging result diagram, Figure 7 In (d) is the PFNet imaging result diagram, Figure 7 In (e) is the imaging result diagram of the present invention. From Figure 6 , Figure 7 it can be seen that the imaging of the present invention is clearer, and the contrast and resolution are better than those of other models.
[0084] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0085] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 block or multiple blocks.
[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 block or multiple blocks.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or one block or a plurality of blocks. Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps of the functions specified in one block or a plurality of blocks.
[0088] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. The obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A polarization image fusion method for a sparse aperture optical system based on deep learning, characterized in that Including: Obtain the original multi-angle polarization angle images collected by the sparse aperture optical system, and generate a degree of linear polarization map, a polarization angle map, and a polarization intensity map; Construct a polarization image fusion model, which includes an encoder, a multi-modal fusion module, and a decoder; the encoder includes a first-branch encoder and a second-branch encoder. The first-branch encoder extracts the global semantic features of the polarization angle map and the polarization intensity map, and the second-branch encoder extracts the local detail features of the degree of linear polarization map and the polarization intensity map; the multi-modal fusion module includes an edge gradient compensation module, a polarization attention mechanism, and a residual aggregation module. The edge gradient compensation module extracts multi-level edge features, the polarization attention mechanism adaptively weights and fuses the features, and the residual aggregation module obtains the aggregated features by retaining the original features through skip connections. The decoder uses multi-core deconvolution to decode the aggregated features; Construct a total loss function in combination with the characteristics of the degree of linear polarization map and the polarization intensity map, and train the polarization image fusion model. Input the degree of linear polarization map, polarization angle map, and polarization intensity map to be fused into the trained polarization image fusion model to obtain a polarization fusion image.
2. The polarization image fusion method of a sparse aperture optical system based on deep learning according to claim 1, wherein: When the second-branch encoder extracts the local detail features of the polarization intensity map, specifically: Perform a 3×3 convolution on the polarization intensity map, and the resulting feature is denoted as F S0_init ; For F S0_init Perform a 3×3 convolution, and the resulting feature is denoted as F S0_16 ; Concatenate F S0_init and F S0_16 by channel, and denote the resulting feature as F S0_32_ ; Perform a 3×3 convolution on F S0_32_ and denote the resulting feature as F S0_32 ; Concatenate F S0_32 , F S0_16 and F S0_init along the channel dimension, and denote the resulting feature as F S0_48_ ; For F S0_48_ Perform a 3×3 convolution, and the resulting feature is denoted as F S0_48 ; Combine F S0_48 、F S0_32 、F S0_16 and F S0_init by channels to obtain the local detail features of the polarization intensity map.
3. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 1, characterized in that: When the edge gradient compensation module extracts multi-level edge features, specifically: Use the Sobel operator to extract the horizontal gradient features and vertical gradient features of the features extracted by the first-branch encoder and the second-branch encoder respectively; Use the Laplace operator to extract the high-frequency detail features in all directions of the features extracted by the first-branch encoder and the second-branch encoder respectively; Concatenate the features extracted by the Sobel operator and the features extracted by the Laplace operator by channel, and obtain multi-level edge features through two 3×3 convolutions.
4. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 1, characterized in that: The encoder further includes a cross-branch feature interaction and fusion module, which fuses the features extracted by the first-branch encoder to generate the fused features of the first-branch encoder, and fuses the features extracted by the second-branch encoder to generate the fused features of the second-branch encoder; When the polarization attention mechanism adaptively weights and fuses the features, it combines channel attention and spatial attention, specifically: Perform channel attention operation on the fused features of the first branch encoder using channel attention, and obtain features through 1×1 convolution and 3×3 convolution, denoted as F enc1 ’; Perform a spatial attention operation on the fused features of the second branch encoder using spatial attention, and obtain features through 1×1 convolution and 3×3 convolution, denoted as F enc2 ’; The features extracted by fusing channel attention and spatial attention are obtained as polarization attention features, denoted as F, where F = concat(F enc1 ’ + F enc2 ’), and concat( ) is the concatenation operation.
5. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 4, characterized in that: When the residual aggregation module obtains the aggregated features by retaining the original features through skip connections, specifically: Denote the feature obtained by performing 1×1 convolution on the polarization attention feature as F’, and the feature obtained by further performing two 3×3 convolutions as ; Denote the aggregated feature as F Agg , F Agg = F edge + F'+ , where F edge is the multi-level edge features extracted by the edge gradient compensation module.
6. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 1, characterized in that: The total loss function is: L = L MSWSSIM + α L mse + β L mae + γ L edge + L penalty , Among them, \(L\) is the total loss function, \(L\) MSWSSIM is the multi-size window loss function, \(L\) mse is the pixel loss function combining the degree of linear polarization map and the polarization intensity map, \(L\) mae is the pixel loss function combining the polarization intensity map, \(L\) edge is the edge loss function, \(L\) penalty is the adaptive penalty term, α 、 β 、 γ are the balance parameters.
7. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 6, characterized in that: The multi-size window loss function is: , Among them, N is the number of window sizes, ni is the window corresponding to the i -th window size, is the weight coefficient, w is { n 1, …, ni, …, nN }, any window in it, ( ) is the structural loss function, S 0 is the polarization intensity map, DoLP is the degree of linear polarization map, is the predicted image of the polarization image fusion model.
8. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 7, characterized in that: The calculation method is as follows: , Among them, is the mean value of which is the region of the polarization intensity map within the window w ; is the mean value of which is the region of the predicted image of the polarization image fusion model within the window w ; is a preset constant, is the covariance of and is the variance of is the variance of The calculation method of the weight coefficient is: , Among them, is the variance of is the region of the linear polarization degree map within the window w . is the correction function.
9. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 7, characterized in that: The pixel loss function that combines the degree of linear polarization map and the polarization intensity map is: , Among them, is the i th pixel in the polarization intensity map, is the i th pixel in the degree of linear polarization map, is the i th pixel in the predicted image of the polarization image fusion model, m is the number of pixels.
10. A polarization image fusion method for a sparse aperture optical system based on deep learning according to claim 7, characterized in that: The pixel loss function that combines the polarization intensity map is: , Among them, is the i th pixel in the polarization intensity map, is the i th pixel in the predicted image of the polarization image fusion model, m is the number of pixels.
Citation Information
Patent Citations
Full-linear polarization image fusion method based on auto-encoder
CN115393233A
Underwater robot vision sharpening method based on multi-modal fusion network
CN119478648A
Detection method using fusion network based on attention mechanism, and terminal device
US11222217B1
Cited By
Polarization image fusion method and system based on global perception and multi-branch heterogeneous attention
CN121120426A
Polarization image fusion method and system with global perception and multi-branch heterogeneous attention
CN121120426B
Underwater target measurement method, system and equipment based on salient region sparse matching
CN121147728A
Polarization image fusion system and method
CN121437292A
Image edge extraction method based on TiS3 / MoS2 heterojunction polarization sensitive photoelectric detector, detector and preparation method of detector
CN122473214A