A full linear polarization image fusion method based on autoencoder

Through the fully linear polarization image fusion method based on the autoencoder, the deep feature extraction and fusion technology is used to solve the problem of poor polarization image fusion effect in the prior art, and more efficient complex object detection and information reconstruction are achieved.

CN115393233BActive Publication Date: 2025-05-23NORTHWEST A & F UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210878279.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-25
Publication Date
2025-05-23
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

The existing polarized image fusion technology is not ideal in complex object detection, especially in outdoor natural environments, which makes it difficult to effectively denoise and retain image details.

Method used

The full linear polarization image fusion method based on the autoencoder is adopted to obtain the light intensity, linear polarization degree and linear polarization angle information through polarization imaging technology, and the convolutional neural network is used to deeply extract and fuse this information to generate a more efficient fusion image.

Benefits of technology

It improves the accuracy and contrast of complex object detection, suppresses polarization noise, and can effectively reconstruct information in complex scenarios, and is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393233B_ABST
    Figure CN115393233B_ABST
Patent Text Reader

Abstract

The present invention discloses a full linear polarization image fusion method based on an autoencoder. First, Stokes solution is performed on different polarization state images obtained by a focal plane polarization sensor to obtain Stokes vectors, which are used to obtain the mapping of linear polarization degree and linear polarization angle to form a new polarization feature image, and combined with the light intensity image to obtain a set of new data paradigms, and then an autoencoder based on a convolutional neural network is used to extract features, fuse features, and reconstruct images from the light intensity information and polarization feature image of the target. The present invention can minimize the polarization blur caused by material properties and lighting environment, and is robust to various scenes. The image fusion method of the autoencoder can fully extract and fuse features, retain and enhance the polarization information of the target, and improve the target detection capability under complex backgrounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of polarization image processing and information fusion, and specifically relates to a full-line polarization image fusion method based on an autoencoder. Background Art

[0002] Compared with traditional imaging technology that can only obtain the intensity image and spectral information of the target, polarization imaging can obtain the target's stress properties, birefringence properties, roughness, illumination information, edge information, surface orientation, etc., and is more suitable for underwater detection, atmospheric remote sensing detection, and camouflaged targets. Therefore, polarization imaging, as an important means of obtaining target image information, is currently widely used in military targets, environmental monitoring, biomedical testing and other fields.

[0003] Polarization imaging fusion technology is a new type of photoelectric imaging detection technology that uses polarization feature information and light intensity information to achieve target detection. It can reveal the multi-dimensional characteristics of the target and enhance the contrast of the target. Polarization features can be used to obtain intensity, linear polarization degree, and linear polarization angle images, and then obtain details such as the structure, roughness, and shadow of the target. At present, the research on polarization image fusion mainly maps the intensity image, linear polarization degree image, and linear polarization angle image to pseudo-color space to provide more polarization information. However, this type of research is suitable for human eye observation and is difficult to extend to target detection in the field of machine vision. Another recent popular related research is to extract target features from light intensity images and linear polarization degree images, and reconstruct and restore images through multi-scale decomposition, sparse representation, or deep learning methods. However, linear polarization images have low brightness and do not have advantages in underwater detection and other fields, making them difficult to be widely used in complex scenes. Since the linear polarization degree map and the linear polarization angle image are sensitive to the lighting environment and are closely related to the physical properties of the target, they are prone to significant polarization blur in natural scenes. Therefore, the current polarization image fusion based on the information dimension first needs to denoise the polarization image, but the effect is still not ideal, and denoising may even lose some image details. It is still helpless in outdoor natural environments (outdoor scene noise is very significant). How to extract polarization information from multiple polarization states of complex targets and fuse it with the light intensity image to obtain good results, thereby improving the target detection capability under complex backgrounds is the key to the application of polarization technology. Summary of the invention

[0004] In order to solve the above-mentioned technical problems, the present invention designs a full linear polarization image fusion method based on autoencoder, which utilizes polarization imaging technology and an autoencoder based on a convolutional neural network to extract deep features of light intensity, linear polarization degree, and linear polarization angle information, thereby realizing the fusion and enhancement of full linear polarization images and improving the target detection performance under complex targets.

[0005] In order to solve the above-mentioned technical problems, the present invention adopts the following solutions:

[0006] A full linear polarization image fusion method based on an autoencoder comprises the following steps:

[0007] Step 1: Perform Stokes solution on the polarization image obtained by polarization imaging technology to obtain four Stokes parameters S 0 , S 1 , S 2 , S 3 , the formula is as follows:

[0008]

[0009] And determine the linear polarization degree image DoLP and the linear polarization angle image AoLP according to the Stokes vector of the polarization image:

[0010]

[0011]

[0012] Among them, S 0 represents the intensity image, S 1 Represents the linear or vertical polarization component image, S 2 Represents the linear polarization component image at 45 degrees or 135 degrees, S 3 Represents the right-handed or left-handed polarization component image in the light beam. In the natural environment, the circular polarization component is very small and can be ignored; I 0 ,I 45 , I 90 ,I 135 Respectively represent the polarization images of the outgoing light intensity at four polarization directions of 0°, 45°, 90°, and 135°, I L and I R Represent left- and right-hand circular polarization images, respectively;

[0013] Step 2, obtain the mapping of the linear polarization degree image DoLP and the linear polarization angle image AoLP to form a new polarization characteristic image, and record the mapping of the linear polarization degree image DoLP and the linear polarization angle image AoLP as L1 and L2 respectively:

[0014] L 1 =DoLP·cos(2AoLP) Formula (4)

[0015] L 2 =DoLP·sin(2AoLP) formula (5);

[0016] Step 3: Use the local histogram equalization algorithm to equalize the intensity image S 0 Perform local equalization processing to adjust the pixel value distribution;

[0017] Step 4: The light intensity image S processed in step 3 is 0 , the polarization characteristic image L obtained by formula (4) and formula (5) 1 and L 2 , is sent to the pre-trained image fusion network based on the autoencoder, the image fusion network includes an encoder, a fusion layer, and a decoder; the intensity image S processed in step 3 0 , L 1 and L 2 The high-dimensional features obtained after feature extraction by the encoder are weighted accordingly in the fusion layer, and then reconstructed by the decoder to generate a fused image.

[0018] Further, the step 1 specifically includes:

[0019] Step 11, using polarization imaging technology to collect polarization images or near-infrared polarization images of the scene; the image is captured based on a focal plane polarization sensor, and its key component is a focal plane array, in which each four micro-polarizers collect polarized light in directions of 0°, 45°, 90°, and 135°, respectively, to form a 2×2 super pixel. These periodic super pixels are arranged on the focal plane, so that the focal plane polarization sensor can simultaneously obtain polarization images in four directions;

[0020] Step 12: Decode the polarization image obtained in step 11 into the outgoing light intensity polarization images I in four polarization directions. 0 ,I 45 ,I 90 ,I 135 In addition, I L and I R Represent left- and right-hand circular polarization images, respectively;

[0021] Step 13, the four polarization image components I obtained in step 12 0 ,I 45 ,I 90 ,I 135 Perform Stokes polarization state solution to obtain the four Stokes parameters S of the polarization image 0 , S 1 , S 2 , S 3 ;

[0022] Step 14: Determine the degree of linear polarization image DoLP and the angle of linear polarization image AoLP according to the Stokes vector of the polarization image.

[0023] Furthermore, the decoding scheme in step 12 adopts Newton interpolation method, bilinear interpolation method or bicubic interpolation method.

[0024] Furthermore, in step 4, the training of the image fusion network based on the autoencoder is implemented on the MS-COCO general dataset. The training data in the training dataset is enhanced by cropping, rotating, and flipping data to generate high-dimensional features of different targets in the encoder respectively, and the reconstructed image is obtained by decoding. The training is continuously adjusted by the loss function until the image generated by the generator is within the allowable difference range with the real image.

[0025] The network design is as follows:

[0026] The encoder, i.e. the feature extraction stage, uses a standard convolutional layer and a cascaded densely connected convolutional block combination to extract rough features and deep features respectively, and obtain the multi-dimensional feature map of the source image of the training set; the cascaded densely connected convolutional block consists of three convolutional layers, and the output of each layer is used as the input of all subsequent layers; all convolutional layers are composed of convolution kernels, batch normalization layers and ReLU activation layers; the size of all convolution kernels in the network is the same, all 3X3, and the dimension of each convolutional layer of the autoencoder is 16, among which the last convolutional layer does not contain a ReLU layer;

[0027] Fusion layer: The mechanism of fusion layer is pixel weighting and does not participate in network training;

[0028] The decoder, i.e. the image reconstruction stage, recovers the final fusion result from the high-level features. The decoder consists of 5 convolutional layers, with dimensions of 64, 32, 16, 8, and 1 respectively.

[0029] The loss function is very important for feature extraction and image reconstruction, and is designed as:

[0030]

[0031] By minimizing Achieve image fusion performance, where weights are determined through experimental attempts

[0032] is 1000, θ is the training parameter in the neural network, C refers to the training dataset MS-COCO, and the loss function consists of two parts: the structural similarity loss function L MS-SSIM And the gradient loss function L G :

[0033]

[0034]

[0035] Among them, MS-SSIM is often used as a full-reference image quality evaluation indicator to drive I f The fused image is continuously sThe input image is close; are the gradients of the fusion result and the source image, respectively. is the Laplace operator, ||·|| represents the Frobenius norm, and H and W are the length and width of the image, respectively.

[0036] Furthermore, the training dataset is implemented on the MS-COCO general dataset, where millions of images are involved in the training, and the images involved in the training include both general images and polarization images.

[0037] Furthermore, the polarization state images acquired by the polarization imaging technology include images from indoor, outdoor, and near-infrared scenes.

[0038] The full linear polarization image fusion method based on autoencoder has the following beneficial effects:

[0039] (1) Compared with the traditional method of mapping images into pseudo-color space, the full linear polarization image fusion method based on the autoencoder of the present invention is more suitable for interpreting the physical properties of the scene and is convenient for collaborating with hardware facilities to realize target detection in the fields of industrial detection, medical diagnosis detection, atmospheric remote sensing detection, underwater imaging, etc.

[0040] (2) Compared with the method of fusing only the intensity map and the linear polarization degree image, the method of fusion of the full linear polarization image based on the autoencoder of the present invention introduces the effective extraction and fusion of the polarization degree and polarization angle image, and suppresses the noise from the polarization state. Even in outdoor scenes where the lighting is uncontrollable and the materials are different, most of the scene information can be reconstructed, effectively improving the target detection accuracy in complex scenes.

[0041] (3) Compared with the conventional multi-scale decomposition polarization image fusion method, the full-linear polarization image fusion method based on the autoencoder of the present invention can reflect the shadow and edge information in the image and will not introduce the pseudo-Gibbs effect, which is more suitable for polarization image fusion.

[0042] (4) The present invention uses a full-line polarization image fusion method based on an autoencoder and uses the general dataset MS-COCO for training, thus avoiding the problem of a small number of polarization image datasets and a lack of real images in the field of polarization fusion.

[0043] (5) Compared with the conventional polarization image fusion method, the full-linear polarization image fusion method based on the autoencoder of the present invention increases the target enhancement of the light intensity map, further increases the target details, improves the contrast, and is more suitable for the observation of non-related researchers and the recognition of scenes.

[0044] (6) The present invention is based on the full linear polarization image fusion method of the autoencoder, which is aimed at the feature extraction and reconstruction of the convolutional neural network. The network structure is simple in design, the size of the convolution kernel is 3×3, and the maximum number of channels does not exceed 64 dimensions. Although it is based on the fusion of three source images, it achieves a good balance between the fusion effect and the computational efficiency.

[0045] (7) The full-linear polarization image fusion method based on the autoencoder of the present invention is suitable for the fusion of multiple images due to its operation of extracting features by target, and can be extended to multi-exposure image fusion, multi-focus image fusion, visible light and infrared image fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 : Flow chart of the full linear polarization image fusion method based on the autoencoder in the embodiment of the present invention;

[0047] Figure 2 : Schematic diagram of the network model of the autoencoder in the embodiment of the present invention;

[0048] Figure 3 : Effect comparison diagram of the present invention and the prior art. DETAILED DESCRIPTION

[0049] The present invention will be further described below in conjunction with the accompanying drawings:

[0050] Figures 1 to 3 A specific implementation of a full linear polarization image fusion method based on an autoencoder of the present invention is shown. Figure 1 is a flow chart of the full linear polarization image fusion method based on the autoencoder in this embodiment; Figure 2 is a schematic diagram of the network model of the autoencoder in this embodiment; Figure 3 It is a comparison diagram of the effects of the present invention and the prior art in this embodiment.

[0051] like Figure 1 As shown, the full linear polarization image fusion method based on the autoencoder in this embodiment includes the following steps:

[0052] Step 1: Perform Stokes solution on the polarization image obtained by polarization imaging technology to obtain four Stokes parameters S 0 , S 1 , S 2 , S 3 , the formula is as follows:

[0053]

[0054] And determine the linear polarization degree image DoLP and the linear polarization angle image AoLP according to the Stokes vector of the polarization image:

[0055]

[0056]

[0057] Among them, S 0 represents the intensity image, S 1 Represents the linear or vertical polarization component image, S 2 Represents the linear polarization component image at 45 degrees or 135 degrees, S 3 Represents the right-handed or left-handed polarization component image in the light beam. In the natural environment, the circular polarization component is very small and can be ignored; I 0 ,I 45 , I 90 ,I 135 Respectively represent the polarization images of the outgoing light intensity at four polarization directions of 0°, 45°, 90°, and 135°, I L and I R Represent left- and right-hand circular polarization images, respectively;

[0058] Step 2, obtain the mapping of the linear polarization degree image DoLP and the linear polarization angle image AoLP to form a new polarization characteristic image, and record the mapping of the linear polarization degree image DoLP and the linear polarization angle image AoLP as L1 and L2 respectively:

[0059] L 1 =DoLP·cos(2AoLP) Formula (4)

[0060] L 2 =DoLP·sin(2AoLP) formula (5);

[0061] The linear polarization degree DoLP and linear polarization angle image AoLP obtained according to formula (2) and formula (3) can reflect information such as illumination, roughness, edge, stress distribution, birefringence, surface orientation, etc. compared with the intensity map. DoLP has high overall contrast, low brightness, and fewer details; while AoLP has high brightness and more details, especially showing greater advantages under low light conditions, but at the same time AoLP amplifies the noise and is overexposed. Therefore, the present invention proposes the above-mentioned mapping of the linear polarization degree and the linear polarization angle;

[0062] Step 3: Use the local histogram equalization algorithm to equalize the intensity image S 0 Perform local equalization processing to adjust the pixel value distribution;

[0063] Step 4: The light intensity image S processed in step 3 is 0 , the polarization characteristic image L obtained by formula (4) and formula (5) 1 and L 2, is sent to the pre-trained image fusion network based on the autoencoder, the image fusion network includes an encoder, a fusion layer, and a decoder; the intensity image S processed in step 3 0 , L 1 and L 2 The high-dimensional features obtained after feature extraction by the encoder are weighted accordingly in the fusion layer, and then reconstructed by the decoder to generate a fused image.

[0064] Preferably, the step 1 specifically includes:

[0065] Step 11, using polarization imaging technology to collect a polarization image or near-infrared polarization image of the scene; the image is captured based on a focal plane polarization sensor, and its key component is a focal plane array, in which each four micropolarizers collect polarized light in directions of 0°, 45°, 90°, and 135°, respectively, to form a 2×2 super pixel. These periodic super pixels are arranged on the focal plane, so that the focal plane polarization sensor can simultaneously obtain polarization images in four directions;

[0066] Step 12: Decode the polarization image obtained in step 11 into polarization images I of the outgoing light intensity in four polarization directions. 0 ,I 45 ,I 90 ,I 135 In addition, I L and I R Represent left- and right-hand circular polarization images, respectively;

[0067] Step 13, the four polarization image components I obtained in step 12 0 ,I 45 ,I 90 ,I 135 Perform Stokes polarization state solution to obtain the four Stokes parameters S of the polarization image 0 , S 1 , S 2 , S 3 ;

[0068] Step 14: determining the degree of linear polarization image DoLP and the angle of linear polarization image AoLP according to the Stokes vector of the polarization image.

[0069] Preferably, the decoding scheme in step 12 adopts Newton interpolation method, bilinear interpolation method or bicubic interpolation method. In this embodiment, Newton interpolation method is adopted.

[0070] Preferably, in step 4, the training of the image fusion network based on the autoencoder is implemented on the MS-COCO general dataset, and the training data in the training data set is enhanced by cropping, rotating, flipping, etc. to generate high-dimensional features of different targets in the encoder respectively, and the reconstructed image is obtained by decoding, and the training is continuously adjusted by the loss function constraint until the image generated by the generator is almost the same as the real image (that is, the image generated by the generator is within the allowable difference range with the real image);

[0071] The network design is as follows:

[0072] The encoder, i.e. the feature extraction stage, uses a standard convolutional layer and a cascaded densely connected convolutional block combination to extract rough features and deep features respectively, and obtain the multi-dimensional feature map of the source image of the training set; the cascaded densely connected convolutional block consists of three convolutional layers, and the output of each layer is used as the input of all subsequent layers; all convolutional layers are composed of convolution kernels, batch normalization layers and ReLU activation layers; the size of all convolution kernels in the network is the same, all 3X3, and the dimension of each convolutional layer of the autoencoder is 16, among which the last convolutional layer does not contain a ReLU layer;

[0073] Fusion layer: The mechanism of fusion layer is pixel weighting and does not participate in network training;

[0074] The decoder, i.e. the image reconstruction stage, recovers the final fusion result from the high-level features. The decoder consists of 5 convolutional layers, with dimensions of 64, 32, 16, 8, and 1 respectively.

[0075] The loss function is very important for feature extraction and image reconstruction, and is designed as:

[0076]

[0077] By minimizing Achieve image fusion performance, where weights are determined through experimental attempts

[0078] is 1000, θ is the training parameter in the neural network, C refers to the training dataset MS-COCO, and the loss function consists of two parts: the structural similarity loss function L MS -SSI and gradient loss function L G :

[0079]

[0080]

[0081] Among them, MS-SSIM is often used as a full-reference image quality evaluation indicator to drive If The fused image is continuously s The input image is close; are the gradients of the fusion result and the source image, respectively. is the Laplace operator, ||·|| represents the Frobenius norm, and H and W are the length and width of the image, respectively. The autoencoder based on deep learning has a fast processing speed on the one hand and a strong feature representation capability on the other hand. Although the present invention realizes the fusion of three images, the processing speed is still very fast and the noise of the original polarization state is avoided, achieving a better fusion effect.

[0082] Preferably, the training data set is implemented on the MS-COCO general data set, in which millions of images are involved in the training, and the images involved in the training include both general images and polarization images.

[0083] Preferably, the polarization state images obtained by the polarization imaging technology include images from indoor, outdoor, and near-infrared scenes. Among them, the linear polarization degree and linear polarization angle noise of outdoor target images are very large, and the polarization blur is most significant. The mapping L proposed in the present invention 1 , L 2 This problem is solved to a large extent, and the high-frequency information of the linear polarization degree and linear polarization angle is retained.

[0084] In this embodiment, Python language is used in conjunction with the TensorFlow deep learning framework to perform full linear polarization image fusion based on autoencoders. According to the pre-constructed loss function formulas (6), (7), and (8), the MS-COCO dataset trains an image fusion network model based on autoencoders, and saves the model and its parameters.

[0085] In this embodiment, Figure 3 As shown, the polarization image or near-infrared polarization image of the scene is collected by using polarization imaging technology; then, the polarization image is decoded into the outgoing light intensity polarization image I in four polarization directions. 0 ,I 45 ,I 90 ,I 135 , and the four polarization image components I 0 ,I 45 , I 90 ,I 135 Perform Stokes polarization state solution to obtain the four Stokes parameters S of the polarization image 0 , S 1 , S 2 , S 3 ,like Figure 3 (a); Then, the linear polarization degree image DoLP and the linear polarization angle image AoLP are determined according to the Stokes vector of the target image, as Figure 3(b) and (c) in FIG. 4 ; then, according to formula (4) and formula (5), the mapping of the linear polarization degree image DoLP and the linear polarization angle image AoLP is obtained to form new polarization characteristic images L1 and L2, as shown in FIG. Figure 3 (d) and (e) in Fig. 2; then, the local histogram equalization algorithm is used to equalize the intensity image S 0 Perform local equalization processing to adjust the pixel value distribution; then, the intensity image S after equalization processing 0 , polarization characteristic image L 1 and L 2 Send it to the pre-trained autoencoder-based image fusion network model, such as Figure 2 As shown in the figure, the pre-trained image fusion network model based on the autoencoder is used to transform the equalized light intensity image S 0 , polarization characteristic diagram L 1 , L 2 The image is sent to the autoencoder to extract the high-dimensional features of each image. After the encoder extracts the features, there are 3×64-dimensional feature maps. The 64-dimensional features from the three source images are weighted accordingly in the fusion layer to obtain a 1×64-dimensional feature map, which is then sent to the decoder. Each convolutional layer gradually restores and reconstructs the final fusion map, as shown in the figure. Figure 3 In (j). Figure 3 (f)-(i) are the polarization fusion results obtained by FEVIP, LPF, PFNet, and DeepFuse using (a) and (b), respectively, and the present invention has better effect. The present invention obtains the mapping of linear polarization degree and linear polarization angle to form a new polarization feature image, and combines it with the light intensity image to obtain a set of new data paradigms, and then uses the autoencoder based on convolutional neural network to extract features, fuse features, and reconstruct images from the light intensity information and polarization feature image of the target. This polarization data mapping can minimize the polarization blur caused by material properties and lighting environment, and is robust to various scenes. The image fusion method of the autoencoder can enable the features to be fully extracted and fused, retaining and enhancing the polarization information of the target, and improving the target detection capability under complex backgrounds.

[0086] The present invention has the following beneficial effects:

[0087] Compared with the traditional method of mapping images into pseudo-color space, the full linear polarization image fusion method based on the autoencoder is more suitable for interpreting the physical properties of the scene and is convenient for collaborating with hardware facilities to realize target detection in the fields of industrial detection, medical diagnosis detection, atmospheric remote sensing detection, underwater imaging, etc.

[0088] Compared with the fusion method of the light intensity map and the linear polarization degree image, the full linear polarization image fusion method based on the autoencoder in the present invention introduces the effective extraction and fusion of the polarization degree and polarization angle images, and suppresses the noise from the polarization state. Even in the outdoor scene where the lighting is uncontrollable and the material is different, most of the scene information can be reconstructed, effectively improving the target detection accuracy in complex scenes.

[0089] Compared with the conventional multi-scale decomposition polarization image fusion method, the full-line polarization image fusion method based on the autoencoder in the present invention can reflect the shadow and edge information in the image without introducing the pseudo-Gibbs effect, and is more suitable for polarization image fusion.

[0090] The present invention is based on the full-line polarization image fusion method of the autoencoder, and uses the general data set MS-COCO for training, which avoids the problem of small polarization image data sets and no real images in the field of polarization fusion.

[0091] Compared with the conventional polarization image fusion method, the full-linear polarization image fusion method based on the autoencoder in the present invention increases the target enhancement of the light intensity map, further increases the target details, improves the contrast, and is more suitable for the observation of non-related researchers and the recognition of scenes.

[0092] The present invention is based on the full linear polarization image fusion method of the autoencoder, which is aimed at the feature extraction and reconstruction of the convolutional neural network. The network structure is simple in design, the size of the convolution kernel is 3×3, and the maximum number of channels does not exceed 64 dimensions. Although it is based on the fusion of three source images, it achieves a good balance between the fusion effect and the computational efficiency.

[0093] The full-linear polarization image fusion method of the present invention is based on the autoencoder. Due to its operation of extracting features by target, the algorithm is suitable for the fusion of multiple images and can be expanded to multi-exposure image fusion, multi-focus image fusion, visible light and infrared image fusion.

[0094] In summary: Based on the powerful feature extraction capability of deep learning, the present invention designs an image fusion network based on autoencoders, constrains the performance of image fusion by a loss function, and designs a new fusion of polarization feature images and light intensity images based on polarization imaging, which increases the details of the fused image, uses polarization features to supplement and enhance the scene's shadows, roughness, edges and other attributes, and suppresses polarization blur caused by illumination and material properties. The present invention is suitable for polarization image fusion and target detection systems in complex scenes.

[0095] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the implementation of the present invention is not limited to the above-mentioned method. As long as various improvements are made using the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.

Claims

1. A full linear polarization image fusion method based on autoencoder, It is characterized in that The following steps are involved: Step 1: Perform Stokes solution on the polarization image obtained by polarization imaging technology to obtain four Stokes parameters S 0 , S 1 , S 2 , S 3 , the formula is as follows: (Formula 1) And determine the linear polarization degree image DoLP and the linear polarization angle image AoLP according to the Stokes vector of the polarization image: (Formula 2) (Formula 3) Among them, S 0 represents the intensity image, S 1 Represents the linear or vertical polarization component image, S 2 Represents the linear polarization component image at 45 degrees or 135 degrees, S 3 Represents the right-handed or left-handed polarization component image in the light beam. In the natural environment, the circular polarization component is very small and can be ignored; I 0 ,I 45 ,I 90 ,I 135 Respectively represent the polarization images of the outgoing light intensity at four polarization directions of 0°, 45°, 90°, and 135°, I L and I R Represent left- and right-hand circular polarization images, respectively; Step 2: Obtain the mapping of the linear polarization degree image DoLP and the linear polarization angle image AoLP to form a new polarization characteristic image. The mapping of the linear polarization degree image DoLP and the linear polarization angle image AoLP are respectively recorded as L 1 , L 2 : L1=DoLP·cos(2AoLP) Formula (4) L2=DoLP·sin(2AoLP) formula (5); Step 3: Use the local histogram equalization algorithm to equalize the intensity image S 0 Perform local equalization processing to adjust the pixel value distribution; Step 4: The light intensity image S processed in step 3 is 0 , the polarization characteristic image L obtained by formula (4) and formula (5) 1 and L 2 , is sent to the pre-trained image fusion network based on the autoencoder, the image fusion network includes an encoder, a fusion layer, and a decoder; the intensity image S processed in step 3 0 , L 1 and L 2 The high-dimensional features obtained after feature extraction by the encoder are weighted accordingly in the fusion layer, and then reconstructed by the decoder to generate a fused image; The specific structure of the image fusion network is as follows: The encoder, i.e. the feature extraction stage, uses a standard convolutional layer and a cascaded densely connected convolutional block combination to extract rough features and deep features respectively, and obtain the multi-dimensional feature map of the source image of the training set; the cascaded densely connected convolutional block consists of three convolutional layers, and the output of each layer is used as the input of all subsequent layers; all convolutional layers are composed of convolution kernels, batch normalization layers and ReLU activation layers; the size of all convolution kernels in the network is the same, all 3X3, and the dimension of each convolutional layer of the autoencoder is 16, among which the last convolutional layer does not contain a ReLU layer; Fusion layer: The mechanism of fusion layer is pixel weighting and does not participate in network training; The decoder, i.e. the image reconstruction stage, recovers the final fusion result from the high-level features; the decoder consists of 5 convolutional layers, and the dimensions of each convolutional layer are 64, 32, 16, 8, and 1 respectively.

2. According to the full linear polarization image fusion method based on autoencoder according to claim 1, It is characterized in that The step 1 specifically includes: Step 11, using polarization imaging technology to collect a polarization image or a near-infrared polarization image of the scene; Step 12: Decode the polarization image obtained in step 11 into the outgoing light intensity polarization images I in four polarization directions. 0 ,I 45 ,I 90 ,I 135 In addition, I L and I R Represent left- and right-hand circular polarization images, respectively; Step 13, the four polarization image components I obtained in step 12 0 ,I 45 ,I 90 ,I 135 Perform Stokes polarization state solution to obtain the four Stokes parameters S of the polarization image 0 , S 1 , S 2 , S 3 ; Step 14: Determine the degree of linear polarization image DoLP and the angle of linear polarization image AoLP according to the Stokes vector of the polarization image.

3. The full linear polarization image fusion method based on autoencoder according to claim 2, It is characterized in that In step 11, the polarization image is captured based on a focal plane polarization sensor, a key component of which is a focal plane array, in which every four micropolarizers collect polarized light in directions of 0°, 45°, 90° and 135° respectively to form a 2×2 super pixel. These periodic super pixels are arranged on the focal plane, so that the focal plane polarization sensor can obtain polarization images in four directions simultaneously.

4. The full linear polarization image fusion method based on autoencoder according to claim 2, It is characterized in that The decoding scheme in step 12 adopts Newton interpolation method, bilinear interpolation method or bicubic interpolation method.

5. The full linear polarization image fusion method based on autoencoder according to claim 1, It is characterized in that In step 4, the training of the image fusion network based on the autoencoder is implemented on the MS-COCO general dataset. The training data in the training data set is enhanced by cropping, rotating, and flipping the data to generate high-dimensional features of different targets in the encoder respectively, and the reconstructed image is obtained by decoding. The training is continuously adjusted by the loss function until the image generated by the generator is within the allowable difference range with the real image. The loss function is very important for feature extraction and image reconstruction, and is designed as: (Formula 6) By minimizing The performance of image fusion is achieved. The weight α is determined to be 1000 through experimental attempts. θ is the training parameter in the neural network. C refers to the training dataset MS-COCO. The loss function consists of two parts: the structural similarity loss function L MS-SSIM And the gradient loss function L G : (Formula 7) (Formula 8) Among them, MS-SSIM is often used as a full-reference image quality evaluation indicator to drive I f The fused image is continuously s The input image is close; , are the gradients of the fusion result and the source image, respectively. is the Laplace operator. ||·|| represents the Frobenius norm. H and W are the length and width of the image, respectively.

6. The full linear polarization image fusion method based on autoencoder according to claim 5, It is characterized in that The training dataset is implemented on the MS-COCO general dataset, where millions of images are involved in the training, including both general images and polarization images.

7. The full linear polarization image fusion method based on autoencoder according to claim 1, It is characterized in that The polarization state images acquired by polarization imaging technology include images from indoor, outdoor, and near-infrared scenes.