Polarization spectrum reconstruction method based on color polarization array
By installing a color polarization array in front of the RGB color camera and combining a deep learning model, the complex and non-real-time problem of spectral polarization image acquisition in the prior art is solved, and low-cost, real-time reconstruction of spectral and polarization information is achieved.
Patent Information
- Application Number
- CN202510401842.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art requires complex optical structures and computational reconstruction algorithms when acquiring four-dimensional spectral polarization images, and lacks real-time performance and cannot dynamically capture spectral and polarization information.
Using an RGB color camera combined with a color polarization array, the mosaic is removed by interpolation and the end-to-end pre-trained deep learning model, especially the convolutional neural network Transformer combined with the U-shaped architecture, to achieve efficient reconstruction of polarization spectroscopy.
It realizes low-cost and real-time acquisition of spectral and polarization information, and the reconstruction effect is better than traditional methods, suitable for dynamic scenarios, reducing the computational complexity and parameter quantity.
Smart Images

Figure CN120339509A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of polarization imaging and spectral imaging, and particularly to an efficient polarization spectrum reconstruction method based on a color polarization array. Background Art
[0002] In the field of computer vision and imaging technology, the wavelength and polarization state of light are inherent properties of light, carrying hidden properties of the target scene. Although this goes beyond the scope of traditional intensity-based imaging applications, capturing and analyzing these characteristics is still crucial for understanding the physical properties of complex scenes. Four-dimensional spectral polarization images can provide rich information about the scene, including 2D spatial information, 1D polarization, and 1D spectral distribution, which is very valuable for many applications such as remote sensing, medical imaging, and astronomy.
[0003] Although there are now some methods for obtaining four-dimensional spectral polarization, their limitations are obvious, often relying on complex optical structure settings and computational reconstruction algorithms. In addition, another major defect of these methods is the lack of real-time performance, unable to capture spectral and polarization information from dynamic scenes.
[0004] Recently, due to the rise of deep learning methods, many advancements have been made in the field of computer vision. In particular, the Transformer model, which initially revolutionized natural language processing (NLP), but with recent explorations of the application of Transformer in computer vision, researchers have found that the Transformer model outperforms traditional methods in many tasks. The multi-head self-attention module in Transformers is good at capturing non-local similarities and long-term dependencies. In the task of reconstructing polarization or spectral images, the Transformer-based model shows excellent performance. However, the traditional Transformer model with a self-attention mechanism cannot effectively utilize the channel correlation in high-dimensional data. In addition, the network parameter quantity of a pure Transformer-based model is often extremely large, thus bringing great difficulties to the training of the model. Summary of the Invention
[0005] Aiming at the defects of the above prior art, the purpose of the present invention is to provide an efficient polarization spectrum reconstruction method based on a color polarization array.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A polarization spectrum reconstruction method based on a color polarization array, the method comprising the following steps:
[0008] S1. Obtain the unprocessed low-resolution mosaic image I0 using an RGB color camera. Among them, a color polarization array is placed in front of the image sensor of the RGB color camera;
[0009] S2. Perform interpolation to remove the mosaic on the unprocessed low-resolution mosaic image I0 to obtain the preliminarily processed image I1;
[0010] S3. Input the image I1 into an end-to-end pre-trained deep learning model for polarization spectrum reconstruction to reconstruct a four-dimensional spectral polarization image.
[0011] Furthermore, in step S1, a color polarization array is placed in front of the image sensor of the RGB color camera. Specifically, the color polarization array is closely attached to the microlens array of the lens module in front of the image sensor.
[0012] Furthermore, in step S2, performing interpolation to remove the mosaic on the unprocessed low-resolution mosaic image I0 specifically includes: aggregating pixel points with the same color and the same polarization angle, and using the spatial similarity of the picture and the redundancy of the channel dimension to first interpolate and reshape to the original resolution; then cascading two-dimensional images of different channels into a normal image, and performing spatial interpolation again to further fuse image features.
[0013] Furthermore, in step S3, the end-to-end pre-trained deep learning model adopts a convolutional neural network Transformer combined with a U-shaped architecture, including a three-dimensional convolutional module and a polarization spectrum attention block based on the Transformer model.
[0014] Furthermore, in step S3, the three-dimensional convolutional module serves as an encoder to layer by layer extract spatial-spectral-polarization joint features.
[0015] Furthermore, in step S3, the polarization spectrum attention block includes a polarization spectrum attention module and a feed-forward network module; the polarization spectrum attention module is used for feature representation of the image; the feed-forward network module is used to predict the final reconstructed image from the extracted features.
[0016] Furthermore, in step S3, the polarization spectrum attention module first performs patch embedding processing on the input image, cuts the entire image into local patches, converts each image patch into a feature vector through linear projection to form a Token sequence, and then the feed-forward network module performs position encoding, Query-Key-Value mapping, and attention weighting calculation on the Token sequence.
[0017] Furthermore, in step S3, the specific processing steps of the polarization spectrum attention module are:
[0018] S31. First, extract the feature representation of the image I1 through a three-dimensional convolutional block;
[0019] S32. For the first polarization channel among the four polarization channels, that is, the polarization channel with the polarization direction at an angle of 0 degrees to the horizontal plane, regard the characteristics of all spectral channels at each spatial position as a group of sequence Tokens. On this basis, set the first polarization channel as the target channel for attention calculation, acting as the value Value, and generate the query Query and the key Key respectively for the other three polarization channels through a group of shared linear transformations;
[0020] S33. According to the dot product mechanism, calculate the dot product similarity between the query Query and the key Key, and obtain the attention weight matrix through normalization by an activation function;
[0021] S34. Apply the attention weight matrix to the target channel to achieve the directed fusion and guided reconstruction of polarization-spectral joint information.
[0022] The present invention uses an easily obtainable RGB color camera and a color polarization array with low cost and easy mass production to collect the RGB polarization low-resolution mosaic measurement values obtained from hyperspectral projection, and then realizes the efficient reconstruction of high-dimensional polarization spectral data through a convolutional neural network Transformer combined with a U-shaped architecture end-to-end deep learning network model. On the one hand, compared with traditional spectral acquisition methods, the present invention can not only achieve low cost and real-time acquisition of spectral information, but also further reconstruct polarization information on the basis of spectral reconstruction, which has important value for downstream applications; on the other hand, the present invention also designs an end-to-end deep learning network model with a low number of parameters and high efficiency, namely a convolutional neural network Transformer combined with a U-shaped architecture, for the high-dimensional optical reconstruction task of polarization spectral reconstruction, and particularly designs a polarization spectral attention block, and the reconstruction effect is far better than previous methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flow block diagram of the method of the present invention.
[0024] Figure 2 It is a schematic diagram of the color polarization array in the embodiment of the present invention.
[0025] Figure 3 It is a schematic diagram of the end-to-end network model in the embodiment of the present invention.
[0026] Figure 4 It is a schematic diagram of the structure of the polarization spectral attention block in the embodiment of the present invention.
[0027] Figure 5 It is a schematic diagram of the reconstructed polarization spectral image in the embodiment of the present invention. It shows the reconstruction results of three wavelength channels of the reconstructed polarization spectral image, as well as the polarization degree and polarization angle of the reconstructed image.
[0028] Figure 6 Schematic diagram of the structure of the imaging system in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] like Figure 1 As shown, this embodiment provides a polarization spectrum reconstruction method based on a color polarization array, comprising the following steps:
[0031] S1, using an RGB color camera to obtain an unprocessed low-resolution mosaic image I0, wherein a color polarization array is placed in front of a CMOS sensor of the RGB color camera;
[0032] S2, performing interpolation preprocessing on the unprocessed low-resolution mosaic image I0 to remove the mosaic, and obtaining a preliminarily processed image I1;
[0033] S3, inputting the image I1 into the end-to-end pre-trained deep learning model for polarization spectrum reconstruction to reconstruct a four-dimensional spectral polarization image.
[0034] The specific implementation process of the above method is shown in the following example.
[0035] 1. Image Acquisition
[0036] Hyperspectral images play a significant role in remote sensing, medical diagnosis, target detection and other fields because they have richer narrowband spectral information than ordinary RGB images. However, traditional spectral acquisition methods not only mostly use push-broom, which is time-consuming and cannot capture dynamic scenes, but also often rely on precision optical systems and complex algorithms, and are very expensive. RGB cameras are relatively easy to obtain. RGB images are images with broadband color information obtained by narrowband spectral projection, so they have the potential to restore hyperspectral images. The present invention places a color polarization array in front of the CMOS sensor of the RGB color camera to obtain polarization-modulated RGB polarization low-resolution image measurement values for reconstruction of polarization spectrum 4-dimensional information. This method has the advantages of miniaturization and low cost, and the color polarization array has a mass production technical route. Figure 2 The schematic diagram of the structure of the color polarization array (CPFA) is shown, which is used to simultaneously collect the color information and polarization information of the image. The 4×4 array in the figure is a minimum unit of the color polarization array, which consists of 16 pixels. The stripes in different directions of each pixel represent its corresponding polarization channel (that is, the angles between the polarization direction and the horizontal plane are 0°, 45°, 90°, and 135° respectively), and different colors mark the color channel (R, G, B) to which it belongs.
[0037] In this embodiment, first, the color polarization array is closely attached to the lens module ([ Figure 5 ) on the microlens array in front of the image sensor to ensure that each pixel receives information of a specific polarization angle and color band, thereby achieving the purpose of simultaneously encoding the polarization information and spectral information of the image. This solution overcomes the defects of complexity, inability to dynamically obtain, and high cost in encoding polarization and spectral information in the prior art.
[0038] Then, a configured RGB color camera is used to obtain a low-resolution mosaic image I0 that has not undergone any ISP (Image Signal Processor) processing. This image is the raw data directly read from the sensor and contains color and polarization joint encoding information. Different from traditional image acquisition methods, the present invention bypasses conventional processing procedures such as white balance, demosaicing, and gamma correction, and retains the complete pixel-level physical response data, so that each sampling point still maintains a one-to-one correspondence with the original color polarization array.
[0039] 2. Interpolation for removing mosaic
[0040] After the unprocessed low-resolution mosaic image I0 is acquired, the present invention performs an interpolation operation on I0 to remove the mosaic. The interpolation method can adopt bilinear interpolation or bicubic interpolation. After the initial interpolation for demosaicing, the checkerboard effect will be reduced, and it is easier to learn global semantic information during network training, and the reconstructed result is also smoother.
[0041] In the color polarization array structure proposed by the present invention, the information received by each pixel not only includes the differences in color channels (R / G / B), but also includes the information encoding of a specific polarization angle. Therefore, when performing image restoration, the interpolation method must maintain physical consistency and optical characteristic matching in both the color and polarization dimensions. In addition, the color polarization image has the characteristics of sparse sampling and non-uniform distribution during the acquisition stage. Based on this, in this embodiment, for the jointly encoded image of color and polarization information, a joint-guided interpolation strategy is designed, which combines color space similarity and polarization angle directionality information, significantly improves the interpolation accuracy and image restoration quality, and has strong innovation. The specific operation is as follows: Aggregate pixel points with the same color and the same polarization angle, and use the spatial similarity of the picture and the redundancy of the channel dimension to first interpolate and reshape to the original resolution; then cascade the two-dimensional images of different channels into a normal image, and perform another spatial interpolation to further fuse the image features and improve the image quality, which is helpful for subsequent reconstruction steps.
[0042] 3. End-to-end network model
[0043] With the development of deep learning in the field of images, the research on hyperspectral image (HSI) reconstruction tasks has also shifted from traditional image processing methods to deep learning methods. Many works have focused on carefully designing various spectral-spatial networks, among which the convolutional neural network (CNN) is one of the most popular architectures.
[0044] With the development of the natural language processing field (NLP), Transformer, which originated from natural language processing, has begun to expand to the computer vision field, and Vision Transformer is one of the typical representatives. Compared with CNN, the Transformer-based model can better capture long-range dependencies. In the task of spectral reconstruction, it can also better utilize the spectral self-similarity of hyperspectral images. However, the Transformer-based model usually has more parameters and higher computational complexity than CNN, which brings higher costs to the training of the network.
[0045] In this embodiment, for the high-dimensional reconstruction task of polarization spectral reconstruction, a convolutional neural network-Transformer joint U-shaped architecture is designed, specifically a U-shaped network structure based on three-dimensional convolutional encoding and Transformer decoding, which is used to achieve high-fidelity reconstruction of polarization-spectral joint images. This structure combines the local structure modeling ability of three-dimensional convolution with the global modeling ability of Transformer in the channel dimension for the first time in view of the characteristics that polarization-spectral images are highly heterogeneous and correlated in the spatial, spectral, and polarization dimensions, and constructs a deep reconstruction framework with multi-dimensional feature perception and global dependence modeling capabilities.
[0046] As Figure 3 shown, the entire network maintains a U-shaped structure. In the downsampling path, three-dimensional convolutional modules are used as encoders to extract spatial-spectral-polarization joint features layer by layer. Among them, the number of three-dimensional convolutional modules is determined according to specific needs and is set to four in this embodiment. Through continuous three-dimensional convolution, the fine-grained information of the input image in the spatial local, spectral neighborhood, and polarization angle is effectively aggregated, so as to obtain a compact representation of high-dimensional sparse observations. At the deepest layer of the encoder, the polarization spectral attention block further extracts deep features from the three-dimensional feature tensor of the encoder and inputs them into the Transformer decoder.
[0047] The decoder consists of polarization spectral attention blocks constructed by multi-layer Transformer models. Each polarization spectral attention block includes a polarization spectral attention module and a feed-forward network module, as Figure 4As shown, the polarization spectrum attention module is used for feature characterization of the image, and the feed-forward network module is used to predict the final reconstructed image from the extracted features. The polarization spectrum attention module first processes the input image through block embedding, cuts the whole image into local blocks, converts each image block into a feature vector through linear projection to form a Token sequence, and then the feed-forward network module performs position encoding, Query-Key-Value mapping and attention weighting calculation on the Token sequence to realize the modeling of the long-range dependence relationship between different polarization channels and spectral channels. Further, a residual connection can be set to improve the stability of training. During the decoding process, the attention mechanism is used to dynamically allocate the information fusion weights between channels, guide the flow and reconstruction of information in the high-dimensional space, and gradually complete the restoration of the image spatial structure. Since the Transformer module has a larger receptive field and the selective expression ability between channels, the semantic dependence relationship between different polarization angles and different bands can be fully mobilized during the decoding stage, improving the global consistency of the final image and the spectral-polarization joint reconstruction quality.
[0048] Meanwhile, the network adopts a skip connection mechanism to directly introduce the local features of each layer of the encoder into the corresponding decoding stage, enhancing the retention of low-level structure information and effectively alleviating the problem of information degradation in reconstruction.
[0049] The proposed three-dimensional convolutional encoding and Transformer decoding structure of the present invention combines the advantages of both the Transformer model and CNN. It not only uses three-dimensional convolution as the encoder of the network to aggregate and extract polarization spectrum depth information at a lower computational cost, but also uses a polarization spectrum attention block based on the Transformer model as the decoder to capture long-range dependencies and the high self-similarity between the polarization channels and spectral channels of the polarization spectrum image, greatly improving the reconstruction effect. Therefore, this model not only has a lower number of parameters, but also has a better reconstruction effect than previous models. The present invention breaks through the limitation of traditional U-Net that only processes in the spatial domain, realizes the joint modeling and full-dimensional restoration of spatial-spectral-polarization three-dimensional information, and is the first network structure that can simultaneously achieve polarization and spectral collaborative reconstruction, showing significant advantages in reconstruction accuracy, network generalization ability and physical consistency.
[0050] In this embodiment, to solve the problem of high-dimensional heterogeneous feature fusion in polarization spectrum reconstruction, a polarization spectrum attention module is proposed, aiming to achieve efficient modeling and fusion of polarization and spectral features through a guided attention mechanism. This module fully combines the differences and correlations in information expression between polarization angles and spectral channels to construct a joint attention calculation framework. Specifically, first, a three-dimensional convolutional block is used to extract the feature representation of the image after interpolation and demosaicing, and the joint attention calculation framework is regarded as a feature map with multiple polarization-spectral combination channels. Subsequently, for the first polarization channel in the four polarization channels of the polarization spectrum attention module (i.e., the polarization channel with a polarization direction angle of 0 degrees with the horizontal plane), for each spatial position, the features of all spectral channels at this position are regarded as a group of sequence tokens. On this basis, a cross-channel attention structure is built, with the first polarization channel set as the target channel for attention calculation, acting as the value (Value), and the other three channels (i.e., the polarization channels with polarization direction angles of 45 degrees, 90 degrees, and 135 degrees with the horizontal plane) are respectively generated as queries (Query) and keys (Key) through a group of shared linear transformations. Subsequently, according to the dot product (ScaledDot-Product Attention) mechanism, the dot product similarity between the query and the key is calculated, and the attention weight matrix is obtained through normalization by the Softmax function, and then it is applied to the value (Value) channel to achieve the directed fusion and guided reconstruction of polarization-spectral joint information. Different from the traditional attention modeling in the spatial domain, this method significantly reduces the computational complexity from O(H*W) to O(H*C), and at the same time utilizes the high correlation between spectral channels in hyperspectral images to effectively improve the spectral feature extraction ability. The output of the polarization spectrum attention module contains polarization-spectral joint features that have been guided and adaptively weighted, which not only enhances the response of key channels but also suppresses invalid or interfering information, thus providing a more discriminative intermediate feature representation for subsequent image reconstruction.
[0051] 4. Pre-trained model
[0052] Compared with traditional image processing methods, deep learning methods have shown significant advantages in many aspects. They can automatically extract complex features at the original pixel level, implement the mapping from input to output in a data-driven manner, and do not require complex preprocessing and postprocessing steps. In this embodiment, the real-world polarization spectrum dataset open-sourced in Spectral and Polarization Vision: Spectro-polarimetric Real-world Dataset published in the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) is used as the benchmark dataset to train and test the end-to-end network model. Through the analysis of a large amount of data, the mapping from the unprocessed low-resolution mosaic image to the full-resolution polarization spectrum image is learned, and a pre-trained model is obtained for subsequent reconstruction.
[0053] The pre-trained model obtained in the present invention has good transferability and module independence, providing a basis for the lightweight deployment and marginal application of subsequent models. With the continuous improvement of model performance and the further deepening of structure optimization, in the future, it may even be considered to solidify the core parameters of the pre-trained model to the hardware level, including integrating them into a dedicated acceleration chip through parameter burning, or further, etching them into an optical quantum chip through an optoelectronic regulation structure, so as to achieve true end-side intelligence and a light-electric integrated reconstruction system without computational burden. This idea is expected to get rid of the dependence on high-performance computing resources, realize the real-time operation and low-power consumption deployment of polarization-spectrum joint reconstruction in scenarios such as space remote sensing, mobile terminals, and unmanned systems, reflecting the forward-looking and engineering feasibility of the present invention in terms of system architecture design and application expansion.
[0054] 5. Reconstruction Results
[0055] In this embodiment, the present invention is compared with other existing methods. (a) 3D Unet (b) Restormer (c) Ours (the present invention) (d) Bicubic interpolation. It can be seen that the method proposed in this embodiment is closest to the true value and achieves the best results in both the reconstruction of polarization degree, polarization angle, and spectral reconstruction.
Claims
1. A polarization spectrum reconstruction method based on a color polarization array, characterized in that The method includes the following steps: S1. Use an RGB color camera to obtain an unprocessed low-resolution mosaic image I0. A color polarization array is placed in front of the image sensor of the RGB color camera. S2. Interpolate and remove the mosaic from the unprocessed low-resolution mosaic image I0 to obtain a preliminarily processed image I1. S3. Input the image I1 into an end-to-end pre-trained deep learning model for polarization spectrum reconstruction to reconstruct a four-dimensional spectral polarization image.
2. The polarization spectrum reconstruction method based on a color polarization array according to claim 1, wherein In step S1, a color polarization array is placed in front of the image sensor of the RGB color camera. Specifically, the color polarization array is closely attached to the microlens array of the lens module in front of the image sensor.
3. The polarization spectrum reconstruction method based on a color polarization array according to claim 1, wherein In step S2, when interpolating and removing the mosaic from the unprocessed low-resolution mosaic image I0, specifically: Aggregate pixel points with the same color and the same polarization angle. Utilize the spatial similarity of the picture and the redundancy of the channel dimension to first interpolate and reshape to the original resolution. Then cascade two-dimensional images of different channels into a normal image and perform another spatial interpolation to further fuse image features.
4. The polarization spectrum reconstruction method based on a color polarization array according to claim 1, characterized in that, In step S3, the end-to-end pre-trained deep learning model adopts a convolutional neural network Transformer combined with a U-shaped architecture, including a three-dimensional convolutional module and a polarization spectrum attention block based on the Transformer model.
5. A polarization spectrum reconstruction method based on a color polarization array according to claim 4, characterized in that, In step S3, the three-dimensional convolutional module serves as an encoder to extract spatial-spectral-polarization joint features layer by layer.
6. The polarization spectrum reconstruction method based on a color polarization array according to claim 4, characterized in that In step S3, the polarization spectrum attention block includes a polarization spectrum attention module and a feed-forward network module. The polarization spectrum attention module is used for feature representation of the image. The feed-forward network module is used to predict the final reconstructed image from the extracted features.
7. A polarization spectrum reconstruction method based on a color polarization array according to claim 6, characterized in that, In step S3, the polarization spectrum attention module first performs a chunk embedding process on the input image, cuts the entire image into local chunks, converts each image chunk into a feature vector through a linear projection to form a Token sequence, and then the feed-forward network module performs position encoding, Query-Key-Value mapping, and attention weighting calculation on the Token sequence.
8. A polarization spectrum reconstruction method based on a color polarization array according to claim 6, characterized in that In step S3, the specific processing steps of the polarization spectrum attention module are as follows: S31. First, extract the feature representation of the image I1 through a three-dimensional convolutional block. S32. For the first polarization channel among the four polarization channels, that is, the polarization channel with a polarization direction having an angle of 0 degrees with the horizontal plane, regard the features of all spectral channels at each spatial position as a group of sequence Tokens. On this basis, set the first polarization channel as the target channel for attention calculation, acting as the value Value, and generate a query Query and a key Key respectively for the other three polarization channels through a group of shared linear transformations. S33. According to the dot product mechanism, calculate the dot product similarity between the query Query and the key Key, and normalize it through an activation function to obtain an attention weight matrix. S34. Apply the attention weight matrix to the target channel to achieve directed fusion and guided reconstruction of polarization-spectral joint information.