Self-supervised highlight removal method and system based on polarization information guidance

By employing a polarization-information self-supervised specular removal method, a specular confidence map and an encoding/decoding network are constructed using polarization imaging characteristics. This solves the problem of existing specular removal methods relying on labeled data, and achieves efficient specular removal and structure preservation under complex lighting conditions.

CN122492530APending Publication Date: 2026-07-31EAST CHINA JIAOTONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EAST CHINA JIAOTONG UNIVERSITY
Filing Date
2026-06-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing specular highlight removal methods rely on a large amount of labeled data, which is not adaptable enough, and the polarization information is not fully utilized, making it difficult to effectively distinguish between specular highlights and diffuse reflection under complex lighting conditions.

Method used

A self-supervised specular removal method based on polarization information is adopted. By acquiring intensity images under multiple polarization directions, the total light intensity, linear polarization degree and linear polarization angle images are calculated to construct a polarization specular confidence map. Spatial and channel attention modulation is performed using an encoder-decoder network, and self-supervised training is carried out in combination with a pseudo specular-free reference image. Various polarization consistency constraint losses are designed.

Benefits of technology

Accurately remove specular highlights under complex lighting conditions, reduce data acquisition and annotation costs, improve the targeting and robustness of specular removal, and preserve diffuse reflection structure and texture details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492530A_ABST
    Figure CN122492530A_ABST
Patent Text Reader

Abstract

This disclosure relates to a self-supervised specular removal method and system based on polarization information guidance, comprising: acquiring intensity images of a scene to be detected under multiple polarization directions; calculating a total intensity image, a linear polarization degree image, and a linear polarization angle image based on the intensity images under multiple polarization directions; constructing a polarization specular confidence map based on the above images; performing spatial attention and channel attention modulation on the intensity images under multiple polarization directions through an encoder-decoder network based on the polarization specular confidence map to obtain a feature map after specular removal; constructing a pseudo-spectrum-free reference image; determining a polarization consistency constraint loss based on the feature map after specular removal and the pseudo-spectrum-free reference image; performing self-supervised training on the encoder-decoder network; and processing the intensity images under multiple polarization directions through the trained encoder-decoder network to obtain a target image with specular removal. This disclosure can improve the robustness of specular removal under complex lighting conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of computer vision and polarization imaging technology, and in particular to a self-supervised specular removal method and system based on polarization information guidance. Background Technology

[0002] Specular reflection is a common challenge in computer vision tasks. The presence of specular highlights often leads to camera oversaturation, resulting in the loss of visual information about object surfaces and adversely affecting advanced processing tasks such as visual recognition, tracking, and stereo reconstruction. Therefore, effectively detecting and removing specular highlights is of great significance in practical applications.

[0003] Existing specular highlight removal methods are mainly divided into two categories: traditional methods and deep learning methods. Traditional methods are typically based on color priors or reflection models, separating diffuse and specular reflection components by analyzing the color distribution or brightness characteristics of an image. However, these methods cannot semantically distinguish between specular highlights and white objects, and are prone to significant errors when both coexist. From an optical perspective, specular reflection has a high degree of polarization, while diffuse reflection has relatively weak polarization characteristics; therefore, polarization information can effectively distinguish different types of reflection components. Some studies have attempted to utilize polarization information for reflection removal, but these often rely on strictly controlled light sources, which has significant limitations under real-world, complex lighting conditions.

[0004] In recent years, researchers have proposed deep learning schemes based on polarization information, such as two-stage reflection removal networks and generative adversarial networks. However, these supervised learning methods typically rely on large amounts of paired training data, while obtaining high-quality specular-free reference images in real-world scenes is extremely difficult, limiting their practical applications. Therefore, how to provide a specular removal technique that can utilize polarization physics priors, does not rely on real specular-free labeled data, and simultaneously achieves specular suppression and structure preservation has become a pressing technical problem for those skilled in the art. Summary of the Invention

[0005] To address the problems of existing specular removal methods, such as reliance on large amounts of labeled data, insufficient adaptability under complex lighting conditions, and inadequate utilization of polarization information, this disclosure proposes a self-supervised specular removal method guided by polarization information to solve these issues.

[0006] According to one aspect of this disclosure, a self-supervised specular removal method guided by polarization information is provided, comprising: S10. Obtain intensity images of the scene to be detected under multiple polarization directions; S20. Calculate the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image based on the intensity images under the multiple polarization directions; S30. Construct a polarization spectro confidence map based on the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image; S40. Based on the polarization specular confidence map, spatial attention and channel attention modulation are applied to the intensity images under the multiple polarization directions through an encoding and decoding network to obtain a feature map after removing specular highlights. S50. Construct a pseudo-highlight-free reference image, determine the polarization consistency constraint loss based on the feature map after removing the highlights and the pseudo-highlight-free reference image, and perform self-supervised training on the encoder-decoder network according to the polarization consistency constraint loss. Process the intensity images under the multiple polarization directions through the trained encoder-decoder network to obtain the target image with the highlights removed.

[0007] Preferably, the total light intensity image, the linear polarization degree image, and the linear polarization angle image are respectively represented as: , , , In the formula, , , and The graphs represent the intensity images at polarization directions of 0°, 45°, 90°, and 135°, respectively. S0 represents the orthogonal radiation intensity of light, obtained from any two orthogonal polarization states. S1 represents the difference between the intensity of horizontally polarized light and the intensity of vertically polarized light. S2 represents the difference between the intensity of linearly polarized light at 45° and the intensity of linearly polarized light at 135°. This is the total light intensity image. This is a linear polarization degree image. This is a linear polarization angle image. These are pixel coordinates.

[0008] Preferably, the process of constructing the polarization spectrophotometry confidence map includes: The total light intensity image is normalized to obtain a normalized total light intensity image; Calculate the local variance of each pixel in the linear polarization angle image to obtain the local variance map of the linear polarization angle; The normalized total light intensity image, the linear polarization degree image, and the local variance map of the linear polarization angle are weighted and fused to obtain a weighted fusion result. The weighted fusion result is mapped to a predetermined numerical range to obtain a polarization hyperspectral confidence map.

[0009] Preferably, the polarization spectrophotometry confidence map is represented as follows: , In the formula, This is a confidence map of polarized spectra. , and These are the weighting coefficients. This is a normalized total light intensity image. This is a linear polarization degree image. This represents the local standard deviation of the linear polarization angle.

[0010] Preferably, spatial attention and channel attention modulation are performed on the intensity images under the multiple polarization directions using an encoding / decoding network, including: A spatial weight map is generated based on the polarization hyperspectral confidence map, and the spatial weight map is used to perform spatial attention modulation on the feature map in the encoder-decoder network to obtain a spatially modulated feature map. The spatially modulated feature map is globally pooled and then transformed by a multilayer perceptron to obtain the channel weight vector. Based on the channel weight vector, channel attention modulation is applied to the spatially modulated feature map to obtain a channel-modulated feature map.

[0011] Preferably, the pseudo-highlight-free reference image is represented as: , In the formula, This is a pseudo-no-highlight reference image. , , and These represent intensity images at polarization directions of 0°, 45°, 90°, and 135°, respectively.

[0012] Preferably, the polarization consistency constraint loss includes reconstruction loss, polarization degree suppression loss, specular residual suppression loss, and specular region smoothing loss, wherein the reconstruction loss is expressed as: , In the formula, To rebuild the losses, The diffuse reflection intensity image output by the network. This is a pseudo-no-highlight reference image. N This represents the number of images input to the network in each iteration. i Here is the sample index in the batch, indicating the number of samples in the current batch. i Zhang Image; The polarization degree suppression loss is expressed as: , In the formula, To suppress the loss of polarization degree, Represents the set of pixels in the highlight area. This represents the number of pixels in that area. These are pixel coordinates; The specular residual suppression loss is expressed as: , , In the formula, To suppress specular residual loss, To take a positive function, The residual ratio threshold, This is the total light intensity image. To predict the residual image; The smoothing loss in the highlight region is expressed as: , In the formula, To smooth out the loss in the highlight areas, and These are the gradient operators for the horizontal and vertical directions, respectively. This is a confidence plot of polarized spectra.

[0013] According to one aspect of this disclosure, a self-supervised specular removal system guided by polarization information is provided, comprising: The image acquisition module acquires intensity images of the scene to be detected under multiple polarization directions; The polarization image calculation module calculates the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image based on the intensity images under the multiple polarization directions. The polarization spectrophotometer construction module constructs a polarization spectrophotometer based on the total light intensity image, the linear polarization degree image, and the linear polarization angle image. The polarization specular perception attention module modulates the intensity images under multiple polarization directions with spatial attention and channel attention through an encoding and decoding network based on the polarization specular confidence map, thereby obtaining a feature map after removing specular highlights. The self-supervised training module constructs a pseudo-highlight-free reference image, determines the polarization consistency constraint loss based on the feature map after highlight removal and the pseudo-highlight-free reference image, and performs self-supervised training on the encoder-decoder network according to the polarization consistency constraint loss. The trained encoder-decoder network processes the intensity images under multiple polarization directions to obtain the target image with highlights removed.

[0014] According to one aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: execute the above-described polarization-information-guided self-supervised specular removal method.

[0015] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program / instructions and a bit stream thereon, wherein the computer program instructions, when executed by a processor, implement the above-described self-supervised specular removal method guided by polarization information to generate the bit stream.

[0016] Compared to the prior art, the beneficial effects of this disclosure are as follows: 1) This disclosure introduces polarization imaging information and utilizes the essential difference in polarization characteristics between specular reflection and diffuse reflection to construct a polarization specular highlight confidence map to explicitly model the highlight region, thereby effectively distinguishing specular highlights from light-colored object surfaces and achieving more accurate highlight localization and suppression under complex lighting conditions.

[0017] 2) This disclosure adopts a self-supervised learning framework and constructs a pseudo diffuse reflection reference image based on the minimum polarization energy assumption. It can complete network training without relying on real non-highlight labeled data, which significantly reduces the cost of data acquisition and labeling and improves the feasibility of actual deployment.

[0018] 3) This disclosure constructs a polarization-guided encoder-decoder network and introduces a polarization-guided specular awareness attention module. This module explicitly suppresses specular regions through a spatial attention branch and enhances feature responses related to polarization changes through a channel attention branch, enabling the network to adaptively focus on specular regions during multi-scale feature learning, thereby improving the targeting and effectiveness of specular removal.

[0019] 4) This disclosure designs various polarization consistency constraint losses, including reconstruction loss, polarization degree suppression loss, specular residual suppression loss, and specular region smoothing loss. By jointly optimizing the above losses, it is possible to maintain diffuse reflection structure and texture details while suppressing specular components, achieving physically consistent specular removal effect.

[0020] 5) This disclosure improves the specular removal effect while maintaining the structure preservation capability by combining polarization physics priors with deep learning models, and maintains good robustness under complex lighting conditions.

[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure.

[0022] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0024] Figure 1 A flowchart of a self-supervised specular removal method based on polarization information guidance in an embodiment of this disclosure is shown; Figure 2 A schematic diagram of a polarization spectrophotometer confidence map in an embodiment of this disclosure is shown; Figure 3 A schematic diagram of the polarization-guided specular removal network structure in an embodiment of this disclosure is shown; Figure 4 A schematic diagram of the polarization specular sensing attention module structure in an embodiment of this disclosure is shown; Figure 5 The following diagram shows a comparison of the highlight removal effects in the embodiments of this disclosure; Figure 6 A block diagram of a self-supervised specular removal system based on polarization information guidance in an embodiment of this disclosure is shown. Detailed Implementation

[0025] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0026] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0027] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0028] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of this disclosure, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0030] Based on the above ideas, this disclosure proposes a self-supervised specular removal method guided by polarization information. Figure 1 A flowchart illustrating a self-supervised specular removal method guided by polarization information is shown. The method includes: S10. Obtain intensity images of the scene to be detected under multiple polarization directions; S20. Calculate the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image based on the intensity images under the multiple polarization directions; S30. Construct a polarization spectro confidence map based on the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image; S40. Based on the polarization specular confidence map, spatial attention and channel attention modulation are applied to the intensity images under the multiple polarization directions through an encoding and decoding network to obtain a feature map after removing specular highlights. S50. Construct a pseudo-highlight-free reference image, determine the polarization consistency constraint loss based on the feature map after removing the highlights and the pseudo-highlight-free reference image, and perform self-supervised training on the encoder-decoder network according to the polarization consistency constraint loss. Process the intensity images under the multiple polarization directions through the trained encoder-decoder network to obtain the target image with the highlights removed.

[0031] This embodiment of the disclosure acquires intensity images of the scene to be detected in four polarization directions: 0°, 45°, 90°, and 135°. Stokes parameters are calculated to obtain total light intensity images, linear polarization degree images, and linear polarization angle images. Based on the above images, a polarization specular highlight confidence map is constructed to explicitly model the highlight region, thereby effectively distinguishing specular highlights from light-colored object surfaces and improving the accuracy and robustness of highlight removal under complex lighting conditions. This embodiment also introduces a polarization specular perception attention module. During the feature extraction stage of the encoding / decoding network, spatial attention weights are generated using a polarization specular confidence map to explicitly suppress specular regions. Simultaneously, the feature response related to polarization changes is enhanced through a channel attention branch, enabling the network to adaptively focus on specular regions. Furthermore, a pseudo diffuse reflection reference image based on the minimum polarization energy assumption is constructed, and the network is self-supervised by combining various polarization consistency constraints such as reconstruction loss, polarization degree suppression loss, specular residual suppression loss, and specular region smoothing loss. This allows the network to maintain diffuse reflection structure and texture details while suppressing specular components, thereby achieving a physically consistent specular removal effect that balances model efficiency and practical deployment feasibility.

[0032] This disclosure further extends the above-described method with detailed possible implementations, specifically including: S10. Obtain intensity images of the scene to be detected under multiple polarization directions.

[0033] In one embodiment, an area-array polarization camera designed using the focal plane segmented polarization imaging principle acquires intensity images of the scene to be detected under multiple polarization directions. The intensity images at four different polarization angles are the intensity images at 0°, 45°, 90°, and 135°, respectively, and are represented as follows: , , and .

[0034] The intensity image is preferably acquired using a camera with a focal plane array polarization sensor, such as the MER2-503-23GC-P polarization camera. This type of polarization camera integrates micro-polarizers of different orientations onto the pixel surface of a traditional CMOS image sensor, enabling the simultaneous acquisition of light intensity information in multiple polarization directions during a single exposure. Specifically, the four polarizers at different angles correspond to polarization components of 0°, 45°, 90°, and 135°, respectively. By combining and calculating adjacent pixels, the polarization state information of the scene light can be obtained for subsequent calculations of the degree of polarization and the polarization angle.

[0035] The aforementioned polarization camera can directly acquire intensity images in multiple polarization directions without requiring multiple shots or mechanical polarizer switching. This avoids the registration errors and system complexity issues caused by time-division acquisition in traditional polarization imaging, and improves the real-time performance and stability of the overall data acquisition process. All intensity images are saved in 8-bit grayscale PNG format with an original resolution of 2448×2048 pixels. To reduce computational overhead and maintain the main structure, all intensity images are bilinearly interpolated and scaled to a uniform size before training; in this embodiment, 256×256 pixels is used.

[0036] S20. Calculate the total light intensity image, linear polarization degree image, and linear polarization angle image based on the intensity images under the multiple polarization directions.

[0037] In this embodiment, the total light intensity image, linear polarization degree image, and linear polarization angle image obtained by Stokes vector calculation are respectively represented as follows: , , , In the formula, , , and The graphs represent the intensity images at polarization directions of 0°, 45°, 90°, and 135°, respectively. S0 represents the orthogonal radiation intensity of light, obtained from any two orthogonal polarization states. S1 represents the difference between the intensity of horizontally polarized light and the intensity of vertically polarized light. S2 represents the difference between the intensity of linearly polarized light at 45° and the intensity of linearly polarized light at 135°. This is the total light intensity image. This is a linear polarization degree image, with values ​​ranging from [0,1]. This is a linear polarization angle image, with values ​​ranging from [0, π). These are pixel coordinates.

[0038] S30. Construct a polarization spectrophotometry confidence map based on the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image.

[0039] In this embodiment, specular reflection and diffuse reflection exhibit significant differences in polarization intensity and stability during polarization imaging. Specular reflection typically has a high degree of linear polarization, and its linear polarization angle changes relatively smoothly in local space. In contrast, diffuse reflection has a lower degree of polarization and a more unstable polarization angle distribution due to multiple scattering. Based on these physical characteristics, this embodiment constructs a polarization specular confidence map to characterize the distribution of regions in the image where specular reflection dominates.

[0040] The process of constructing the polarization spectro confidence map includes: normalizing the total intensity image to obtain a normalized total intensity image; calculating the local variance of each pixel in the linear polarization angle image to obtain a linear polarization angle local variance map; weightedly fusing the normalized total intensity image, the linear polarization angle image, and the linear polarization angle local variance map to obtain a weighted fusion result; and mapping the weighted fusion result to a predetermined numerical range to obtain the polarization spectro confidence map. Figure 2 A schematic diagram of constructing a polarization spectrophotometer confidence map, wherein, Figure 2 (a) in the image is the input image. Figure 2 (b) in the image is a polarization spectrophotometry confidence map. Figure 2 (c) in the image represents the highlight mask. Figure 2 In the example, (d) is the final image after removing highlights using the method in this embodiment. First, the total light intensity image is normalized, and then the local variance of the linear polarization angle image is calculated. The normalized total light intensity image, the linear polarization degree image, and the local variance of the linear polarization angle are weighted and fused. After mapping by the Sigmoid function, a polarization highlight confidence map with values ​​in the range [0,1] is obtained, which is used to characterize the confidence of the specular reflection region in the image.

[0041] Specifically, the total light intensity image is first normalized to obtain a normalized total light intensity image, which is represented as follows: , In the formula, This is a normalized total light intensity image. This is the total light intensity image.

[0042] The local variance map of AoLP is calculated to characterize the spatial stability of the polarization angle. For pixel position x, the polarization angle variance map of its neighborhood (using a 5×5 window in this embodiment) is defined as follows: , In the formula, For neighborhood windows, The total number of neighboring pixels. Indicates the neighborhood The mean, Let be the linear polarization angle at pixel position y. Normalizing to [0,1] yields the normalized local variance plot of the polarization angle. .

[0043] The normalized total light intensity image, the linear polarization degree image, and the local variance map of the linear polarization angle are weighted and fused together: , , In the formula, For the weighted fusion result, This is a confidence map of polarized spectra. , and For weighting coefficients, in this embodiment, they are all set to 1.0. This is the Sigmoid function, with a coefficient of 4 used to enhance contrast. This is the normalized local variance map of the polarization angle. Through the above mapping, the values ​​are restricted to the (0,1) interval, and the closer the value is to 1, the greater the probability that the pixel belongs to the specular reflection-dominated region.

[0044] In this embodiment, the polarization spectrophotometer is... As a core link connecting polarization physics priors and deep learning networks, it plays a crucial role, providing the network with explicit specular localization priors. In the polarization specular perception attention module, It directly participates in generating the spatial weight map, modulating the feature map pixel by pixel, so that the feature response of high-confidence regions is explicitly weakened while the features of low-confidence regions are preserved; at the same time... It also participates as a weighting factor in the calculation of polarization degree suppression loss and highlight region smoothing loss, constraining the network to apply stronger smoothness and lower polarization response in the highlight region. Through the above methods, By effectively embedding polarization physics priors into the feature learning and loss optimization process of the network, the network's behavior is highly consistent with the physical intuition of "highlight suppression and diffuse reflection preservation," thereby significantly improving the targeting and physical consistency of highlight removal.

[0045] S40. Based on the polarization specular confidence map, spatial attention and channel attention modulation are performed on the intensity images under the multiple polarization directions through an encoding and decoding network to obtain a feature map after specular removal.

[0046] In this embodiment, a schematic diagram of the polarization-guided specular removal network structure is shown below. Figure 3As shown, the network is used for specular suppression and diffuse reflection reconstruction of polarized images. The network input consists of intensity images in four polarization directions: 0°, 45°, 90°, and 135°. A polarization feature extraction module calculates the total intensity image, linear polarization degree image, and linear polarization angle image based on Stokes parameters. Simultaneously, the total intensity image is normalized, and the local variance of the linear polarization angle image is calculated. The normalized total intensity image, linear polarization degree image, and local variance of the linear polarization angle are weighted and fused, and after activation function mapping, a polarization specular confidence map is obtained. This polarization specular confidence map characterizes the probability distribution of specular regions. The intensity images in the four polarization directions and the polarization specular confidence map are concatenated along the channel dimension and input into an encoder-decoder network. The encoder-decoder network adopts a downsampling-upsampling U-shaped structure, supplemented with skip connections to preserve spatial detail information. A polarization specular perception attention module is embedded in each stage of the encoder's feature extraction process, guided by the polarization specular confidence map, to finally output a diffuse reflection intensity image after specular removal.

[0047] The input consists of intensity images in four polarization directions, which are then concatenated with the polarization spectro confidence map along the channel dimension, as follows: , The encoder consists of four progressively downsampled feature extraction stages, each containing two convolutional blocks. Each convolutional block is composed of a 3×3 convolution, batch normalization, and a SiLU activation function. The number of channels in each encoding stage is 32, 64, 128, and 256, respectively, and the spatial resolution is progressively reduced through max pooling.

[0048] At the lowest level, the network employs a bottleneck module to further model high-level semantic features. Subsequently, the decoder performs progressive upsampling through transposed convolutions and makes skip connections with feature maps of corresponding scales in the encoder to recover spatial detail. The decoding stage uses a dual convolutional structure to fuse and reconstruct the concatenated features. Finally, the network maps the features to a single-channel output using 1×1 convolutions and uses the Sigmoid function to constrain the prediction results to the [0,1] interval, obtaining a polarization spectrophotometry confidence map.

[0049] To further enhance the network's feature extraction capability in highlight regions, this embodiment introduces a polarization-spectrum-aware attention module (PHAM) during the feature extraction stage, such as... Figure 4As shown, this module receives two inputs simultaneously: a feature map from the encoder-decoder network and a polarization spectrophotometry confidence map constructed from polarization physics priors. The module is divided into parallel channel attention branches and spatial attention branches, ultimately achieving adaptive modulation of the input features through element-wise multiplication. The channel attention branch processes the input feature map: first, it aggregates global information of the features through adaptive pooling, then passes it through two 1×1 convolutions and corresponding activation functions to generate a channel weight vector matching the feature channel dimensions. This weight vector is used to recalibrate the importance of each channel in the feature map.

[0050] The spatial attention branch uses the polarization specular confidence map as guiding information: the polarization specular confidence map is subjected to 3×3 convolution, activation function, 1×1 convolution and activation function to generate a spatial weight map. This weight map can adaptively modulate different spatial positions of the feature map according to the distribution of specular regions. Finally, the channel weights and spatial weights are jointly weighted on the input feature map to obtain the output feature map enhanced by specular perception attention, achieving the suppression of specular region features and the enhancement of non-spectral region features.

[0051] Spatial attention and channel attention modulation are performed on intensity images under multiple polarization directions using an encoding-decoding network, including: generating a spatial weight map based on the polarization spectrophotometer confidence map; using the spatial weight map to perform spatial attention modulation on the feature map in the encoding-decoding network to obtain a spatially modulated feature map; performing global pooling on the spatially modulated feature map and transforming it through a multilayer perceptron to obtain a channel weight vector; and performing channel attention modulation on the spatially modulated feature map based on the channel weight vector to obtain a channel modulated feature map.

[0052] The spatial attention branch generates a spatial weight map based on the polarization spectrophotometry confidence map and modulates the feature map pixel by pixel. First, a convolutional smoothing operation is performed on the feature map to enhance its spatial continuity. Then, the sigmoid function is used to map it to the [0,1] interval to obtain the spatial weight map, represented as: , In the formula, For spatial weighting, This is a convolution operation.

[0053] This spatial weight map is used for pixel-by-pixel modulation of the feature map, and is represented as follows: , In the formula, ⊙ represents element-wise multiplication. This design reflects the physical intuition of "suppressing highlights and preserving non-highlight areas." For the current level feature map, This is a spatially modulated feature map. When... When it is large, When the value is close to 1, the characteristic response at the corresponding position is explicitly weakened.

[0054] Channel attention branch on spatially modulated feature map Perform global average pooling: , In the formula, Here, H represents the feature map after global pooling, and W represents the height and width of the feature map.

[0055] The channel response is modeled using a multilayer perceptron consisting of two fully connected layers: , In the formula, This is the channel weight vector. For activation function, This is the weight matrix of the second fully connected layer. This is the weight matrix of the first fully connected layer.

[0056] Channel attention is applied to the spatially suppressed feature map, resulting in a channel-modulated feature map, represented as follows: , In the formula, This is a feature map after channel modulation.

[0057] S50. Construct a pseudo-highlight-free reference image, determine the polarization consistency constraint loss based on the feature map after removing the highlights and the pseudo-highlight-free reference image, and perform self-supervised training on the encoder-decoder network according to the polarization consistency constraint loss. Process the intensity images under the multiple polarization directions through the trained encoder-decoder network to obtain the target image with the highlights removed.

[0058] Since it is difficult to obtain paired "highlight-no-highlight" supervision data in real-world scenarios, this embodiment constructs a polarization consistency constraint loss based on polarization physical consistency. This loss function consists of four parts: reconstruction loss, polarization degree suppression loss, specular residual suppression loss, and specular region smoothing loss.

[0059] The reconstruction loss measures the difference between the diffuse image output by the network and the pseudo-diffuse reference image. Based on the minimum polarization energy assumption, the pixel with the minimum value in the four polarization directions is closer to the diffuse component, thus constructing a pseudo-no-spectrum reference image, denoted as: , In the formula, This is a pseudo-no-highlight reference image. , , and These represent intensity images at polarization directions of 0°, 45°, 90°, and 135°, respectively.

[0060] The reconstruction loss is expressed as: , In the formula, To rebuild the losses, The diffuse reflection intensity image output by the network. This is a pseudo-no-highlight reference image. N This represents the number of images input to the network in each iteration. i Here is the sample index in the batch, indicating the number of samples in the current batch. i Zhang Image; In polarization imaging, the specular reflection region typically has a high degree of linear polarization. The polarization degree is lower in the diffuse reflection region. Based on this physical principle, this embodiment designs a polarization degree suppression loss to effectively suppress the polarization degree response of the output image in the highlight region. Specifically, it utilizes the input polarization image's polarization... As a weight, the polarization degree suppression loss is expressed as reducing the brightness of the output in the highlight region: , In the formula, To suppress the loss of polarization degree, Represents the set of pixels in the highlight area. This represents the number of pixels in that area. These are pixel coordinates; To enhance the network's ability to destigmatize in high-light regions, a prediction residual image is constructed, which is the difference between the total intensity image and the diffuse reflection intensity image output by the network. A specular residual suppression loss is set to suppress intensity values ​​in the prediction residual image that exceed those in the total intensity image. The specular residual suppression loss is expressed as follows: (The constraint is applied at a multiple of the pixel count.) , , In the formula, To suppress specular residual loss, To take a positive function, The residual ratio threshold, This is the total light intensity image. To predict the residual image; To maintain the spatial continuity of the output image and reduce artifacts, a highlight region-weighted smoothing loss is introduced, which is expressed as: , In the formula, To smooth out the loss in the highlight areas, and These are the gradient operators for the horizontal and vertical directions, respectively. For polarized specular confidence maps, this loss imposes a stronger smoothing constraint on specular regions, as these regions are prone to artifacts, while allowing natural textures to be preserved in non-spectral regions.

[0061] Figure 5 This is a schematic diagram comparing the highlight removal effect of the present disclosure embodiment with that of existing methods, wherein, Figure 5 (a) in the image is the input image. Figure 5 Image (b) shows the image result based on Umeyama's specular removal method. Figure 5 Image (c) in the image shows the result based on the Zhu specular removal method. Figure 5 In the image, (d) represents the image result based on the Lei specular removal method. Figure 5 Image (e) in the image represents the result based on Tang's specular removal method. Figure 5 In the diagram, (f) represents the image result of the highlight removal method in this embodiment. Figure 5 As can be seen, the method in this embodiment can preserve the surface texture and structural details of an object while suppressing specular highlights.

[0062] The polarization-guided specular removal network described in this embodiment can collaboratively model the physical information in multi-angle polarized images without relying on traditional methods based on color priors or large amounts of labeled data. The key idea of ​​this method is to explicitly locate specular reflection regions by constructing a polarization specular confidence map, adaptively modulate specular features using an encoder-decoder network combined with a polarization specular perception attention module, and achieve self-supervised training through polarization consistency constraint loss. Using this technical solution, the accuracy and robustness of specular removal under complex lighting conditions can be significantly improved, while preserving diffuse reflection structures and texture details, balancing model efficiency and practical deployment feasibility.

[0063] As another aspect of this disclosure, a self-supervised specular removal system 100 guided by polarization information is also provided, such as... Figure 6 As shown, it includes: Image acquisition module 1 acquires intensity images of the scene to be detected under multiple polarization directions; The polarization image calculation module 2 calculates the total light intensity image, the linear polarization degree image, and the linear polarization angle image based on the intensity images under the multiple polarization directions. The polarization spectro confidence map construction module 3 constructs a polarization spectro confidence map based on the total light intensity image, the linear polarization degree image, and the linear polarization angle image. The polarization specular perception attention module 4 modulates the intensity images under the multiple polarization directions with spatial attention and channel attention through an encoding and decoding network based on the polarization specular confidence map to obtain a feature map after removing specular highlights. The self-supervised training module 5 constructs a pseudo-highlight-free reference image, determines the polarization consistency constraint loss based on the feature map after highlight removal and the pseudo-highlight-free reference image, and performs self-supervised training on the encoder-decoder network according to the polarization consistency constraint loss. The trained encoder-decoder network processes the intensity images under multiple polarization directions to obtain the target image with highlights removed.

[0064] Without causing contradictions, the above-described modules in the system of the present disclosure embodiments can implement any of the above-described methods.

[0065] Based on the description of the above embodiments, it can be seen that the embodiments of this disclosure can achieve the following technical effects: 1) This embodiment of the present disclosure introduces polarization imaging information and utilizes the essential difference in polarization characteristics between specular reflection and diffuse reflection to construct a polarization specular highlight confidence map to explicitly model the highlight region, thereby effectively distinguishing specular highlights from light-colored object surfaces and achieving more accurate highlight positioning and suppression under complex lighting conditions.

[0066] 2) The embodiments of this disclosure adopt a self-supervised learning framework and construct a pseudo diffuse reflection reference image based on the assumption of minimum polarization energy. The network training can be completed without relying on real non-highlight labeled data, which significantly reduces the cost of data acquisition and labeling and improves the feasibility of actual deployment.

[0067] 3) This embodiment of the present disclosure constructs a polarization-guided encoder-decoder network and introduces a polarization-guided specular awareness attention module. This module explicitly suppresses specular regions through a spatial attention branch and enhances feature responses related to polarization changes through a channel attention branch, enabling the network to adaptively focus on specular regions during multi-scale feature learning, thereby improving the targeting and effectiveness of specular removal.

[0068] 4) This disclosure incorporates various polarization consistency constraint losses, including reconstruction loss, polarization degree suppression loss, specular residual suppression loss, and specular region smoothing loss. By jointly optimizing these losses, diffuse reflection structure and texture details can be maintained while suppressing specular components, resulting in physically consistent specular removal.

[0069] 5) This embodiment of the present disclosure combines polarization physics priors with deep learning models to improve the specular removal effect while maintaining the structure preservation capability, and maintains good robustness under complex lighting conditions.

[0070] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured for the aforementioned self-supervised specular removal method guided by polarization information. The electronic device can be provided as a terminal, a server, or other type of device.

[0071] This disclosure also proposes a computer-readable storage medium storing a computer program / instructions and a bitstream thereon. When the computer program / instructions are executed by a processor, they implement the aforementioned polarization-information-guided self-supervised specular removal method to generate the bitstream. The computer-readable storage medium can be a non-volatile computer-readable storage medium.

[0072] Those skilled in the art will understand that, in the above-described self-supervised specular removal method and system based on polarization information in specific embodiments, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0074] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A self-supervised specular removal method guided by polarization information, characterized in that, include: S10. Obtain intensity images of the scene to be detected under multiple polarization directions; S20. Calculate the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image based on the intensity images under the multiple polarization directions; S30. Construct a polarization spectro confidence map based on the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image; S40. Based on the polarization specular confidence map, spatial attention and channel attention modulation are applied to the intensity images under the multiple polarization directions through an encoding and decoding network to obtain a feature map after removing specular highlights. S50. Construct a pseudo-highlight-free reference image, determine the polarization consistency constraint loss based on the feature map after removing the highlights and the pseudo-highlight-free reference image, and perform self-supervised training on the encoder-decoder network according to the polarization consistency constraint loss. Process the intensity images under the multiple polarization directions through the trained encoder-decoder network to obtain the target image with the highlights removed.

2. The method according to claim 1, characterized in that, The total light intensity image, linear polarization degree image, and linear polarization angle image are respectively represented as follows: , , , In the formula, , , and The graphs represent the intensity images at polarization directions of 0°, 45°, 90°, and 135°, respectively. S0 represents the orthogonal radiation intensity of light, obtained from any two orthogonal polarization states. S1 represents the difference between the intensity of horizontally polarized light and the intensity of vertically polarized light. S2 represents the difference between the intensity of linearly polarized light at 45° and the intensity of linearly polarized light at 135°. This is the total light intensity image. This is a linear polarization degree image. This is a linear polarization angle image. These are pixel coordinates.

3. The method according to claim 1, characterized in that, The process of constructing the polarization spectrophotometry confidence map includes: The total light intensity image is normalized to obtain a normalized total light intensity image; Calculate the local variance of each pixel in the linear polarization angle image to obtain the local variance map of the linear polarization angle; The normalized total light intensity image, the linear polarization degree image, and the local variance map of the linear polarization angle are weighted and fused to obtain a weighted fusion result. The weighted fusion result is mapped to a predetermined numerical range to obtain a polarization hyperspectral confidence map.

4. The method according to claim 3, characterized in that, The polarization spectro confidence map is represented as follows: , In the formula, This is a confidence map of polarized spectra. , and These are the weighting coefficients. This is a normalized total light intensity image. This is a linear polarization degree image. This represents the local standard deviation of the linear polarization angle.

5. The method according to claim 1, characterized in that, Spatial attention and channel attention modulation are performed on the intensity images under the multiple polarization directions using an encoding / decoding network, including: A spatial weight map is generated based on the polarization hyperspectral confidence map, and the spatial weight map is used to perform spatial attention modulation on the feature map in the encoder-decoder network to obtain a spatially modulated feature map. The spatially modulated feature map is globally pooled and then transformed by a multilayer perceptron to obtain the channel weight vector. Based on the channel weight vector, channel attention modulation is applied to the spatially modulated feature map to obtain a channel-modulated feature map.

6. The method according to claim 1, characterized in that, The pseudo-highlight-free reference image is represented as follows: , In the formula, This is a pseudo-no-highlight reference image. , , and These represent intensity images at polarization directions of 0°, 45°, 90°, and 135°, respectively.

7. The method according to claim 1, characterized in that, The polarization consistency constraint loss includes reconstruction loss, polarization degree suppression loss, specular residual suppression loss, and specular region smoothing loss, wherein the reconstruction loss is expressed as: , In the formula, To rebuild the losses, The diffuse reflection intensity image output by the network. This is a pseudo-no-highlight reference image. N This represents the number of images input to the network in each iteration. i Here is the sample index in the batch, indicating the number of samples in the current batch. i Zhang Image; The polarization degree suppression loss is expressed as: , In the formula, To suppress the loss of polarization degree, Represents the set of pixels in the highlight area. This represents the number of pixels in that area. These are pixel coordinates; The specular residual suppression loss is expressed as: , , In the formula, To suppress specular residual loss, To take a positive function, The residual ratio threshold, This is the total light intensity image. To predict the residual image; The smoothing loss in the highlight region is expressed as: , In the formula, To smooth out the loss in the highlight areas, and These are the gradient operators for the horizontal and vertical directions, respectively. This is a confidence plot of polarized spectra.

8. A self-supervised specular removal system guided by polarization information, characterized in that, include: The image acquisition module acquires intensity images of the scene to be detected under multiple polarization directions; The polarization image calculation module calculates the total light intensity image, the degree of linear polarization image, and the angle of linear polarization image based on the intensity images under the multiple polarization directions. The polarization spectrophotometer construction module constructs a polarization spectrophotometer based on the total light intensity image, the linear polarization degree image, and the linear polarization angle image. The polarization specular perception attention module modulates the intensity images under multiple polarization directions with spatial attention and channel attention through an encoding and decoding network based on the polarization specular confidence map, thereby obtaining a feature map after removing specular highlights. The self-supervised training module constructs a pseudo-highlight-free reference image, determines the polarization consistency constraint loss based on the feature map after highlight removal and the pseudo-highlight-free reference image, and performs self-supervised training on the encoder-decoder network according to the polarization consistency constraint loss. The trained encoder-decoder network processes the intensity images under multiple polarization directions to obtain the target image with highlights removed.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the self-supervised specular removal method based on polarization information as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program / instructions and a bit stream thereon, characterized in that, When the program / instruction is executed by the processor, it implements the self-supervised specular removal method based on polarization information as described in any one of claims 1 to 7 to generate the bitstream.