Flare-removing water surface target segmentation method and system based on polarization feature fusion
By using a polarization feature fusion method based on GFNet and CDDFuse networks, the problem of target blurring caused by strong reflection interference in complex waters was solved, achieving high-precision water surface target segmentation and improving robustness and adaptability.
Patent Information
- Application Number
- CN202511687896.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-01-09
AI Technical Summary
Existing polarization imaging technology is insufficient in robustness, accuracy, and adaptability when faced with strong reflection interference in complex waters, making it difficult to achieve high-precision automatic identification and segmentation of water surface targets.
A polarization feature fusion-based approach is adopted, which separates the reflection and illumination components through the GFNet enhancement network, and combines the CDDFuse fusion network to perform high and low frequency feature decomposition to generate a flare-free water surface target image. Semantic segmentation is then performed using a deep learning model.
It significantly improves the ability to detect and segment water surface targets under conditions of strong glare and water wave disturbance, and enhances the separability of targets from the background and the segmentation accuracy.
Smart Images

Figure CN121305083A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of remote sensing, and particularly relates to a sun flare removing water surface target segmentation method and system based on polarization feature fusion. BACKGROUND
[0002] With the continuous development of ocean science and intelligent sensing technology, unmanned underwater vehicles (UUV) are increasingly applied in the fields of ocean exploration, environmental monitoring, underwater archaeology, and marine security. UUVs have the ability to operate autonomously for a long time, greatly expanding the space for human cognition and utilization of the ocean. However, in near-shore shallow water and other environments, there are many challenges in automatically identifying and segmenting UUVs using remote sensing technology, especially in the presence of strong sunlight reflection (solar flares) and other abnormal lighting interference. Traditional optical imaging methods are often limited by the high reflectivity and complex background of water bodies, resulting in blurred targets, unclear boundaries, and even complete masking of targets by strong light, making it difficult to achieve high-precision automatic identification and segmentation.
[0003] In recent years, polarization imaging technology has received a lot of attention due to its ability to obtain polarization state information. As an important physical dimension of light waves, polarization properties can reflect key information such as object material, structure, and surface characteristics. Typical polarization image parameters such as linear degree of polarization (DOLP), polarization angle (AOP), and parallel / vertical polarization radiation (PPR / VPR) can effectively enhance the texture and boundary features of objects, improve distinguishability in homogeneous backgrounds, and especially suppress water surface glare under certain lighting conditions, showing significant potential for better performance than ordinary imaging in complex water environments. Therefore, polarization imaging is considered one of the important physical means to improve the detection accuracy of water surface and shallow water targets.
[0004] Multi-modal image fusion (MMIF) aims to combine information from different source images to generate more discriminative and robust fusion images. The main focus of deep fusion research on polarization images is on the fusion of DOLP and intensity images, or joint feature extraction of DOLP and AOP. However, these methods are mostly not optimized for strong specular reflection interference scenarios such as solar flares, making it difficult to stably output high-quality segmentation results in actual complex water conditions.
[0005] In summary, although current polarization imaging fusion methods provide important technical support for water surface target detection, there are still obvious shortcomings in robustness, accuracy, and adaptability when facing complex water reflection interference scenarios. SUMMARY
[0006] This invention provides a flare removal water surface target segmentation method and system based on polarization feature fusion, to solve the problem that existing technologies still have significant shortcomings in robustness, accuracy and adaptability when facing real-world scenarios such as strong reflection interference in complex water areas. To address the aforementioned technical problems, the present invention discloses the following technical solutions: One aspect of the present invention provides a flare-removing water surface target segmentation method based on polarization feature fusion, comprising: Simultaneously acquire polarization images of water surface targets at multiple preset polarization angles; Based on the polarization image, generate a DOLP polarization degree image, an AOP polarization angle image, and a PPR parallel polarization radiation image; The AOP image and PPR image are input into a preset GFNet enhancement network to obtain a PPRen enhanced parallel polarization image. The GFNet enhancement network decomposes the image into reflection and illumination components and outputs an enhanced image after removing reflections. The PPRen image and the DOLP image are input into a preset CDDFuse fusion network to obtain I. fused The fused image uses a high-low frequency decomposition mechanism to preserve high-frequency details and suppress low-frequency background. to I fused The image is semantically segmented to obtain the segmentation results of the water surface target.
[0007] Optionally, generating the DOLP polarization degree image, AOP polarization angle image, and PPR parallel polarization radiation image based on the polarization image includes: For each pixel position corresponding to different polarization images, the following processing is performed: Based on the light intensity value at the pixel location of each polarized image, the Stokes vector at the pixel location is calculated. The Stokes vector consists of the total light intensity, the horizontal-vertical polarized light component, and the diagonal polarized light component. Calculate the degree of polarization DOLP, polarization angle AOP, and parallel polarization radiation PPR at the pixel location based on the Stokes vector; Based on the degree of polarization (DOLP), polarization angle (AOP), and parallel polarization radiation (PPR) at all pixel locations, generate DOLP, AOP, and PPR images respectively.
[0008] Optionally, the method further includes: The GFNet augmentation network is trained end-to-end based on a pre-acquired training dataset and a preset loss function. The GFNet augmentation network consists of a feature encoder, a degraded polarization feature subnetwork, a PPR augmentation feature subnetwork, an AOP feature subnetwork, a fusion subnetwork, and an augmentation decoder. The CDDFuse fusion network is trained based on the output data of the trained GFNet augmented network. The CDDFuse fusion network consists of a Restormer encoder, a low-frequency feature branch, a high-frequency feature branch, and a Restormer decoder.
[0009] Optionally, the step of inputting the AOP image and PPR image into a preset GFNet enhancement network to obtain a PPRen enhanced parallel polarization image includes: High-dimensional features are obtained by fusing PPR and AOP images based on a feature encoder. The high-dimensional features are fed into the degraded polarization feature subnet, the PPR enhancement feature subnet, and the AOP feature subnet, and the degraded polarization features, PPR enhancement features, and AOP features are extracted respectively. The PPR enhancement features and AOP features are fused based on the fusion subnet to obtain a fused image; The fused image is fed into the enhancement decoder to generate the enhanced PPRen image.
[0010] Optionally, the step of inputting the PPRen image and the DOLP image into a preset CDDFuse fusion network to obtain I fused The merged images include: The PPRen and DOLP images are encoded separately using the Restormer encoder to obtain the shallow features of the PPRen and DOLP images. The shallow features of the PPRen image and the shallow features of the DOLP image are respectively passed to the low-frequency feature branch and the high-frequency feature branch to extract the high-frequency and low-frequency features of the PPRen image, as well as the high-frequency and low-frequency features of the DOLP image. The two sets of high-frequency and low-frequency features mentioned above are fed into the Restormer decoder to obtain the fused IF. fused image.
[0011] Optionally, the simultaneous acquisition of polarization images of the water surface target at multiple preset polarization angles includes: Polarization images were acquired simultaneously using a polarization camera at polarization angles of 0°, 45°, 90°, and 135°.
[0012] Optionally, the method further includes: For each pixel location, the Stokes vector at that pixel location is calculated using the following method: in, This represents the total light intensity value. The horizontal-vertical polarized light component; The diagonal polarized light component; This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 0°. This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 45°. This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 90°. This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 135°.
[0013] Optionally, the step of calculating the degree of polarization (DOLP), polarization angle (AOP), and parallel polarization radiation (PPR) at the pixel location based on the Stokes vector includes: For each pixel location, DOLP, AOP, and PPR at that pixel location are calculated using the following formula: .
[0014] Optionally, the I fused The image undergoes semantic segmentation to obtain the segmentation results of the water surface target, including: Will I fused The image is input into the KNet segmentation network or the FastSCNN segmentation network, and the obtained water surface target segmentation mask is used as the segmentation result.
[0015] Another aspect of the present invention provides a flare-free water surface target segmentation system based on polarization feature fusion, which applies the flare-free water surface target segmentation method based on polarization feature fusion provided in the foregoing aspect.
[0016] This invention discloses a flare-removing water surface target segmentation method and system based on polarization feature fusion. First, a segmentation framework based on deep fusion of multiple polarization features is constructed, jointly utilizing multiple physical features acquired through polarization imaging, including degree of polarization (DOLP), angle of polarization (AOP), and parallel polarization radiation (PPR). These polarization features are fused using a deep learning model, effectively enhancing the separability between the target and the background. Particularly under complex conditions such as strong glare and water wave disturbance, it significantly improves the detection and segmentation capabilities for small targets such as unmanned underwater vehicles.
[0017] Secondly, a GFNet network based on Retinex theory for glare suppression and feature enhancement is constructed. This network encodes, separates, and fuses the input PPR and AOP images, extracting the target reflection component and glare illumination component, thereby effectively suppressing the bright areas in the image caused by sunlight reflection. Simultaneously, GFNet enhances the target contour and texture details at the feature level, and the output enhanced image (PPRen) provides a high-quality input foundation for subsequent image fusion and target segmentation.
[0018] Finally, a correlation-driven feature decomposition strategy and a two-stage training mechanism are proposed. The CDDFuse fusion network performs shallow encoding, frequency domain decomposition, and deep fusion on the enhanced image output by GFNet and the DOLP image, extracting high-frequency details and low-frequency background features respectively, thus effectively distinguishing the target from the background. The entire system adopts a two-stage training process: the first stage independently trains GFNet to optimize glare suppression and target feature enhancement; the second stage trains the CDDFuse network based on the GFNet output to achieve the optimal configuration of fusion and feature decomposition. This step-by-step optimization strategy allows each module to fully utilize its strengths, ultimately achieving end-to-end collaborative optimization and significantly improving the overall system's segmentation accuracy and robustness.
[0019] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify essential or necessary features of this disclosure, nor is it intended to limit the scope of this disclosure. Attached Figure Description
[0020] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.
[0021] Figure 1 This is a flowchart illustrating a flare removal method for water surface target segmentation based on polarization feature fusion, provided in an embodiment of the present invention. Figure 2 An implementation provided by an embodiment of the present invention Figure 1 A flowchart illustrating step S200; Figure 3 A schematic diagram of a GFNet enhanced network provided in an embodiment of the present invention; Figure 4 A schematic diagram of the framework of a CDDFuse fusion network provided in an embodiment of the present invention; Figure 5 An implementation provided by an embodiment of the present invention Figure 1 A flowchart illustrating step S300; Figure 6 An implementation provided by an embodiment of the present invention Figure 1 A flowchart of step S400. Detailed Implementation
[0022] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0023] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0024] In water environments with strong sunlight, glare and solar flares often make targets difficult to identify, especially when shooting on the water surface. Strong light reflection can submerge targets, making the outlines of small targets such as unmanned underwater vehicles (UUVs) blurry or even completely invisible, which seriously affects the subsequent automatic detection and image segmentation results.
[0025] Traditional water surface imaging and segmentation methods typically only acquire single-type information such as light intensity. When there are complex environmental interferences such as water waves, floating objects, and sunlight reflection, the target and background are easily confused, leading to a decrease in target recognition accuracy. This single perception method is inadequate in complex natural scenes and cannot meet the actual needs of high-precision recognition.
[0026] Although current polarization imaging techniques can provide more physical-level image features, existing image fusion algorithms have not been specifically optimized for strong glare interference in water surface environments. Their fusion effects are limited, making it difficult to effectively preserve target information in bright areas, resulting in target areas being easily obscured by strong reflections. Therefore, this invention discloses a glare-reducing water surface target segmentation method and system based on polarization feature fusion to suppress glare effects and significantly improve the accuracy of small target segmentation on the water surface.
[0027] Figure 1 This is a flowchart illustrating a flare removal method for water surface target segmentation based on polarization feature fusion, as provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: Step S100: Simultaneously acquire polarization images of the water surface target at multiple preset polarization angles.
[0028] In one embodiment of this invention, to obtain comprehensive polarization information of a water surface target at different polarization angles, a high-performance polarization camera (such as a SALSA camera) is used to simultaneously acquire data on the target area. This type of camera has four-channel synchronous imaging capability, enabling it to acquire polarized light intensity images with polarization angles of 0°, 45°, 90°, and 135° at the same time point, effectively avoiding timing deviations caused by target movement or water surface fluctuations. These polarization images provide rich physical information, facilitating subsequent analysis and utilization of the reflection differences between the target and the background.
[0029] In this embodiment of the invention, the acquired polarization image resolution can reach 1024×1024 pixels, ensuring the integrity of image details and spatial resolution; the frame rate can be stably maintained at 15 frames / second, meeting the real-time requirements in dynamic water surface scenes.
[0030] The acquisition polarization angles and frequencies listed above are merely illustrative examples. In practical applications, other suitable acquisition polarization angles and frequencies can be selected, which will not be elaborated here.
[0031] Step S200: Generate a DOLP polarization degree image, an AOP polarization angle image, and a PPR parallel polarization radiation image based on the polarization image.
[0032] After a high-performance polarization camera simultaneously acquires four polarization images at four directions (0°, 45°, 90°, and 135°), these polarization images have the same image resolution and pixel arrangement. For example, for a 1024×1024 pixel image, each pixel position (x, y) contains light intensity values at four different polarization angles.
[0033] In one embodiment of the present invention, step S200 can be implemented in the following manner: For each pixel position corresponding to different polarization images, such as Figure 2 As shown, the following processing is performed: Step S201: Calculate the Stokes vector at the pixel location based on the light intensity value at the pixel location of each polarization image.
[0034] The Stokes vector is the standard form for describing the polarization state of light. It can fully express the physical properties of linearly polarized light and consists of the total light intensity, the horizontal-vertical polarized light components, and the diagonal polarized light components.
[0035] Using each pixel location as a processing unit, the light intensity value of that pixel location in each polarization image is jointly analyzed. Taking polarization angles of 0°, 45°, 90°, and 135° as examples, the Stokes vector at any pixel location is calculated using the following method: in, This is the total light intensity value, reflecting overall brightness information; The horizontal-vertical polarization component represents the difference between the horizontal and vertical polarization components. The diagonal polarized light component reflects the intensity contrast of diagonally polarized light. This represents the light intensity value at this pixel location in a polarized image with a polarization angle of 0°. This represents the light intensity value at this pixel location in a polarized image with a polarization angle of 45°. This represents the light intensity value at this pixel location in a polarized image with a polarization angle of 90°. This represents the light intensity value at this pixel location in a polarized image with a polarization angle of 135°.
[0036] Step S202: Calculate the degree of polarization DOLP, polarization angle AOP, and parallel polarization radiation PPR at the pixel position based on the Stokes vector.
[0037] Based on the obtained Stokes vector, the three types of polarization feature quantities at each pixel location are further calculated.
[0038] The first category is the degree of polarization (DOLP), which describes and measures the degree of linear polarization of light. The numerical range is [0,1], where 0 represents completely unpolarized light and 1 represents completely polarized light. DOLP can reflect surface roughness, unevenness, and material properties, and can distinguish the texture features of water surfaces from those of targets.
[0039] The DOLP at each pixel location is calculated using the following formula: .
[0040] The second type is the polarization angle (AOP), which is the principal direction of the photoelectric vector, measured in degrees (°), and ranging from 0° to 180°. Objects of different materials and surface orientations exhibit characteristic responses to AOP, which is helpful for target contour extraction and segmentation.
[0041] The AOP at each pixel location is calculated using the following formula: .
[0042] The third type is parallel polarized radiation (PPR), which is mainly used to characterize the reflectivity of an object in the direction of illumination. It combines the total light intensity and the horizontal polarization component. At a certain solar zenith angle (e.g., >35°), PPR has a strong ability to resist solar flares on water surfaces, can highlight target signals, and suppress strong reflections.
[0043] The PPR at each pixel location is calculated using the following formula: .
[0044] Step S203: Based on the degree of polarization DOLP, polarization angle AOP, and parallel polarization radiation PPR at all pixel locations, generate DOLP image, AOP image, and PPR image respectively.
[0045] The calculated DOLP, AOP, and PPR values at each pixel location are reconstructed according to their two-dimensional spatial positions, forming three feature images of the same size as the original polarization image: the DOLP image, the AOP image, and the PPR image. These images not only preserve the spatial structure information of the original image but also introduce rich physical semantic features. Specifically, the DOLP image highlights regions with strong light polarization, helping to suppress reflective backgrounds; the AOP image enhances the differences between structural boundaries and object surface properties; and the PPR image preserves the brightness contours of the target in highly reflective scenes.
[0046] In one embodiment of the present invention, the GFNet enhanced network and the CDDFuse fusion network are trained in advance using a phased and progressively optimized approach.
[0047] (a) GFNet Enhanced Network First, based on a pre-collected and labeled polarization image training dataset, and combined with a pre-designed multivariate loss function, the GFNet augmentation network is trained end-to-end. The SOED-P high-resolution polarization water surface target dataset can be used for training, containing real-world polarization images under various lighting, weather, and interference conditions.
[0048] The GFNet enhancement network jointly encodes PPR and AOP images, combines Retinex theory to separate the reflection component (target information) and illumination component (abnormally bright glare), and enhances the target polarization features, providing high-quality input for subsequent segmentation.
[0049] The Retinex decomposition theory can be understood using the following formula: Where I is the input image (such as a PPR or AOP image); R is the reflection component, which reflects the inherent properties of the object (texture, color, material) and is independent of illumination; L is the illumination component, which reflects the intensity and distribution of light, mainly including ambient brightness and abnormal highlights (such as solar glare).
[0050] The GFNet enhanced network, with the help of Retinex decomposition, can separate R (target body information) and suppress glare noise in L.
[0051] The GFNet enhancement network disclosed in this embodiment of the invention consists of multiple functional modules, including: a feature encoder, a degraded polarization feature subnet, a PPR enhancement feature subnet, an AOP feature subnet, a fusion subnet, and an enhancement decoder.
[0052] The feature encoder employs a multi-layer convolutional structure (e.g., 4 layers of 3×3 convolutions with LeakyReLU activation) to perform deep feature extraction on the input PPR and AOP images. The degraded polarization feature sub-network is mainly used to model and separate illumination degradation information from interference sources such as water surface reflections, including glare areas. The PPR enhancement feature sub-network enhances the details and structure of the target region in the PPR image, improving target separability. The AOP feature sub-network extracts directional and material information, enhancing the ability to recognize target contours and edges. After processing the three branches, the features are integrated through a fusion sub-network (composed of 3 convolutions) to extract key fusion features. These features are then reconstructed into a single-channel PPR enhanced image (PPRen) by the enhancement decoder (composed of 4 layers of 3×3 convolutions, with the last layer using Tanh activation). This image exhibits higher target contrast and background robustness.
[0053] The following explains the feature encoding and enhancement process of the GFNet enhanced network: 1. The following methods are used to extract mixed features: in, Φ is the encoder, which fuses information from the PPR and AOP domains; Φ is the high-dimensional feature after fusion.
[0054] 2. The following method is used to complete the output of the three-branch feature subnet: in, For the degraded polarization feature subnet, extract the illumination-related components; Enhance the feature subnet of PPR to improve target information; For AOP feature subnets, supplement material and structural information; These are the degraded polarization characteristics, enhanced PPR characteristics, and polarization angle AOP characteristics of the corresponding subnet outputs, respectively.
[0055] 3. Feature fusion and decoding are performed using the following methods: in, To merge subnets, enhanced PPR features are combined with AOP features; For the decoder, restore to the enhanced version image, This is the final anti-glare enhanced feature map output. During the fusion process, the Retinex decomposition mechanism separates the illumination / reflection components in the feature space, suppresses abnormal highlights, and outputs a glare-free and enhanced target feature map.
[0056] Figure 3 This is a schematic diagram of the GFNet enhancement network, which aims to suppress solar flares on the water surface through parallel polarization equivalent radiation (PPR) and fuse target material information contained in AOP images.
[0057] The trained GFNet enhancement network can separate the reflection component (target intrinsic information) and the illumination component (glare interference) in the feature space, and suppress abnormally bright areas. Among them, the three-branch feature subnetwork (degraded polarization feature subnetwork, PPR enhancement feature subnetwork, and AOP feature subnetwork) are trained in an end-to-end manner to enhance the target separability and anti-interference ability of PPRen.
[0058] The loss functions used in the GFNet augmented network include reconstruction loss, illumination smoothing, mutual information constraint, and perceptual loss, with weights that can be optimized experimentally. The following are the loss functions for the GFNet augmented network. One specific implementation, but other loss functions can also be used in practical applications: in, The reconstruction loss of the PPRen image is used to constrain the output image to be consistent with the target. This is a degraded lighting smoothing loss used to improve the model's ability to decompose anomalous lighting. Mutual information constraints are used to ensure that the feature fusion information is maximized and free of redundancy. The perceptual loss is used to improve the subjective visual quality of the output image; λ1, λ2, λ3, and λ4 are the weights of each loss term, used to control the training balance.
[0059] (ii) CDDFuse converged network After training the GFNet augmentation network, the trained GFNet augmentation network is used to infer the original training dataset, generating augmented images (PPRen) in batches, which serve as one of the training inputs for the CDDFuse fusion network. During the CDDFuse network training phase, PPRen images and the original DOLP images are used as input data. Cross-modal fusion is performed on the PPRen and DOLP images, where PPRen images are the images with removed and prominent target features, and DOLP images supplement information such as object details, texture, and roughness. A high- and low-frequency decomposition method is used to capture global background and local details respectively, and correlation loss is used to guide information complementarity, improving the ability to distinguish between the target and the background.
[0060] In the embodiments disclosed in this invention, the CDDFuse fusion network mainly consists of four key modules: a Restormer encoder, a low-frequency feature branch, a high-frequency feature branch, and a Restormer decoder.
[0061] The Restormer encoder, based on a Transformer architecture, utilizes window-based attention to model features of the input image, preserving long-range dependencies. After encoding, it branches into two parallel sub-branches: a low-frequency feature branch focuses on fusing background structure and illumination information to suppress redundant interference; and a high-frequency feature branch focuses on preserving edge, texture, and detail information, aiding in the recovery and enhancement of the target contour. Finally, the fused features are reconstructed by the Restormer decoder, outputting a fused image (I1) with clear details, balanced illumination, and a prominent target. fused ).
[0062] The loss function of the CDDFuse fusion network includes reconstruction loss, decomposition loss, global intensity loss, and texture gradient loss, etc. The specific weights can be fine-tuned experimentally. The following is a specific example of the CDDFuse fusion network loss function; other loss functions can also be used in practical applications: CDDFuse employs a two-stage loss function mechanism in its fusion network training, which is explained below: Stage I: Source Image Reconstruction and Decomposition Stage In the first stage, the loss function The network is guided to retain key source image information and initially establish the ability to decouple high- and low-frequency features. At this stage, CDDFuse has not yet output the final fused image. Instead, it uses the PPRen image output by GFNet and the input DOLP image to guide the network to learn to retain the content of key source images and establish the ability to decompose in the frequency domain, so that the encoder can effectively distinguish the structural levels in the image.
[0063] Stage II: Image Fusion Optimization Stage In the second stage, the loss function Optimize the final output fused image (I fused This process ensures the network possesses brightness consistency, detail preservation, and frequency domain separation. At this stage, the network has acquired preliminary extraction and fusion capabilities, and the focus shifts to optimizing the output I... fused Improving image quality enhances its segmentation friendliness, perceptible detail representation, and global brightness consistency, thereby providing optimal input for subsequent object detection and segmentation networks (such as KNet or FastSCNN).
[0064] in, , To compensate for the reconstruction loss of PPRen and DOLP, the source information is preserved; The feature decomposition loss is used to promote the effective separation of high- and low-frequency information; To minimize overall intensity loss, the brightness of the fused images is constrained to be consistent. Gradient loss is used to enhance texture and edge details; α1, α2, α3, and α4 are weight coefficients, respectively.
[0065] The following explains the mathematical modeling and physical significance of the CDDFuse fusion network. 1. Shallow Feature Extraction (Restormer Encoder): ResEnc( () is a Restormer encoder used to enhance the parallel polarization map of the input. Shallow feature extraction is performed using the polarization degree map DOLP. Restormer is an image restoration model based on the Transformer structure, suitable for modeling long-distance dependencies and cross-modal feature coupling, and can simultaneously extract local texture and global structural information.
[0066] Input data: PPRen image output by GFNet and original DOLP image.
[0067] Output data: shallow feature maps are obtained respectively. and This is used for subsequent frequency domain decomposition.
[0068] This stage extracts the most basic geometric contours, illumination boundaries, and polarization features from the original image, which form the basis for subsequent frequency domain modeling.
[0069] 2. High- and low-frequency decomposition (Transformer-CNN structure): This stage performs structural decoupling on shallow features, dividing them into two types of subdomain features: (1) Low-frequency feature extraction: BaseEnc Belonging to the Transformer branch, it captures low-frequency, global structural information, such as background water, distant areas, and parts with gradually changing lighting.
[0070] The output data is a low-frequency feature map. and .
[0071] (2) High-frequency feature extraction: DetailEnc It belongs to the Invertible Neural Network branch and captures high-frequency details, such as object edges and textures.
[0072] The output data is a high-frequency feature map. and .
[0073] 3. Related driving losses: This loss encourages correlation (similarity) between low-frequency features to ensure consistent background information; at the same time, high-frequency features remain complementary to improve the ability to distinguish object boundaries.
[0074] 4. Final fusion output: in,[ ] represents feature concatenation; ResDec is the Restormer decoder.
[0075] After all subdomain features have been modeled, high and low frequency information is concatenated at the channel level and input into the Restormer decoder (ResDec) to generate the final fused image I. fused .
[0076] In this process, four feature maps are stacked along the channel dimension to form a feature representation containing multi-scale and multi-modal information. Leveraging Restormer's global modeling capabilities, a clear image with balanced texture and contrast is restored. The fused image not only suppresses glare and reflection interference but also preserves edge and target information, providing a clear background and well-defined boundaries for target segmentation.
[0077] Figure 4 This is a schematic diagram of the CDDFuse fusion network framework, which separates the water surface (low-frequency signal) from the moving unmanned underwater vehicle (UUV) target (high-frequency signal) based on correlation-driven loss.
[0078] Throughout the training process, the CDDFuse fusion network is trained end-to-end using multi-objective optimization strategies such as perceptual loss, structural similarity constraints (e.g., SSIM loss), and high-low frequency consistency loss. This allows the output image to retain realism while achieving higher segmentation performance, providing high-quality input images for subsequent segmentation networks.
[0079] Step S300: Input the AOP image and PPR image into the preset GFNet enhancement network to obtain the PPRen enhanced parallel polarization image.
[0080] The GFNet enhancement network can decompose an image into reflection and illumination components and output an enhanced image after removing reflections.
[0081] In one embodiment of the present invention, such as Figure 5 As shown, step S300 can be implemented in the following manner: Step S301: Based on the feature encoder, fuse the PPR image and the AOP image to obtain high-dimensional features.
[0082] The GFNet augmentation network first receives registered PPR and AOP images as input. The feature encoder module fuses these two images to extract local texture information and global semantic features. The PPR image provides reflection intensity information, and the AOP image provides directional structure information; the combination of these two images generates a set of high-dimensional feature maps.
[0083] Step S302: Input the high-dimensional features into the degraded polarization feature subnet, the PPR enhancement feature subnet, and the AOP feature subnet, and extract the degraded polarization features, PPR enhancement features, and AOP features respectively.
[0084] After feature encoding, the resulting high-dimensional features are input into three sub-networks with different functions: a degenerate polarization feature sub-network, a PPR enhancement feature sub-network, and an AOP feature sub-network. The degenerate polarization feature sub-network can identify abnormal regions formed by strong light reflection from the water surface, such as glare and sunspots; the PPR enhancement feature sub-network focuses on enhancing the representation of real structures such as object edges and reflective areas, improving the perception of image details; the AOP feature sub-network further extracts directional information from the target's geometric structure, making the network more sensitive to the target's boundaries and orientation. The three sub-networks work together to separate the intrinsic target features from interference components in the image.
[0085] Step S303: Based on the fusion subnet, fuse the PPR enhancement features and AOP features to obtain the fused image.
[0086] The PPR enhancement features and AOP features obtained from the three subnets are fed into the fusion subnet, which not only preserves the clarity of the target structure but also takes into account the ability to suppress complex lighting backgrounds. The fused features have stronger discriminative power and can accurately represent the target's structure and edge information even under strong reflective interference.
[0087] Step S304: Pass the fused image into the enhancement decoder to generate the enhanced PPRen image.
[0088] The intermediate feature maps output from the fusion subnet are fed into the enhancement decoder to generate the final enhanced image, PPRen. The decoder's role is to restore the spatial structure and original resolution of the image while retaining the high-quality features after enhancement and noise reduction. This results in an output image that is significantly superior to the original PPR image in terms of detail, edge visibility, and target visibility, making it suitable for subsequent fusion and segmentation tasks.
[0089] Step S400: Input the PPRen image and the DOLP image into the preset CDDFuse fusion network to obtain I fused Image fusion.
[0090] The CDDFuse fusion network employs a high-low frequency decomposition mechanism to preserve high-frequency details and suppress low-frequency background.
[0091] In one embodiment of the present invention, such as Figure 6 As shown, step S400 can be implemented in the following manner: Step S401: Encode the PPRen image and DOLP image respectively based on the Restormer encoder to obtain the shallow features of the PPRen image and the shallow features of the DOLP image.
[0092] The CDDFuse fusion network first receives a PPRen image enhanced by a GFNet network and a DOLP polarization image as input, which are then fed into a Restormer encoder for shallow feature extraction. Restormer is an image restoration network based on an improved Transformer, whose encoder possesses powerful global modeling capabilities, fully extracting long-range dependencies and multi-scale structures from the input image. In this process, the PPRen image provides rich texture and structural information, while the DOLP image contains polarization features and illumination consistency information. After encoding, both output their corresponding shallow feature representations.
[0093] Step S402: Input the shallow features of the PPRen image and the shallow features of the DOLP image into the low-frequency feature branch and the high-frequency feature branch respectively, and extract the high-frequency and low-frequency features of the PPRen image, as well as the high-frequency and low-frequency features of the DOLP image.
[0094] The shallow features output from step S401 are then input into the high-frequency feature branch and the low-frequency feature branch for further processing. In the low-frequency branch, a Transformer module (BaseEnc) is used to globally model the shallow features, extracting low-frequency information such as background, water, and overall illumination, maintaining the brightness consistency and structural integrity of the fused image. In the high-frequency branch, a reversible neural network (INN) structure (DetailEnc) is used to model the shallow features, thereby extracting high-frequency information such as edges, textures, and details. This branching structure can effectively achieve frequency domain decoupling of polarization image information, improve feature representation capabilities, and avoid boundary blurring or texture loss during the fusion process.
[0095] Step S403: Input the two sets of high-frequency and low-frequency features above into the Restormer decoder to obtain the fused IF. fused image.
[0096] High-frequency and low-frequency features from two input images (PPRen and DOLP) are concatenated and fed into the Restormer decoder to reconstruct the fused image. The Restormer decoder restores the original resolution and enhances image contrast, detail, and target visibility by integrating multi-scale features and performing deconvolution operations. The concatenation operation ensures that the fused result not only performs well in terms of global illumination and structural consistency but also preserves important visual cues such as detailed features and edge textures. The final output image is... fused The image combines the advantages of both modalities, possessing strong glare suppression and high target discrimination capabilities, making it suitable for subsequent automatic target detection and segmentation tasks.
[0097] Step S500: For I fusedThe image is semantically segmented to obtain the segmentation results of the water surface target.
[0098] In one embodiment of the present invention, I can be used fused The image is input into a KNet or FastSCNN segmentation network, and the obtained water surface target segmentation mask is used as the segmentation result. Of course, I can also be used. fused The image can be input into other segmentation networks; no restrictions are imposed here.
[0099] In one embodiment of the present invention, image I is fused. fused After glare suppression and detail enhancement, it can be used as input to subsequent segmentation networks to achieve accurate detection and segmentation of small targets on the water surface (such as unmanned underwater vehicles, buoys, and floating objects). Specifically, mainstream semantic segmentation network models, such as KNet (Kernel-based Network) or FastSCNN (Fast Semantic Segmentation Convolutional Neural Network), can be used for end-to-end training and prediction. KNet is an instance segmentation network based on dynamic kernel generation, which has the ability to model complex target morphologies and is suitable for accuracy-priority scenarios. Its dynamic convolutional kernel mechanism improves the segmentation ability of small targets. FastSCNN, on the other hand, is a lightweight and fast semantic segmentation network, suitable for real-time and embedded applications. Its lightweight structure makes it easy to deploy.
[0100] Finally, after inference by the segmentation network, a pixel-level segmentation mask of the water surface target is output, in which each pixel is labeled as either "target" or "background".
[0101] To objectively evaluate the performance of segmentation networks, standard semantic segmentation evaluation metrics can be used, including Intersection over Union (IoU) and the Dice coefficient. IoU measures the degree of overlap between the predicted and ground truth regions; a higher value indicates more accurate prediction. The Dice coefficient is another metric that measures the similarity between two sets, particularly suitable for evaluating the accuracy of small targets or target edges. Statistical analysis of the IoU and Dice between the predicted mask and the manually labeled ground truth mask allows for a comprehensive evaluation of the contribution of the fusion enhancement mechanism to improving target separability and guides the optimization of network structure and parameters.
[0102] The method provided in this embodiment of the invention can stably achieve high-precision automatic segmentation of small targets (such as UUVs) on the water surface under complex lighting and strong glare interference. The fused image outperforms the comparative methods in terms of structure, texture, and visual fidelity, with an inference speed of <0.08 seconds per frame, making it suitable for practical engineering deployment.
[0103] It has the following advantages: 1. Significantly improves the segmentation accuracy of small targets on the water surface, with strong robustness. By deeply fusing three polarization features—degree of polarization (DOLP), angle of polarization (AOP), and parallel polarization radiation (PPR)—this algorithm fully utilizes the complementarity of each parameter, enabling accurate segmentation of small, slow-moving targets (such as UUVs) on the water surface even in complex environments (such as strong solar glare, waves, and raindrop interference). Experimental results show that, when combined with mainstream segmentation networks (such as KNet and FastSCNN), the average intersection-over-union ratio (mIoU) reaches 86.39 and 70.47, respectively, with significantly higher segmentation accuracy than algorithms using only a single polarization feature or traditional multimodal fusion algorithms.
[0104] 2. Effectively suppresses sunlight glare and improves target visibility. By designing a glare-reducing enhancement network (GFNet) and introducing a Retinex decomposition mechanism at the feature level, abnormally bright areas caused by strong reflections from the water surface can be eliminated to the greatest extent possible without losing the intrinsic information of the target. Compared with existing technologies, this invention can better restore the true target boundaries and details obscured by glare, effectively solving the problem of significant decrease in segmentation accuracy under strong lighting interference in traditional methods.
[0105] 3. Multi-branch related-driven integration, balancing details and consistency. By introducing a correlation-driven feature decomposition and fusion network (CDDFuse), which separates high- and low-frequency features and constrains correlation loss, this invention ensures both the full preservation of high-frequency details such as target boundaries and textures, and the structural consistency of the background region. Compared with traditional simple image fusion methods, this invention outperforms traditional methods in terms of target highlighting and background suppression, and significantly improves quantitative indicators such as structural similarity, information entropy, and texture gradient of the fused image.
[0106] 4. End-to-end two-stage optimization to adapt to complex scenarios The entire process employs a two-stage deep learning training approach: first optimizing glare reduction and enhancement, then optimizing feature fusion, ensuring targeted improvements at each stage, and finally coordinating end-to-end optimization of overall performance. This invention not only possesses excellent generalization capabilities but also demonstrates superior adaptability and stability under various time periods, weather conditions, and water surface disturbances.
[0107] 5. Highly efficient in computation and easy to deploy in practice. With its simple structure and fast inference speed, it achieves a single-frame processing time of less than 0.08 seconds on a standard GPU and can also meet near real-time requirements on a CPU, significantly outperforming some traditional and deep learning fusion methods. It also has low memory overhead, facilitating practical deployment in embedded systems or marine monitoring terminals.
[0108] 6. Possesses broad application prospects It can be widely used in surface / underwater target monitoring, marine security, and aquatic unmanned system perception, and is especially suitable for practical scenarios with extremely high requirements for environmental robustness and segmentation accuracy, providing key technical support for improving the automation level of marine intelligent monitoring and unmanned operations.
[0109] This invention also discloses a flare-removing water surface target segmentation system based on polarization feature fusion. This system applies the flare-removing water surface target segmentation method based on polarization feature fusion disclosed in the foregoing embodiments, which will not be described again here.
[0110] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A flare-free water surface target segmentation method based on polarization feature fusion, characterized in that, include: Simultaneously acquire polarization images of water surface targets at multiple preset polarization angles; Based on the polarization image, generate a DOLP polarization degree image, an AOP polarization angle image, and a PPR parallel polarization radiation image; The AOP image and PPR image are input into a preset GFNet enhancement network to obtain a PPRen enhanced parallel polarization image. The GFNet enhancement network decomposes the image into reflection and illumination components and outputs an enhanced image after removing reflections. The PPRen image and the DOLP image are input into a preset CDDFuse fusion network to obtain I. fused The fused image uses a high-low frequency decomposition mechanism to preserve high-frequency details and suppress low-frequency background. to I fused The image is semantically segmented to obtain the segmentation results of the water surface target.
2. The method according to claim 1, characterized in that, The generation of the DOLP polarization degree image, AOP polarization angle image, and PPR parallel polarization radiation image based on the polarization image includes: For each pixel position corresponding to different polarization images, the following processing is performed: Based on the light intensity value at the pixel location of each polarized image, the Stokes vector at the pixel location is calculated. The Stokes vector consists of the total light intensity, the horizontal-vertical polarized light component, and the diagonal polarized light component. Calculate the degree of polarization DOLP, polarization angle AOP, and parallel polarization radiation PPR at the pixel location based on the Stokes vector; Based on the degree of polarization (DOLP), polarization angle (AOP), and parallel polarization radiation (PPR) at all pixel locations, generate DOLP, AOP, and PPR images respectively.
3. The method according to claim 1, characterized in that, The method further includes: The GFNet augmentation network is trained end-to-end based on a pre-acquired training dataset and a preset loss function. The GFNet augmentation network consists of a feature encoder, a degraded polarization feature subnetwork, a PPR augmentation feature subnetwork, an AOP feature subnetwork, a fusion subnetwork, and an augmentation decoder. The CDDFuse fusion network is trained based on the output data of the trained GFNet augmented network. The CDDFuse fusion network consists of a Restormer encoder, a low-frequency feature branch, a high-frequency feature branch, and a Restormer decoder.
4. The method according to claim 3, characterized in that, The step of inputting the AOP image and PPR image into a preset GFNet enhancement network to obtain a PPRen enhanced parallel polarization image includes: High-dimensional features are obtained by fusing PPR and AOP images based on a feature encoder. The high-dimensional features are fed into the degraded polarization feature subnet, the PPR enhancement feature subnet, and the AOP feature subnet, and the degraded polarization features, PPR enhancement features, and AOP features are extracted respectively. The PPR enhancement features and AOP features are fused based on the fusion subnet to obtain a fused image; The fused image is fed into the enhancement decoder to generate the enhanced PPRen image.
5. The method according to claim 3, characterized in that, The process involves inputting the PPRen image and the DOLP image into a preset CDDFuse fusion network to obtain I. fused The merged images include: The PPRen and DOLP images are encoded separately using the Restormer encoder to obtain the shallow features of the PPRen and DOLP images. The shallow features of the PPRen image and the shallow features of the DOLP image are respectively passed to the low-frequency feature branch and the high-frequency feature branch to extract the high-frequency and low-frequency features of the PPRen image, as well as the high-frequency and low-frequency features of the DOLP image. The two sets of high-frequency and low-frequency features mentioned above are fed into the Restormer decoder to obtain the fused IF. fused image.
6. The method according to claim 2, characterized in that, The simultaneous acquisition of polarization images of water surface targets at multiple preset polarization angles includes: Polarization images were acquired simultaneously using a polarization camera at polarization angles of 0°, 45°, 90°, and 135°.
7. The method according to claim 6, characterized in that, The method further includes: For each pixel location, the Stokes vector at that pixel location is calculated using the following method: in, This represents the total light intensity value. The horizontal-vertical polarized light component; The diagonal polarized light component; This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 0°. This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 45°. This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 90°. This refers to the light intensity value at the pixel location in a polarized image with a polarization angle of 135°.
8. The method according to claim 7, characterized in that, The calculation of the degree of polarization (DOLP), polarization angle (AOP), and parallel polarization radiation (PPR) at the pixel location based on the Stokes vector includes: For each pixel location, DOLP, AOP, and PPR at that pixel location are calculated using the following formula: 。 9. The method according to claim 1, characterized in that, The above to I fused The image undergoes semantic segmentation to obtain the segmentation results of the water surface target, including: Will I fused The image is input into the KNet segmentation network or the FastSCNN segmentation network, and the obtained water surface target segmentation mask is used as the segmentation result.
10. A flare-free water surface target segmentation system based on polarization feature fusion, characterized in that, The system applies the flare removal water surface target segmentation method based on polarization feature fusion as described in any one of claims 1 to 9.