A Method for Extracting Mariculture Information of Seawater Raft Cultivation Based on an Interpretable Polarization Deep Learning Network
Through an interpretable polarization deep learning network, combined with polarization scattering prior knowledge and deep learning technology, the interpretability and reliability problems of floating raft farming information extraction in marine aquaculture monitoring are solved, and high-precision segmentation and boundary recognition are achieved, which improves the efficiency and reliability of marine aquaculture monitoring.
Patent Information
- Application Number
- CN202510323261.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In the monitoring of marine aquaculture, the explanationability and reliability of the information extraction method for seawater floating raft farming is poor, and segmentation accuracy and boundary identification are difficult in complex marine environments.
The interpretable polarized deep learning network is adopted, combined with polarized scattering prior knowledge and deep learning technology, and the polarized scattering class activation mapping module, interpretable segmentation consistency learning module, and refined class activation and edge guidance network are improved to improve segmentation accuracy and interpretability.
High-precision seawater floating raft farming area segmentation is achieved in complex marine environments, significantly improving the interpretability and boundary consistency of segmentation results, and providing an efficient, transparent and reliable monitoring solution.
Smart Images

Figure CN119851158B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - technical field of marine remote sensing and artificial intelligence, and relates to a method for extracting seawater raft aquaculture information by an interpretable polarization deep learning network. Background Art
[0002] Marine aquaculture is of great significance for global seafood demand, coastal economic support, and alleviating the pressure on wild fish resources. However, unregulated expansion may lead to environmental problems such as harmful algal blooms, eutrophication, and biodiversity degradation. Traditional on - site survey methods are time - consuming and inefficient, while remote sensing technology can provide large - scale observations and timely data. Polarimetric Synthetic Aperture Radar (PolSAR), as a powerful tool, has the ability to operate all - weather, high spatial resolution, and strong penetration characteristics, and can effectively carry out marine target monitoring. In addition, PolSAR can accurately distinguish water bodies from aquaculture structures, such as floating rafts and cages, due to its inherent multi - polarization channels. The polarization scattering characteristics of these aquaculture structures provide valuable prior information for subsequent classification tasks. Therefore, it is crucial to develop effective PolSAR image processing techniques to achieve accurate extraction of seawater aquaculture.
[0003] In the ocean aquaculture monitoring task, due to the complex and variable environment of the aquaculture area, including the influence of various factors such as waves, wind waves, and ocean currents, as well as the interaction between aquaculture facilities and the surrounding environment, the scattering characteristics of radar signals are complex and diverse, thus posing challenges to the extraction performance and reliability of the model. Chinese Invention Patent (Publication No. CN110910494B) proposes a method of simulating the sea surface distribution under different wind speeds, performing three-dimensional modeling in combination with Rhino software, and using the electromagnetic simulation software FEKO to solve the backscattering field to generate an ISAR image. This method is based on the Freeman and Yamaguchi decomposition theory, analyzes the polarization scattering characteristics of the target, distinguishes different scattering mechanisms of the sea surface and underwater parts, such as surface scattering, dihedral angle scattering, etc., and evaluates the results through the equivalent number of looks and Fresnel reflection coefficients, so as to effectively analyze the backscattering characteristics of seawater raft aquaculture under random sea conditions. However, this method only analyzes the backscattering characteristics of seawater raft aquaculture and does not actually apply the backscattering characteristics of seawater raft aquaculture to the deep learning model to constrain the behavior of the network model. The data-driven deep learning model can mine the characteristics of seawater aquaculture targets from the original SAR data. Chinese Invention Patent (Publication No. CN119398540A) is based on the original UNet neural network model, modifies its structure to increase complexity, forms a nested UNet neural network, and adds a channel attention mechanism (SE module) to build the U2-Net_Se2 model, and finally obtains excellent raft segmentation results. However, due to the "black box" characteristic of this model, it cannot feedback important decision-making regions and features to people. Summary of the Invention
[0004] In view of the above problems existing in the prior art, the present invention proposes a method for extracting seawater raft aquaculture information based on an interpretable polarization deep learning network. The present invention can solve the problems of poor interpretability and reliability of the intelligent interpretation method for seawater aquaculture, improve the interpretability during the extraction process while ensuring the extraction accuracy of raft aquaculture, and clearly point out which features play a key role in the final segmentation result. The core innovation of the present invention lies in: proposing an interpretable polarization deep learning network for semantic segmentation of marine aquaculture areas. The interpretable polarization deep learning network can achieve high-precision segmentation and interpretable analysis of aquaculture areas by integrating prior knowledge of polarization scattering and deep learning techniques. Specifically, the interpretable polarization deep learning network includes three key components: a polarization scattering class activation mapping module, an interpretable segmentation consistency learning module, and a refined class activation and edge guidance network. The polarization scattering class activation mapping module generates high-quality class activation mapping diagrams by combining the polarization scattering characteristics of full-polarization SAR images and gradient weights, which can effectively highlight the features related to aquaculture areas and suppress the noise in complex backgrounds. The interpretable segmentation consistency learning module uses the generated class activation mapping diagrams to guide the model to pay more attention to the features of aquaculture areas during training, improve the interpretability of the model, and enhance the model's attention to aquaculture areas. The refined class activation and edge guidance network adopts a dual-branch structure to perform semantic segmentation and edge supervision respectively, which can simultaneously focus on the semantic features and edge details of aquaculture areas and generate accurate segmentation results. The synergistic effect of these three components enables the interpretable polarization deep learning network to achieve high-precision segmentation of aquaculture areas in complex marine environments, and at the same time significantly improves the interpretability of the model, providing an efficient, transparent and reliable solution for marine aquaculture monitoring tasks.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A method for extracting seawater raft aquaculture information based on an interpretable polarization deep learning network, the method for extracting seawater raft aquaculture information includes the following steps:
[0007] In the first step, collect PolSAR image data of remote sensing satellites, preprocess the PolSAR image data, and prepare a data set, specifically as follows:
[0008] Step 1.1, first collect PolSAR image data of remote sensing satellites through satellite resource websites. These image data usually contain echo signals (HH, HV, VH, VV polarizations) in different polarization states. Then preprocess the PolSAR image data.
[0009] Furthermore, the preprocessing refers to using the PolSARproV.6.0.3 software to perform radiometric calibration, geometric calibration, and denoising preprocessing on the image data to ensure the quality and accuracy of the image data.
[0010] Step 1.2: Represent the polarization information using the covariance matrix of the PolSAR image data obtained in Step 1.1. Perform polarization decomposition on the polarization information to extract different types of scattering components from the PolSAR image data, including dihedral scattering, volume scattering, and surface scattering.
[0011] Furthermore, the methods used for the polarization decomposition include the Yamaguchi four-component decomposition method or the Freeman-Durden decomposition method.
[0012] Step 1.3: Combine the dihedral scattering, volume scattering, and surface scattering obtained in Step 1.2 to obtain a PolSAR pseudo-color image, and obtain a seawater aquaculture dataset and labels based on the PolSAR pseudo-color image.
[0013] Furthermore, use the labelme software to manually annotate the PolSAR pseudo-color image to obtain a seawater aquaculture dataset and labels.
[0014] The second step is to design a refined class activation and edge guidance network to solve the problems that the scattering characteristics of the sea surface affect the recognition of aquaculture rafts, and when multiple rafts are very close or the movement caused by wave action leads to the adhesion of raft targets and jagged boundaries in the image. The refined class activation and edge guidance network adopts an encoder-dual decoder architecture (see Step 2.1), including a convolutional backbone network, a semantic decoder, and an edge decoder; the semantic decoder uses a refined class activation attention module to enhance the semantic features of the raft area in the PolSAR pseudo-color image (see Step 2.2); the edge decoder uses an edge-driven feature enhancement module, which focuses on accurately capturing the boundary information of the raft area in the PolSAR pseudo-color image, thereby further optimizing the boundary consistency of the segmentation result (see Step 2.3). These two parts cooperate with each other to generate the final segmentation output through feature fusion and a convolutional segmentation head. Specifically as follows:
[0015] Step 2.1: Design a refined class activation and edge guidance network based on the encoder-dual decoder architecture, including a convolutional backbone network, a semantic decoder, and an edge decoder, and achieve accurate semantic segmentation through feature fusion. The convolutional backbone network uses ResNet50 to extract the feature map of the PolSAR image data, gradually capturing multi-level spatial information. The refined class activation and edge guidance network uses a learnable parameter α to weight and fuse the feature maps of the convolutional backbone network, the semantic decoder, and the edge decoder, and its formula is expressed as:
[0016] F = α·EF+(1 - α)·Upsample(DF)(1);
[0017] Among them, EF represents the features extracted by the convolutional backbone network, DF represents the features obtained by gradually decoding the semantic decoder or the edge decoder, F represents the weighted fusion features, and Upsample(·) represents the upsampling operation.
[0018] After the PolSAR pseudo-color image is input into the refined class activation and edge guidance network, the convolutional backbone network ResNet50 first extracts features from it, and then sends the extracted features into the semantic decoder and the edge decoder respectively. After the features pass through the two decoders, semantic features and edge features will be generated respectively. The two types of features are fused by pixel-level summation. The fused features are further processed in the final convolutional segmentation head to output the final segmentation result. In order to promote the collaborative optimization of the network, the refined class activation and edge guidance network adopts a multi-task loss function including semantic segmentation loss, edge loss and fusion loss.
[0019] Step 2.2, design a refined class activation attention module. The goal is to dynamically adjust the importance of each channel in the feature map by combining local and global context information, so as to improve the focusing ability of the refined class activation and edge guidance network on the raft area and suppress background interference. The core function of the refined class activation attention module is to enhance the feature expression of the target category through the attention mechanism and re-calibrate the channels using global context information. Specifically:
[0020] Step 2.2.1, given the feature map X ∈ R C×H×W , where C, H, and W represent the total number of feature channels, height, and width respectively. First, extract the global context information z through global average pooling. Then, dynamically generate channel weights through two fully connected networks and activation functions, as shown in formula (2):
[0021]
[0022] Among them, W1 and W2 are the weights of the fully connected layers; δ is the ReLU activation function; σ is the Sigmoid activation function; v c is the weight of channel c, used to adjust the importance of the channel; is the calibrated feature map; X c is the feature map before calibration. This process can dynamically adjust the importance of each channel in the feature map, strengthen category-related features, and at the same time suppress irrelevant information.
[0023] Step 2.2.2, in order to improve the category activation ability of the feature map in the spatial dimension, the feature map is divided into N pA number of non - overlapping patches, denoted as X (p) , and the local class features in each patch are calculated by formula (3):
[0024]
[0025] where represents the local class feature of patch p; represents the class weight; σ s represents the spatial softmax, which is used to highlight the significant regions within the patch; represents the class activation map within patch p. Local feature aggregation can enhance the feature expression ability in a small range and is especially suitable for semantic segmentation in complex backgrounds.
[0026] Step 2.2.3, integrate the local features of all patches to generate global class features, as shown in formula (4):
[0027]
[0028] where F gl represents the global class feature; f p is the weight of patch p, which is used to dynamically adjust the contribution of different patches to the global feature.
[0029] Step 2.2.4, optimize the global features obtained in Step 2.2.3 using a set of attention mechanisms to make them more suitable for class discrimination in the segmentation task. The final output of the refined class activation attention module is obtained by calculating the similarity of each pixel to the class, as shown in formula (5):
[0030]
[0031] and combine it with the global features to generate the final enhanced feature map, as shown in formula (6):
[0032]
[0033] where P (p) represents the class similarity of pixel block p; W q , W k , W v are linear transformation matrices for feature alignment, W q represents the learnable linear transformation matrix of the input features, W k represents the learnable linear transformation matrix of the local class features, W v represents the learnable linear transformation matrix of the global class features; σ c represents the class softmax; X (p)Represents a patch of the input feature; Is a patch in the enhanced input feature.
[0034] Step 2.3, design an edge-driven feature enhancement module to address the segmentation problem of complex or weak boundaries, especially when dealing with densely distributed marine aquaculture rafts in PolSAR pseudo-color images. Since edge information plays a crucial role in semantic segmentation, this edge-driven feature enhancement module effectively captures and strengthens boundary features by combining channel attention, spatial attention, and Sobel edge detection, generating more accurate and clear segmentation results, as follows:
[0035] Step 2.3.1, The channel attention mechanism dynamically adjusts the importance of each channel in the input feature map. Specifically, define the input feature map as X ∈ R C×H×W , generate a channel descriptor z ∈ R C through global average pooling, and then generate channel weights u through two fully connected operations, and use the weights for element-wise weighted channel enhancement. The formula is:
[0036]
[0037] where W3 and W4 are both convolutional weights; δ and σ are the ReLU and Sigmoid activation functions respectively; X l represents the l-th sub-feature map in the feature map; u l represents the weight of the l-th sub-feature map; represents the result after weighting the l-th sub-feature map; C represents the total number of channels of the feature. In this way, the channel attention mechanism can highlight the feature expressions of significant channels while suppressing the information of irrelevant channels, and finally combine all the weighted sub-feature maps to obtain the channel attention enhanced feature X channel .
[0038] Step 2.3.2, The spatial attention mechanism enhances the information in the key spatial regions of the feature map by generating a spatial weight map. Specifically, the input feature map first generates two spatial descriptors X max ∈ R H×W and X avg ∈ R H×W through max pooling and average pooling operations in the channel dimension. Subsequently, these two spatial descriptors are concatenated in the channel dimension and a spatial weight map A s is generated through a convolutional operation. Finally, the spatial weight map A s is multiplied element-wise with the input feature map to generate the spatial enhanced feature map X spatial .
[0039] Step 2.3.3, in order to further extract explicit boundary information, this edge-driven feature enhancement module combines the classical Sobel edge detection method to generate preliminary boundary features by calculating the horizontal and vertical gradients of the input feature map X. First, through the Sobel convolution kernels K x and K y calculate the gradients, as shown in Equation (8):
[0040] G x =X*K x , G y =X*K y (8);
[0041] where, G x represents the horizontal direction gradient, and G y represents the vertical direction gradient;
[0042] Then, calculate the gradient magnitude G, as shown in Equation (9):
[0043]
[0044] To further refine the edge features, first normalize G and activate it with the ReLU function. Then, the features are further processed using depthwise separable convolution to obtain the enhanced edge features, denoted as Sobel boundary enhanced feature X sobel .
[0045] Step 2.3.4, the edge-driven feature enhancement module fuses the channel attention enhanced feature X channel , the spatial attention enhanced feature 0X spatial and the Sobel boundary enhanced feature X sobel by element-wise addition to generate the final output feature map, as shown in Equation (10):
[0046] X output =X channel +X spatial +X sobel (10);
[0047] Step 3: Design the polarization scattering class activation mapping module. The design of the polarization scattering class activation mapping module aims to improve the interpretability and target area localization ability of the neural network in the PolSAR pseudo-color image segmentation task. By combining the prior knowledge of polarization scattering with the gradient weight information of deep learning, this polarization scattering class activation mapping module not only generates class activation maps with clear physical meanings, but also directly participates in the training process to guide the convolutional backbone network to extract more representative features, significantly enhancing the segmentation interpretability. Aquaculture areas in PolSAR pseudo-color images usually exhibit significant dihedral scattering characteristics, which are characteristic signals formed by multiple reflections of radar waves between rafts or floats and the water surface. The polarization scattering class activation mapping module utilizes this physical property to optimize the generation of activation maps through dihedral scattering prior information, focusing on the target area and suppressing background interference in complex marine environments. Specifically as follows:
[0048] Step 3.1, the polarization scattering class activation mapping module first generates an initial segmentation mask M for the input PolSAR pseudo-color image through the refined class activation and edge guidance network in Step 2.1, marking potential target areas (such as raft aquaculture areas). At the same time, using the dihedral scattering component D extracted by Freeman decomposition in Step 1.2, combine it with the segmentation mask to generate a corrected mask M D , as shown in Equation (11):
[0049] M D = Me D(11);
[0050] where, e represents the element-wise multiplication operation. This corrected mask M D highlights the dihedral scattering characteristics within the region of interest to strengthen the attention to the target area. Since the resolution of the feature map A k extracted by the convolutional backbone network in Step 2.1 is different from that of the dihedral scattering component D, it is necessary to perform a downsampling operation on M D to obtain a mask M D ' with the same resolution as the feature map, as shown in Equation (12):
[0051] M D ' = Resize(M D ,(H,W))(12);
[0052] where, the Resize(·) function adjusts the mask to the spatial dimensions (H,W) of the feature map.
[0053] Step 3.2, the polarization scattering class activation mapping module measures the correlation between the feature map A k and the corrected mask M D ' by calculating the cosine similarity, as shown in Equation (13):
[0054]
[0055] Among them, v k and are the vectorized representations of the feature map A k and the modified mask M D ′ respectively, and ‖·‖ is the Euclidean norm. The cosine similarity reflects the alignment degree between the feature map and the target region.
[0056] Step 3.3. To further emphasize the importance of the target region, the polarization scattering class activation mapping module calculates the prediction score Y c of category c through global average pooling for the gradient weight of the feature map A k as shown in formula (14):
[0057]
[0058] The importance information of the feature map for the target category is captured through the gradient weight . To balance the contributions of the similarity and the gradient weight, both are normalized respectively, as shown in formula (15):
[0059]
[0060] Among them: s c represents the vector composed of the correlation score weights of all channels, α c represents the vector composed of the gradient weights of all channels, min(·) represents finding the minimum term in the weight vector, and max(·) represents finding the maximum term in the weight vector. The normalized correlation score weight and gradient weight are combined to generate the final channel weight, as shown in formula (16):
[0061]
[0062] Among them, represents the channel weight; γ is a hyperparameter used to adjust the overall influence of the weight. The channel weight is used to weight the feature map to generate the final class activation map, as shown in formula (17):
[0063]
[0064] Among them, the ReLU activation function ensures that only the features with positive contributions to the target region are retained.
[0065] The class activation map generated by the polarization scattering class activation mapping module has significant physical meaning and interpretability, which can clearly highlight the floating raft structure in the aquaculture area and effectively suppress background interference.
[0066] Step 4: Design an interpretable segmentation consistency learning module. To improve the interpretability of the refined class activation and edge guidance network and enhance the consistency of segmentation feature extraction, the present invention designs an interpretable segmentation consistency learning module. The design goal of this interpretable segmentation consistency learning module is to use the high-quality class activation map generated by the polarization scattering class activation mapping module to guide the model learning, making it more focused on the target area (such as the aquaculture raft area) and effectively suppressing the interference of the background area, thereby improving the interpretability of the refined class activation and edge guidance network itself. Specifically as follows:
[0067] Step 4.1: Generate a class activation map CAM through the polarization scattering class activation mapping module to highlight the attention area of the convolutional backbone network of the refined class activation and edge guidance network for the target class. The generation formula of the class activation map is as follows:
[0068]
[0069] where, CAM i (x,y) represents the class activation value of class i at spatial position (x,y); A k (x,y) is the feature map of the kth channel; is the weight of the kth channel corresponding to class i.
[0070] Step 4.2: The feature maps extracted from the convolutional backbone network of the refined class activation and edge guidance network are summed through the channel dimension to generate a comprehensive activation map SAM to comprehensively reflect the feature intensity of the target area, as shown in formula (19):
[0071]
[0072] where, SAM(x,y) represents the sum of all channel feature maps at spatial position (x,y). To ensure that the comprehensive activation map (SAM) is consistent with the class activation map (CAM) in spatial distribution, the interpretable segmentation consistency learning module constrains the two through a consistency loss function. The consistency loss is calculated based on the mean squared error (MSE), as shown in formula (20):
[0073]
[0074] where, SAM′(x,y) is the normalized comprehensive activation map, CAM′ i(x, y) is the normalized class activation map, where H and W represent the height and width of the feature map respectively.
[0075] The introduction of the interpretable segmentation consistency learning module effectively improves the interpretability of the refined class activation and edge guidance network, making the feature extraction of the target area more accurate and suppressing the interference of the background area. In addition, in complex aquaculture scenarios (such as dense raft areas or complex water surface reflections), this interpretable segmentation consistency learning module can significantly enhance the model's focusing ability on the target area, improving the segmentation accuracy and boundary consistency.
[0076] In the fifth step, a multi-optimization objective loss function L is designed to optimize the entire network, aiming to simultaneously improve the segmentation accuracy, edge detail description ability, and model decision interpretability. Through the semantic segmentation loss L seg and the fusion loss L fuse optimize pixel-level classification and accurately distinguish the target from the background using the cross-entropy function. Combine the edge detection loss L of weighted binary cross-entropy and intersection over union loss edge to strengthen the spatial consistency modeling of complex boundaries. At the same time, introduce the interpretable consistency loss L con to constrain the alignment of the convolutional backbone network features and the class activation map, promote the feature to focus on the target object and suppress noise interference, as shown in Equation (21):
[0077] L = L seg + L edge + L fuse + λL con (21);
[0078] where λ is used to control the weight of the interpretability loss. Through end-to-end optimization, finally, while ensuring the extraction accuracy of floating raft aquaculture, the interpretability of the extraction process is improved.
[0079] The beneficial effects of the present invention are:
[0080] (1) When existing PolSAR image segmentation methods are used to handle the task of marine aquaculture monitoring, there are problems such as obvious background interference, difficult boundary recognition, and insufficient interpretability. These methods are difficult to effectively utilize the physical characteristics of polarimetric scattering, resulting in inaccurate target area positioning, blurred boundaries, and low reliability of segmentation results, especially in complex environments where aquaculture rafts are densely distributed. In view of the above problems, the present invention proposes an innovative interpretable polarimetric deep learning network, which significantly improves the segmentation accuracy and interpretability during the segmentation task. First, through the design of the polarimetric scattering class activation mapping module, the prior knowledge of polarimetric scattering is incorporated into the segmentation network, enabling the features extracted by the convolutional backbone network to accurately focus on the aquaculture raft area while effectively suppressing background interference, fundamentally solving the problem of inaccurate target area positioning. Second, the interpretable segmentation consistency learning module guides the network learning through high-quality class activation maps, which better conforms to the characteristics of the target area and significantly reduces the occurrence of background misjudgment. In addition, to solve the problem of difficult boundary recognition, a refined class activation and edge guidance network is proposed to accurately extract target boundary information and ensure clear and detailed segmentation results.
[0081] (2) The present invention not only effectively solves the limitations of the existing technology but also comprehensively improves the segmentation accuracy, boundary refinement ability, and model interpretability in complex marine environments. Its technical framework has strong applicability and provides not only an efficient and reliable solution for the precise monitoring of aquaculture areas but also can be widely applied to target recognition and segmentation tasks in other remote sensing scenarios. Brief Description of the Drawings
[0082] Figure 1 It is a general schematic diagram of the polarimetric interpretable segmentation network. Figure 1 Among them, the first part is the structural diagram of the polarimetric scattering class activation mapping module, the second part is the structural diagram of the interpretable segmentation consistency learning module, and the third part is the structural diagram of the refined class activation and edge guidance network;
[0083] Figure 2 It is a schematic diagram of the refined class activation and edge guidance network;
[0084] Figure 3 It is a schematic diagram of the refined class activation attention module;
[0085] Figure 4 It is a schematic diagram of the edge-driven feature enhancement module; Detailed Description of the Invention
[0086] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0087] As Figure 1As shown, the present invention provides a method for extracting seawater raft aquaculture information that can interpret a polarized deep learning network. This embodiment is compiled under the Windows 11 system using Python 3.8.10, PyTorch 1.11.1, and CUDA 11.3, and runs on a GPU with RTX 2080ti. When preprocessing the PolSAR image, the PolSARpro V.6.0.3 software is used for processing. It includes the following steps:
[0088] First step, collect the PolSAR image data of remote sensing satellites, preprocess the PolSAR image data, and prepare a dataset, specifically as follows:
[0089] Step 1.1, radiometric calibration: Convert the digital image data into actual physical quantities such as reflectivity to eliminate the radiometric bias of the PolSAR image data.
[0090] Step 1.2, image enhancement: Use Lee filtering with a window size of 7×7 to mitigate the impact of speckle noise in the PolSAR image data.
[0091] Step 1.3, polarization decomposition: Calculate the coherence matrix of the PolSAR image data to capture the polarization state and the correlation between channels. Then perform Freeman decomposition on the coherence matrix to separate the radar backscattering into three main scattering mechanisms: dihedral scattering, surface scattering, and volume scattering. Combine these three scattering components to generate a PolSAR pseudo-color image.
[0092] Step 1.4, make a dataset: After the above preprocessing, the PolSAR pseudo-color image is cropped into small images of 512×512 pixels, and a total of 190 small pieces are obtained. Use the labelme software to manually annotate these small PolSAR pseudo-color images. Then these image patches and their ground truth maps are pre-augmented by vertical and horizontal flipping, so that the total number of image patches reaches 570. All image patches are divided into a training set, a validation set, and a test set, with proportions of 6:1:3 respectively.
[0093] Step 2: Design a refined class activation and edge guidance network to address the issues that the scattering characteristics of the sea surface affect the recognition of aquaculture rafts, and when multiple rafts are very close or the movement caused by wave action leads to the adhesion of raft targets and jagged boundaries in the image. The refined class activation and edge guidance network adopts an encoder-dual decoder architecture (see Step 2.1), including a convolutional backbone network, a semantic decoder, and an edge decoder; the semantic decoder uses a refined class activation attention module to enhance the semantic features of the raft area in the PolSAR pseudo-color image (see Step 2.2); the edge decoder uses an edge-driven feature enhancement module, which focuses on accurately capturing the boundary information of the raft area in the PolSAR pseudo-color image, thereby further optimizing the boundary consistency of the segmentation result (see Step 2.3). These two parts cooperate with each other to generate the final segmentation output through feature fusion and a convolutional segmentation head. The specific details are as follows:
[0094] Step 2.1: Design a refined class activation and edge guidance network based on the encoder-dual decoder architecture, including a convolutional backbone network, a semantic decoder, and an edge decoder, to achieve accurate semantic segmentation through feature fusion. The convolutional backbone network uses ResNet50 to extract the feature map of the PolSAR image data, gradually capturing multi-level spatial information. The refined class activation and edge guidance network uses a learnable parameter α to weight and fuse the feature maps of the convolutional backbone network, the semantic decoder, and the edge decoder. Its formula is expressed as:
[0095] F = α·EF+(1-α)·Upsample(DF)(1);
[0096] where EF represents the features extracted by the convolutional backbone network, DF represents the features gradually decoded by the semantic decoder or the edge decoder, F represents the weighted and fused features, and Upsample(·) represents the upsampling operation.
[0097] After the PolSAR pseudo-color image is input into the refined class activation and edge guidance network, the convolutional backbone network ResNet50 first extracts features from it, and then sends the extracted features into the semantic decoder and the edge decoder respectively. After the features pass through the two decoders, semantic features and edge features will be generated respectively. The two types of features are fused through element-wise summation. The fused features are further processed in the final convolutional segmentation head to output the final segmentation result. To promote the collaborative optimization of the network, the refined class activation and edge guidance network adopts a multi-task loss function that includes a semantic segmentation loss, an edge loss, and a fusion loss. As Figure 2 shown in the schematic diagram of the refined class activation and edge guidance network structure;
[0098] Step 2.2, Design a refined class activation attention module. The goal is to dynamically adjust the importance of each channel in the feature map by combining local and global context information, thereby enhancing the focusing ability of the refined class activation and edge guidance network on the floating raft area and suppressing background interference. The core function of the refined class activation attention module is to enhance the feature expression of the target category through the attention mechanism and recalibrate the channels using global context information. Specifically:
[0099] Step 2.2.1, Given a feature map \(X\in\mathbb{R}^{C\times H\times W}\), where \(C\), \(H\), and \(W\) represent the total number of feature channels, height, and width respectively. First, extract the global context information \(z\) through global average pooling. Then, dynamically generate channel weights through two fully connected layers and activation functions, as shown in Equation (2): C×H×W , where \(C\), \(H\), and \(W\) respectively represent the total number of feature channels, height, and width. First, extract the global context information \(z\) through global average pooling. Then, dynamically generate channel weights through two fully connected layers and activation functions, as shown in Equation (2):
[0100]
[0101] where \(W_1\) and \(W_2\) are the weights of the fully connected layers; \(\delta\) is the ReLU activation function; \(\sigma\) is the Sigmoid activation function; \(v_c\) is the weight of channel \(c\) for adjusting the importance of the channel; c is the calibrated feature map; \(X\) is the feature map before calibration. This process can dynamically adjust the importance of each channel in the feature map, strengthen category-related features, and suppress irrelevant information simultaneously. is the calibrated feature map; \(X\) is the feature map before calibration. This process can dynamically adjust the importance of each channel in the feature map, strengthen category-related features, and suppress irrelevant information simultaneously. c is the calibrated feature map; \(X\) is the feature map before calibration. This process can dynamically adjust the importance of each channel in the feature map, strengthen category-related features, and suppress irrelevant information simultaneously.
[0102] Step 2.2.2, To enhance the category activation ability of the feature map in the spatial dimension, divide the feature map into \(N\) non-overlapping patches, denoted as \(X_p\), and calculate the local category features in each patch through Equation (3): p is the calibrated feature map; \(X\) is the feature map before calibration. This process can dynamically adjust the importance of each channel in the feature map, strengthen category-related features, and suppress irrelevant information simultaneously. (p) is the calibrated feature map; \(X\) is the feature map before calibration. This process can dynamically adjust the importance of each channel in the feature map, strengthen category-related features, and suppress irrelevant information simultaneously.
[0103]
[0104] where represents the local category feature of patch \(p\); represents the category weight; \(\sigma\) s represents the spatial softmax for highlighting the significant regions within the patch; represents the category activation map within patch \(p\). Local feature aggregation can enhance the feature expression ability within a small range and is particularly suitable for semantic segmentation in complex backgrounds.
[0105] Step 2.2.3, Integrate the local features of all patches to generate global category features, as shown in Equation (4):
[0106]
[0107] Among them, F gl is the global category feature; f p is the weight of patch p, which is used to dynamically adjust the contribution of different patches to the global feature.
[0108] Step 2.2.4, optimize the global feature obtained in Step 2.2.3 by using a set of attention mechanisms to make it more suitable for class discrimination in the segmentation task. The final output of the refined class activation attention module is obtained by calculating the similarity of each pixel to the class, as shown in Equation (5):
[0109]
[0110] And combine the global feature to generate the final enhanced feature map, as shown in Equation (6):
[0111]
[0112] Among them, P (p) represents the class similarity of pixel block p; W q , W k , W v are linear transformation matrices for feature alignment, and W q represents the learnable linear transformation matrix of the input feature, W k represents the learnable linear transformation matrix of the local class feature, and W v represents the learnable linear transformation matrix of the global class feature; σ c represents the class softmax; X (p) represents a patch of the input feature; is a patch in the enhanced input feature. As Figure 3 shown is the schematic diagram of the refined class activation attention module structure.
[0113] Step 2.3, design an edge-driven feature enhancement module to handle the segmentation problems of complex or weak boundaries, especially when processing densely distributed marine aquaculture rafts in PolSAR pseudo-color images. Since edge information plays a key role in semantic segmentation, this edge-driven feature enhancement module effectively captures and strengthens boundary features by combining channel attention, spatial attention, and Sobel edge detection, generating more accurate and clear segmentation results, specifically as follows:
[0114] Step 2.3.1, the channel attention mechanism dynamically adjusts the importance of each channel in the input feature map. Specifically, define the input feature map as X ∈ R C×H×W , and generate the channel descriptor z ∈ R C, and then generates channel weights u through two fully connected operations, and the weights are used for element-wise weighted channel enhancement. The formula is:
[0115]
[0116] where W3 and W4 are both convolutional weights; δ and σ are the ReLU and Sigmoid activation functions respectively; X l represents the l-th sub-feature map in the feature map; u l represents the weight of the l-th sub-feature map; represents the result after weighting the l-th sub-feature map; C represents the total number of feature channels. In this way, the channel attention mechanism can highlight the feature expressions of significant channels while suppressing the information of irrelevant channels, and finally combine all the weighted sub-feature maps to obtain the channel attention enhanced feature X channel .
[0117] Step 2.3.2, the spatial attention mechanism enhances the information in the key spatial regions of the feature map by generating a spatial weight map. Specifically, the input feature map first generates two spatial descriptors X max ∈R H×W and X avg ∈R H×W through max-pooling and average-pooling operations in the channel dimension. Subsequently, these two spatial descriptors are concatenated in the channel dimension and a spatial weight map A s is generated through a convolutional operation. Finally, the spatial weight map A s is multiplied element-wise with the input feature map to generate the spatially enhanced feature map X spatial .
[0118] Step 2.3.3, in order to further extract explicit boundary information, this edge-driven feature enhancement module combines the classical Sobel edge detection method to generate preliminary boundary features by calculating the horizontal and vertical gradients of the input feature map X. First, calculate the gradients through the Sobel convolution kernels K x and K y in the horizontal and vertical directions, as shown in formula (8):
[0119] G x = X * K x , G y = X * K y (8);
[0120] where G x represents the horizontal direction gradient, and G y represents the vertical direction gradient;
[0121] Then, calculate the gradient magnitude G, as shown in formula (9):
[0122]
[0123] To further refine the edge features, G is normalized and activated using the ReLU function. Then, the features are further processed using depthwise separable convolutions to obtain enhanced edge features, denoted as the Sobel boundary enhanced feature X sobel 。
[0124] Step 2.3.4, the edge-driven feature enhancement module combines the channel attention enhanced feature X channel , the spatial attention enhanced feature 0X spatial and the Sobel boundary enhanced feature X sobel through element-wise addition to generate the final output feature map, as shown in Equation (10):
[0125] X output =X channel +X spatial +X sobel (10);
[0126] Such as Figure 4 shown is the schematic diagram of the edge-driven feature enhancement module structure.
[0127] Thirdly, design the polarization scattering class activation mapping module. The design of the polarization scattering class activation mapping module aims to improve the interpretability and target area localization ability of the neural network in the PolSAR pseudo-color image segmentation task. By combining the prior knowledge of polarization scattering with the gradient weight information of deep learning, the polarization scattering class activation mapping module not only generates class activation maps with clear physical meanings. It also directly participates in the training process, guiding the convolutional backbone network to extract more representative features and significantly improving the segmentation interpretability. The mariculture areas in PolSAR pseudo-color images usually exhibit significant dihedral angle scattering characteristics, which are characteristic signals formed by the multiple reflections of radar waves between rafts or floats and the water surface. The polarization scattering class activation mapping module utilizes this physical property to optimize the generation of activation maps through the prior information of dihedral angle scattering, focusing on the target area and suppressing background interference in complex marine environments. Specifically as follows:
[0128] Step 3.1, the polarization scattering class activation mapping module first generates an initial segmentation mask M for the input PolSAR pseudo-color image through the refined class activation and edge guidance network in Step 2.1, marking potential target areas (such as raft aquaculture areas). At the same time, using the dihedral angle scattering component D extracted in Step 1.2, it combines it with the segmentation mask to generate a corrected mask M D , as shown in Equation (11):
[0129] M D= Me D(11);
[0130] where e represents the element-wise multiplication operation. The corrected mask M D highlights the dihedral angle scattering characteristics within the region of interest to enhance the attention to the target area. Since the feature map A k extracted by the convolutional backbone network in step 2.1 has a different resolution from the dihedral angle scattering component D, it is necessary to D perform a downsampling operation on M D to obtain a mask M
[0131] M D ′ with the same resolution as the feature map, as shown in formula (12): D M
[0132] ′ = Resize(M
[0133] , (H, W)) (12); k where the Resize(·) function adjusts the mask to the spatial dimensions (H, W) of the feature map. D In step 3.2, the polarization scattering class activation mapping module measures the correlation of the feature map to the target area by calculating the cosine similarity between the feature map A
[0134]
[0135] and the corrected mask M k and as shown in formula (13): k where v D ′ are the vectorized representations of the feature map A and the corrected mask M
[0136] ′ respectively, and ‖·‖ is the Euclidean norm. The cosine similarity c reflects the alignment degree of the feature map with the target area. k In step 3.3, to further emphasize the importance of the target area, the polarization scattering class activation mapping module calculates the prediction score Y
[0137]
[0138] of class c for the gradient weight of the feature map A through global average pooling, as shown in formula (14):
[0139]
[0140] where: sc denotes the vector composed of the relevance score weights of all channels, α c denotes the vector composed of the gradient weights of all channels, min(·) represents finding the minimum term in the weight vector, and max(·) represents finding the maximum term in the weight vector. The normalized relevance score weights and gradient weights are combined to generate the final channel weights, as shown in Equation (16):
[0141]
[0142] where denotes the channel weight; γ is a hyperparameter used to adjust the overall influence of the weights. The channel weight is used to weight the feature map to generate the final class activation map, as shown in Equation (17):
[0143]
[0144] where the ReLU activation function ensures that only the features with positive contributions to the target region are retained.
[0145] The class activation map generated by the polarimetric scattering class activation mapping module has significant physical meaning and interpretability, and can clearly highlight the floating raft structure in the aquaculture area while effectively suppressing background interference. As Figure 1 shown in the first part of
[0146] Step 4: Design an interpretable segmentation consistency learning module. To improve the interpretability of the refined class activation and edge guidance network and enhance the consistency of segmentation feature extraction, the present invention designs an interpretable segmentation consistency learning module. The design goal of this interpretable segmentation consistency learning module is to use the high-quality class activation map generated by the polarimetric scattering class activation mapping module to guide the model learning, making it more focused on the target region (such as the aquaculture raft area) and effectively suppressing the interference of the background region, thereby improving the interpretability of the refined class activation and edge guidance network itself. Specifically as follows:
[0147] Step 4.1, generate a class activation map CAM through the polarimetric scattering class activation mapping module, which is used to highlight the attention region of the convolutional backbone network of the refined class activation and edge guidance network for the target class. The generation formula of the class activation map is as follows:
[0148]
[0149] where CAM i (x,y) represents the class activation value of class i at the spatial position (x,y); A k (x,y) is the feature map of the k-th channel; is the weight of the k-th channel corresponding to class i.
[0150] Step 4.2, the feature maps extracted from the convolutional backbone network of the refined class activation and edge guidance network are summed through the channel dimension to generate a comprehensive activation map SAM, which comprehensively reflects the feature intensity of the target region, as shown in Equation (19):
[0151]
[0152] where SAM(x,y) represents the sum of all channel feature maps at the spatial position (x,y). To ensure that the comprehensive activation map (SAM) is consistent with the class activation map (CAM) in spatial distribution, the interpretable segmentation consistency learning module constrains the two through a consistency loss function. The consistency loss is calculated based on the Mean Squared Error (MSE), as shown in Equation (20):
[0153]
[0154] where SAM′(x,y) is the normalized comprehensive activation map, and CAM′ i (x,y) is the normalized class activation map, and H and W represent the height and width of the feature map respectively.
[0155] The introduction of the interpretable segmentation consistency learning module effectively improves the interpretability of the refined class activation and edge guidance network, making its feature extraction of the target region more accurate and suppressing the interference of the background region. In addition, in complex aquaculture scenarios (such as dense raft areas or complex water surface reflections), this interpretable segmentation consistency learning module can significantly enhance the model's focusing ability on the target region, improving the segmentation accuracy and boundary consistency. As Figure 1 shown in the third part of [], the design scheme of this interpretable segmentation consistency learning module is presented.
[0156] Fifthly, design a multi-optimization objective loss function L to optimize the entire network, aiming to simultaneously improve the segmentation accuracy, edge detail description ability and model decision interpretability. Through the semantic segmentation loss L seg and the fusion loss L fuse to optimize pixel-level classification and accurately distinguish the target from the background using the cross-entropy function. Combine the edge detection loss L edge of weighted binary cross-entropy and intersection over union loss to strengthen the spatial consistency modeling of complex boundaries. At the same time, introduce the interpretable consistency loss L con to constrain the alignment of the convolutional backbone network features and the class activation map, promote the feature to focus on the target object and suppress noise interference, as shown in Equation (21):
[0157] L = L seg+L edge +L fuse +λL con (21);
[0158] Among them, λ is used to control the weight of the interpretability loss. Through end-to-end optimization, while finally ensuring the extraction accuracy of raft culture, the interpretability of the extraction process is improved.
[0159] The above-described embodiments only represent the implementation manners of the present invention, but should not be construed as limiting the scope of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A method for extracting seawater raft cultivation information of an interpretable polarization deep learning network, characterized in that, The interpretable polarization deep learning network in the above-mentioned method for extracting seawater raft culture information includes three key components: a polarization scattering class activation mapping module, an interpretable segmentation consistency learning module, and a refined class activation and edge guidance network. The polarization scattering class activation mapping module generates a class activation mapping graph by combining the polarization scattering characteristics and gradient weights of the full polarization SAR image. The interpretable segmentation consistency learning module uses the generated class activation mapping graph to guide the model to focus on the characteristics of the culture area during the training process. The refined class activation and edge guidance network adopts a dual-branch architecture, performing semantic segmentation and edge supervision respectively, and can simultaneously focus on the semantic characteristics and edge details of the culture area to generate accurate segmentation results. The method for extracting seawater raft culture information includes the following steps: First, collect the PolSAR image data of remote sensing satellites and preprocess it. Second, design a refined class activation and edge guidance network. The refined class activation and edge guidance network adopts an encoder-dual decoder architecture, including a convolutional backbone network, a semantic decoder, and an edge decoder. The semantic decoder uses a refined class activation attention module to enhance the semantic characteristics of the raft area in the PolSAR pseudo-color image. The edge decoder uses an edge-driven feature enhancement module to accurately capture the boundary information of the raft area in the PolSAR pseudo-color image. They cooperate with each other to generate the final segmentation output through feature fusion and a convolutional segmentation head. Third, design a polarization scattering class activation mapping module. The polarization scattering class activation mapping module generates a class activation graph with clear physical meaning by combining the polarization scattering prior knowledge and the gradient weight information of deep learning, and directly participates in the training process to guide the convolutional backbone network to extract representative features. The polarization scattering class activation mapping module optimizes the generation of the activation graph through the dihedral angle scattering prior information. Fourth, design an interpretable segmentation consistency learning module. The design goal of the interpretable segmentation consistency learning module is to use the class activation graph generated by the polarization scattering class activation mapping module to guide the model learning. Specifically: Step 4.1, generate a class activation map CAM through the polarization scattering class activation mapping module, which is used to highlight the attention area of the convolutional backbone network of the refined class activation and edge guidance network for the target class. The generation formula of the class activation map is as follows: Among them, CAM i (x, y) represents the class activation value of class i at the spatial position (x, y); A k (x, y) is the feature map of the k-th channel; is the weight of the k-th channel corresponding to class i; Step 4.2, the feature maps extracted from the convolutional backbone network of the refined class activation and edge guidance network are summed through the channel dimension to generate a comprehensive activation map SAM, which comprehensively reflects the feature intensity of the target area, as shown in formula (19): Among them, SAM(x,y) represents the sum of all channel feature maps at the spatial position (x,y). To ensure that the comprehensive activation map and the class activation map are consistent in spatial distribution, the interpretable segmentation consistency learning module constrains the two through a consistency loss function. The consistency loss is calculated based on the mean square error MSE, as shown in formula (20): Among them, SAM′(x, y) is the normalized comprehensive activation map, and CAM′ i (x, y) is the normalized class activation map, where H and W respectively represent the height and width of the feature map; Fifth, design a multi-optimization objective loss function L to optimize the entire network. Specifically as follows: Through the semantic segmentation loss L seg and the fusion loss L fuse Optimize pixel-level classification, accurately distinguish the target from the background using the cross-entropy function; the edge detection loss L combining weighted binary cross-entropy and intersection over union loss edge , strengthen the spatial consistency modeling of complex boundaries; at the same time introduce the interpretable consistency loss L con , constrain the alignment of the convolutional backbone network features and the class activation map, promote the feature to focus on the target object and suppress noise interference, as shown in Equation (21): L = L seg + L edge + L fuse + λL con (21); Among them, λ is used to control the weight of the interpretability loss.
2. The method for extracting seawater raft culture information of an interpretable polarization deep learning network according to claim 1, wherein The specific steps of the first step are as follows: Step 1.1: Collect PolSAR image data of remote sensing satellites through satellite resource websites, and preprocess the PolSAR image data, including radiometric correction, geometric correction, and denoising preprocessing. Step 1.2: Use the covariance matrix of the PolSAR image data obtained in Step 1.1 to represent polarization information; perform polarization decomposition on the polarization information to extract different types of scattering components from the PolSAR image data, including dihedral angle scattering, volume scattering, and surface scattering. Step 1.3: Combine the dihedral angle scattering, volume scattering, and surface scattering obtained in Step 1.2 to obtain a PolSAR pseudo-color image, and obtain a seawater aquaculture dataset and labels based on the PolSAR pseudo-color image.
3. The method for extracting seawater raft culture information of an interpretable polarization deep learning network according to claim 2, characterized in that In Step 1.2, the methods used for polarization decomposition include the Yamaguchi four-component decomposition method or the Freeman-Durden decomposition method.
4. The method for extracting seawater raft culture information of an interpretable polarization deep learning network according to claim 1, characterized in that The specific steps of the second step are as follows: Step 2.1: Design a refined class activation and edge guidance network based on the encoder-dual decoder architecture, including a convolutional backbone network, a semantic decoder, and an edge decoder, and achieve accurate semantic segmentation through feature fusion; the convolutional backbone network uses ResNet50 to extract feature maps of PolSAR image data and gradually captures multi-level spatial information; the refined class activation and edge guidance network uses a learnable parameter α to weight and fuse the feature maps of the convolutional backbone network, the semantic decoder, and the edge decoder, and its formula is expressed as: F = α·EF+(1 - α)·Upsample(DF)(1); where EF represents the features extracted by the convolutional backbone network, DF represents the features gradually decoded by the semantic decoder or the edge decoder, F represents the weighted fused features, and Upsample(·) represents the upsampling operation. After the PolSAR pseudo-color image is input into the refined class activation and edge guidance network, the convolutional backbone network ResNet50 first extracts features from it, and then sends the extracted features into the semantic decoder and the edge decoder respectively. After the features pass through the two decoders, semantic features and edge features are generated respectively, and the two types of features are fused through element-wise summation; the fused features are processed by the final convolutional segmentation head to output the final segmentation result. Step 2.2: Design a refined class activation attention module to enhance the feature expression of the target category through the attention mechanism and recalibrate the channels using global context information; specifically: Step 2.2.1, given the feature map X ∈ R C×H×W , where C, H, and W represent the total number of feature channels, height, and width respectively. First, extract the global context information z through global average pooling; then, dynamically generate channel weights through a two-layer fully connected network and activation function, as shown in formula (2): Among them, both W1 and W2 are the weights of the fully connected layer; δ is the ReLU activation function; σ is the Sigmoid activation function; v c is the weight of channel c, used to adjust the importance of the channel; is the calibrated feature map; X c is the feature map before calibration; Step 2.2.2, divide the feature map into N p non-overlapping patches, denoted as X (p) , and the local class features in each patch are calculated by formula (3): Among them, represents the local class feature of patch p; represents the class weight; σ s represents the spatial softmax, which is used to highlight the significant regions within the patch; represents the class activation map within patch p; Step 2.2.3: Integrate the local features of all patches to generate global category features, as shown in formula (4): Among them, F gl is the global category feature; f p is the weight of patch p; Step 2.2.4: Optimize the global features obtained in Step 2.2.3 using a set of attention mechanisms; the final output of the refined class activation attention module is obtained by calculating the similarity of each pixel to the category, as shown in formula (5): And combine the global features to generate the final enhanced feature map, as shown in formula (6): Among them, P (p) represents the class similarity of pixel block p; W q , W k , W v are linear transformation matrices for feature alignment, and W q represents the learnable linear transformation matrix of the input feature, W k represents the learnable linear transformation matrix of the local class feature, W v represents the learnable linear transformation matrix of the global class feature; σ c represents the class softmax; X (p) represents a patch of the input feature; is a patch in the enhanced input feature; Step 2.3: Design an edge-driven feature enhancement module to generate a more accurate and clear segmentation result, specifically as follows: Step 2.3.1, the channel attention mechanism dynamically adjusts the importance of each channel in the input feature map; specifically, define the input feature map as X ∈ R C×H×W , generate a channel descriptor z ∈ R C , then generate channel weights u through two fully connected operations, and use the weights for element-wise weighted channel enhancement; its formula is: Among them, W3 and W4 are both convolutional weights; δ and σ are the ReLU and Sigmoid activation functions respectively; X l represents the l-th sub-feature map in the feature map; u l represents the weight of the l-th sub-feature map; represents the result after weighting the l-th sub-feature map; C represents the total number of channels of the features; combining all the weighted sub-feature maps together gives the channel attention enhanced feature X channel ; Step 2.3.2, the spatial attention mechanism enhances the information in the key spatial regions of the feature map by generating a spatial weight map; specifically, the input feature map first generates two spatial descriptors X max ∈R H×W and X avg ∈R H×W through max-pooling and average-pooling operations in the channel dimension; subsequently, these two spatial descriptors are concatenated in the channel dimension and a spatial weight map A s is generated through a convolution operation; finally, the spatial weight map A s is multiplied element-wise with the input feature map to generate a spatially enhanced feature map X spatial ; Step 2.3.3, the edge-driven feature enhancement module combines the Sobel edge detection method to generate preliminary boundary features by calculating the horizontal and vertical gradients of the input feature map X. First, the Sobel convolution kernels K x and K y are used to calculate the gradients as shown in Equation (8): G x = X * K x , G y = X * K y (8); Among them, G x represents the horizontal gradient, and G y represents the vertical gradient; Then, calculate the gradient magnitude G as shown in Equation (9): To further refine the edge features, first normalize G and activate it using the ReLU function; then, process the features using depthwise separable convolution to obtain enhanced edge features, denoted as the Sobel boundary enhanced feature X sobel ; Step 2.3.4, the edge-driven feature enhancement module fuses the channel attention enhanced feature X channel , the spatial attention enhanced feature X spatial and the Sobel edge enhanced feature X sobel by element-wise addition to generate the final output feature map, as shown in Equation (10): X output = X channel + X spatial + X sobel (10).
5. The method for extracting seawater raft culture information of an interpretable polarization deep learning network according to claim 4, wherein The specific steps of the third step are as follows: Step 3.1, the polarization scattering class activation mapping module first generates an initial segmentation mask M for the input PolSAR pseudo-color image through the refined class activation and edge guidance network in Step 2.1 to mark potential target areas; meanwhile, the dihedral angle scattering component D after Freeman decomposition extracted in Step 1.2 is used to combine with the segmentation mask to generate a corrected mask M D , as shown in Equation (11): M D = M ⊙ D(11); Among them, ⊙ represents the element-wise multiplication operation; the corrected mask M D Highlights the dihedral angle scattering characteristics within the region of interest to enhance the attention to the target area; Since the feature map A extracted by the convolutional backbone network in step 2.1 k has a different resolution from the dihedral angle scattering component D, it is necessary to perform a downsampling operation on M D to obtain a mask M D ′ with the same resolution as the feature map, as shown in formula (12): M D M' = Resize(M D , (H, W))(12); Among them, the Resize(·) function adjusts the mask to the spatial dimensions (H, W) of the feature map; Step 3.2, the polarization scattering class activation mapping module measures the correlation of the feature map to the target region by calculating the cosine similarity between the feature map A k and the corrected mask M D ′, as shown in formula (13): Among them, v k and are respectively the vectorized representations of the feature map A k and the corrected mask M D '; ‖·‖ is the Euclidean norm; the cosine similarity reflects the alignment degree between the feature map and the target region; Step 3.3, to further emphasize the importance of the target region, the polarization scattering class activation mapping module calculates the prediction score Y of class c through global average pooling c for the feature map A k of the gradient weight, as shown in formula (14): By gradient weights captures the importance information of the feature map for the target class; to balance the contributions of similarity and gradient weights, both are normalized separately, as shown in Equation (15): where: s c represents the vector composed of the relevance score weights of all channels, and α c represents the vector composed of the gradient weights of all channels. min(·) represents finding the minimum term in the weight vector, and max(·) represents finding the maximum term in the weight vector; combining the normalized relevance score weights and gradient weights to generate the final channel weights, as shown in formula (16): Among them, represents the channel weight; γ is a hyperparameter used to adjust the overall influence of the weight; the channel weight is used to weight the feature map to generate the final class activation map, as shown in formula (17): Among them, the ReLU activation function ensures that only the features with positive contributions to the target region are retained.
Citation Information
Patent Citations
Analysis Methods for Backscattering Mechanism in Seawater Raft Aquaculture under Random Sea States
CN110910494B
Breeding area prediction method based on improved nested UNet neural network
CN119398540A
SAR image buoyant raft cultivation information extraction method based on semi-supervised cyclic consistency generative adversarial network
CN115578645A
Diving battery feature identification method based on interpretable neural network
CN116298915A