A Method for Extracting SAR Seawater Aquaculture Information Based on a Completely Unsupervised Transformer Network

Through the self-supervised pre-training of fully unsupervised Transformer network and the optimization of pseudo-label generator, the problem of accurate extraction of marine aquaculture targets in marine remote sensing images is solved, and efficient segmentation effect is achieved under label-free data.

CN119851157BActive Publication Date: 2025-07-08DALIAN UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510323232.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The prior art is difficult to accurately extract seawater aquaculture targets in marine remote sensing images, especially in complex scenarios, where traditional unsupervised methods are limited and require time-consuming and labor-intensive supervision methods that rely on a large number of labeled data.

Method used

Using a fully unsupervised Transformer network, the self-supervised pre-training and global optimal pseudo-label generator and fine-grained discrimination and completion modules are automatically generated and high-quality pseudo-labels are trained to realize unsupervised segmentation of seawater aquaculture information.

Benefits of technology

Under labelless data, the segmentation accuracy and robustness of marine aquaculture targets are improved, the dependence on labeled data is reduced, and the segmentation performance in complex scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851157B_ABST
    Figure CN119851157B_ABST
Patent Text Reader

Abstract

The present invention provides a method for extracting SAR mariculture information based on a fully unsupervised Transformer network, belonging to the cross technical field of marine remote sensing and artificial intelligence. In the first stage of the present invention, general features of the target are extracted from a large number of unlabeled SAR mariculture images to provide initial weights for the downstream segmentation network. In the second stage, a globally optimal pseudo-label generator and a fine-grained discrimination and completion module are designed: the former captures globally similar semantic features to generate initial pseudo-labels, eliminating the interference of confusing targets while enhancing semantic relevance; the latter improves the local continuity of the pseudo-labels by complementing the detailed and continuous semantic features of the target in regions. The present invention can solve the problems of being unable to deeply mine the intrinsic information of a large amount of unlabeled mariculture and the limited segmentation performance of traditional unsupervised methods in complex scenarios. By using continuously optimized pseudo-labels as ground truth to train the segmentation network, a fully unsupervised mariculture segmentation task can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of marine remote sensing and artificial intelligence, and relates to a method for extracting SAR seawater aquaculture information based on a completely unsupervised Transformer network. Background Art

[0002] China is rich in marine fishery resources, and the output of marine aquaculture accounts for more than 50% of the total global marine aquaculture output. However, unreasonable marine aquaculture planning may lead to serious water quality deterioration, causing significant losses to the marine ecological balance and the aquaculture industry. Therefore, obtaining accurate and comprehensive marine aquaculture location information is crucial, which provides a scientific basis for monitoring. Synthetic Aperture Radar (SAR) is widely used in marine monitoring tasks because it can work continuously during the day and has high resolution, and can obtain a large amount of marine remote sensing data. The scattering characteristics of seawater aquaculture rafts are significantly different from those of seawater: seawater aquaculture rafts mainly exhibit double - bounce scattering characteristics, while seawater mainly presents surface scattering characteristics. Therefore, in SAR images, seawater aquaculture targets are usually brighter than the seawater background.

[0003] Although existing supervised methods can extract seawater aquaculture targets more accurately, they usually rely on a large amount of labeled data. For example, the Chinese invention patent designs a method for extracting seawater aquaculture areas based on deep learning (Chinese invention patent CN112766155A), which requires manual creation of ground - truth maps for each remote - sensing image to train the network, and obtaining high - quality labeled data is time - consuming and laborious. At the same time, traditional unsupervised methods are affected by speckle noise in SAR images and it is difficult to accurately extract aquaculture areas in complex seawater aquaculture images. For example, the Chinese invention patent designs a SAR image semantic segmentation method based on a deep convolutional network and weak - supervised learning (Chinese invention patent CN109344837A), which uses the target features extracted by the convolutional neural network to generate pseudo - labels as ground - truth to train the network, while the convolutional neural network ignores the global semantic information of the target. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention proposes a method for extracting SAR mariculture information based on a completely unsupervised Transformer network, which can solve the problems of being unable to deeply mine the intrinsic information of a large amount of unlabeled mariculture and the limited segmentation performance of traditional unsupervised methods in complex scenarios. In the first stage of the present invention, general features of the target are extracted from a large number of unlabeled SAR mariculture images to provide initial weights for the downstream segmentation network. In the second stage, a globally optimal pseudo-label generator and a fine-grained discrimination and completion module are designed: the former captures globally similar semantic features to generate initial pseudo-labels, eliminating the interference of confusing targets while enhancing semantic relevance; the latter improves the local continuity of the pseudo-labels by complementing the details and continuous semantic features of the target in regions. Finally, the continuously optimized pseudo-labels are used as ground truth to train the segmentation network, realizing a completely unsupervised mariculture segmentation task.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A method for extracting SAR mariculture information based on a completely unsupervised Transformer network, the SAR mariculture information extraction method including two stages: upstream self-supervised pre-training and downstream unsupervised pseudo-label training. First, in the upstream stage, a large amount of unlabeled SAR image data is input into the self-supervised Transformer network for pre-training to obtain the trained network weights, and the self-supervised Transformer network is defined here as the self-supervised Transformer network. Next, in the downstream stage, a globally optimal pseudo-label generator and a fine-grained discrimination and completion module are designed. Specifically: First, the network weights obtained from the upstream self-supervised pre-training are used to initialize the Teacher network, and the downstream training data is input into the Teacher network, and initial pseudo-labels of the training images are generated through the globally optimal pseudo-label generator. Secondly, the initial pseudo-labels are used as constraints to guide the training of the Student network. The image is segmented by the Student network, and its output result is combined with the features learned by the Teacher network and jointly input into the fine-grained discrimination and completion module. The initial pseudo-labels are finely optimized through the fine-grained discrimination and completion module to refine the detailed features of the mariculture target and enhance its continuous information, thereby generating optimized pseudo-labels. Finally, the optimized pseudo-labels are used to replace the initial pseudo-labels to constrain the segmentation network, and the output of the segmentation network in turn optimizes the pseudo-labels. Through the iterative optimization mechanism of the pseudo-labels, the quality of the pseudo-labels and the performance of the segmentation network are continuously improved. Through this mutually optimizing method, the completely unsupervised extraction of SAR mariculture information is finally realized. Specifically, it includes the following steps:

[0007] The first step is to construct an unlabeled mariculture dataset, which specifically includes the following steps:

[0008] Step 1.1, obtain SAR remote sensing image data from the satellite resource platform, and preprocess the SAR remote sensing image, including radiometric calibration, image enhancement, and terrain correction.

[0009] Step 1.2, filter the SAR remote sensing image processed in Step 1.1 to reduce the influence of speckle noise.

[0010] Step 1.3, crop the filtered SAR remote sensing image to a fixed size to generate a SAR mariculture dataset, and divide it into an upstream self-supervised pre-training dataset and a downstream unsupervised pseudo-label training dataset according to a ratio. Further divide the downstream unsupervised pseudo-label training dataset into a training set and a test set, and make labels for the test set.

[0011] Second step, upstream pre-training stage, use a self-supervised Transformer network to train the upstream self-supervised pre-training dataset, specifically as follows:

[0012] Step 2.1, the self-supervised Transformer network consists of and combined, the network parameters of are and the network parameters of are . First, randomly crop each image in the upstream self-supervised pre-training dataset to obtain a global view and a local view . Input the global view into for training, input the local view into for training, and then obtain the two output probabilities of as and .

[0013] Step 2.2, calculate the cross-entropy loss between and using formula (1), and as shown in formula (2), use the calculated cross-entropy loss as the loss function of for backpropagation to update the parameters of , and formulas (1) and (2) are as follows:

[0014] (1);

[0015] (2);

[0016] Among them, represents the length of the output probability distribution vector of the upstream pre-trained model; represents the predicted probability on represents the predicted probability on the represents the predicted probability on the logarithm; represents the gradient of the cross-entropy loss function with respect to parameter for optimizing parameters; is the gradient of the cross-entropy loss with respect to output probability representing the probability and the is output probability with respect to parameter partial derivative.

[0017] Step 2.3, meanwhile, use the network parameters to update the network parameters by the method of Exponential Moving Average (EMA), as shown in formula (3):

[0018] (3);

[0019] Among them, is the smoothing factor of EMA, used to control the influence degree of historical parameters.

[0020] Step 2.4, after the upstream self-supervised Transformer network training is completed, freeze the parameters .

[0021] The third step is to perform downstream unsupervised pseudo-label training. The downstream training network consists of a Teacher network and a Student network. First, freeze the parameters in the second step As the initial weights of the Teacher network, the weights of the Teacher network , , and then use the Teacher network to extract features from the images in the downstream unsupervised pseudo-label training dataset obtained in step 1.3. After that, the features output by the Teacher network are input into the globally optimal pseudo-label generator to generate initial pseudo-labels, and use the initial pseudo-labels as constraints to train the Student network, specifically as follows:

[0022] Step 3.1, use the parameters of the pre-trained and frozen upstream to initialize the parameters of the Teacher network. The Teacher network consists of N multi-head self-attention modules, and each multi-head self-attention module consists of H self-attention heads, which can extract features and perform global modeling on the input data. In the multi-head self-attention module, there are query (Query, ), key (Key, ), and value (Query, ). In the last multi-head self-attention module, there is a special classification vector for marking the overall classification information.

[0023] Step 3.2, define the training images in the downstream unsupervised pseudo-label training dataset obtained in step 1.3 as , divide into P image patches, input them into the Teacher network, and after feature extraction by N multi-head self-attention modules, obtain the last multi-head self-attention feature representation , . The image patch feature corresponds to the feature of the th image patch in the input image. , where and respectively represent the query and key of the th image patch in the th self-attention head.

[0024] Step 3.3, input the last multi-head self-attention feature obtained in step 3.2 into the globally optimal pseudo-label generator, and use formula (4) to calculate the affinity vector between the classification vector and the image patch feature as . Use formula (5) to calculate the affinity matrix between the image patch feature and the image patch feature as . Formulas (4) and (5) are as follows:​

[0025] (4);

[0026] (5);

[0027] Among them, represents the affinity vector between the classification vector and the image patch feature under the self-attention head ; is the dimensional space of the calculation result, where is the number of image patches, ; represents the query vector of the classification vector in the self-attention head ; represents the -th image patch's key in the self-attention head taking the transpose; represents the currently calculated self-attention head; represents the similarity matrix between image patch features under the self-attention head ; is the dimensional space of the calculation result; represents the -th image patch's key in the self-attention head and taking its transpose.

[0028] Step 3.4, then define the image patch with the highest affinity in the affinity vector as the key image patch , . Perform the calculation sequentially for each self-attention head to obtain the key image patch set , is the number of self-attention heads. Among them, the key image patch contains the most discriminative features in the training target. Then, diffuse the key image patch of each self-attention head to the global image patch set , where the global image patch set represents all the image patches with high similarity to the key image patch in the -th self-attention head, as shown in formula (6):

[0029] (6);

[0030] Among them, represents the currently calculated image patch index; Represents a set of image patch indices, which is the number of image patches; Represents in the th self-attention head, the similarity between the key image patch and the image patch is extracted from the similarity matrix .

[0031] Step 3.5, Subsequently, using the -means clustering algorithm (K-Means Clustering) for calculating semantic similarity, cluster the global image patch set of each self-attention head obtained in Step 3.4 into two categories: aquaculture and seawater. Define the image patch set with the category of aquaculture as , and define the image patch set with the category of seawater as . Calculate the patch-level pseudo-labels for each self-attention head image patch using formula (7), where represents the th image patch, and the patch-level pseudo-label of the th self-attention head. After generating each patch-level pseudo-label, the whole-image-level pseudo-label can be stitched together according to the position of the image patch in the entire image . is the whole-image-level pseudo-label obtained from the th self-attention head. Thus, the initial pseudo-label set is obtained. , , where

[0032] (7);

[0033] Step 3.6, Screen the initial pseudo-label set obtained in Step 3.5. Calculate the density value of the whole-image-level pseudo-label corresponding to each attention head using formula (8), and then calculate the image-level pseudo-label with the maximum density value as the final initial pseudo-label using formula (9). Formulas (8) and (9) are as follows:

[0034] (8);

[0035] (9);

[0036] where, is the number of image patches; is the image patch The average distance of the image patch within its category; For the image patch The average distance from the image patches of different categories; Indicates taking and The maximum value in; The density value The value range is [-1, 1].

[0037] In the fourth step, the segmentation result of the Student network and the feature map output by the Teacher network are input into the fine-grained discrimination and completion module to further optimize the pseudo-label. The Student network consists of an encoder and a decoder. The encoder part is based on a multi-scale Transformer architecture for extracting the features of the input data, and the decoder adopts the UPerHead structure (Unified Perceptual Parsing Head), including a feature fusion and multi-scale information aggregation unit. The initial parameters of the Student network are , specifically as follows:

[0038] Step 4.1, the downstream unsupervised pseudo-label training dataset images obtained in Step 1.3 Through data augmentation methods such as rotation, scaling, and linear stretching, the enhanced images are obtained. The enhanced images are input into the encoder of the Student network to extract multi-scale features. The extracted multi-scale features are input into the decoder, and after step-by-step upsampling and feature fusion operations, the feature map is output. Subsequently, the feature map is binarized after convolution processing, and finally the output segmentation result is obtained. At the same time, the enhanced images are input into the Teacher network. The Teacher network learns the correlation between different spatial positions based on the multi-head self-attention module to generate the saliency feature map . The saliency feature map is binarized after convolution processing, and finally the saliency feature result is output.

[0039] Step 4.2, then, the segmentation result of the Student network and the saliency feature result output by the Teacher network are input into the fine-grained discrimination and completion module. First, the Canny edge detection operator is used to extract the boundary information of the target in the segmentation result , and the boundary information , the boundary information Represented in a set form, each element in the set represents an independent target boundary in the segmentation result, . According to each boundary generate a set of minimum bounding boxes based on the coordinates of the four points: the upper left corner, the upper right corner, the lower left corner, and the lower right corner of each boundary , where each minimum bounding box is as shown in formula (10):

[0040] (10);

[0041] Among them, represents the minimum bounding box of the th target boundary; the bounding box contains all pixel points that meet the conditions ; represents a pixel point in the image, whose abscissa is , and the ordinate is ; represents the minimum value of the abscissas of all pixel points in the boundary ; represents the maximum value of the abscissas of all pixel points in the boundary ; represents the minimum value of the ordinates of all pixel points in the boundary ; represents the maximum value of the abscissas of all pixel points in the boundary ; represents the number of aquaculture areas in the image, which is also the number of boundaries.

[0042] Step 4.3, the number of boundaries and their corresponding bounding boxes in each image is determined by the number of targets in the image and is represented in a set form. Obtain the feature within the bounding box using formula (11), and obtain the feature within the bounding box using formula (12). Formulas (11) and (12) are as follows:

[0043] (11);

[0044] (12);

[0045] Among them, contains all pixel values that meet ; ; contains all pixel values that meet ; ;

[0046] Step 4.4. Then, from Step 4.3, a set of feature extraction results can be obtained, specifically: the feature set of the segmentation result output by the Student network within the bounding box and the feature set of the saliency result output by the Teacher network within the bounding box , where is the number of bounding boxes obtained in Step 4.2. In the discrimination stage of the fine-grained discrimination and completion module, by inputting and features within the same bounding box, IoU (Intersection over Union) calculation is performed, as shown in Equation (13):

[0047] (13);

[0048] where represents performing IoU calculation on and ; represents calculating the intersection between and pixels; represents calculating the union between and pixels;

[0049] Step 4.4. Finally, the fine-grained discrimination and completion module operates on the features within each bounding box independently. When the IoU between and exceeds the IoU threshold , then the features between and are fused. Otherwise, the features of are retained until the completion operation is performed on each box, and finally the optimized pseudo-label is obtained, as shown in Formulas (14) and (15):

[0050] (14);

[0051] (15);

[0052] where represents the feature result obtained in the bounding box ; represents fusing the features between and ; represents the threshold of IoU, and its value range is [0, 1].

[0053] Step 5. Establish the objective function and network iterative optimization process of the downstream unsupervised pseudo-label training network. Specifically as follows:

[0054] Step 5.1. Calculate the initial pseudo-labels for global optimal pseudo-label generation using Equation (16) and the segmentation result output by the Student network to obtain the cross-entropy loss . Similarly, use Equation (17) to calculate the optimized pseudo-labels and the segmentation result output by the Student network to obtain the cross-entropy loss . Equations (16) and (17) are as follows:

[0055] (16);

[0056] (17);

[0057] where represents the total number of pixels of the downstream unsupervised pseudo-label training image .

[0058] Step 5.2. The total objective function of the Student network is , and are coefficients and are related to the number of training rounds of the downstream unsupervised pseudo-label training network. When , ; when , . After that, use the objective function as the loss of the Student network, and use Equation (18) to update the weight parameters of the Student network in the reverse direction; at the same time, use Equation (19) to update the weight parameters of the Teacher network by means of exponential weighted moving average (EMA) using the weight parameters of the Student network. Equations (18) and (19) are as follows:

[0059] (18);

[0060] (19);

[0061] where is the learning rate; is the loss function with respect to the parameters of the network; is the smoothing factor for EMA, used to control the influence degree of historical parameters.

[0062] Step 5.3, when steps 3.1 to 5.2 are completed, one round of training is completed. The training of the downstream unsupervised pseudo-label training network affects the Student network and the Teacher network through pseudo-labels, and at the same time, the pseudo-labels are optimized by the prediction results of the Student network and the Teacher network, so as to realize the iterative update of the pseudo-labels. When the number of training rounds is reached, the downstream unsupervised pseudo-label training stops training, and the segmentation output by the Student network is used as the final output of the network.

[0063] The beneficial effects of the present invention are as follows:

[0064] (1) The present invention includes two stages: upstream self-supervised pre-training and downstream unsupervised pseudo-label training. The upstream self-supervised pre-training can extract the intrinsic features in the data from a large amount of unlabeled seawater aquaculture datasets. The downstream unsupervised pseudo-label training generates pseudo-labels with high-quality semantic features for training images through the iterative optimization of the Teacher network and the Student network.

[0065] (2) The present invention designs a globally optimal semantic pseudo-label generator. It uses the feature learning capabilities of different attention heads to spread from the most discriminative regions of the target to the global environment. This process captures similar semantic features in the entire region and generates initial pseudo-labels, while enhancing semantic relevance and eliminating the interference of confusing targets.

[0066] (3) The present invention constructs a fine-grained discrimination and completion module to improve the local continuity of pseudo-labels. The fine-grained discrimination and completion module calculates the overlapping part of the segmentation result of the Student network and the saliency result of the Teacher network in the discrimination stage, so as to identify the regions that need to be optimized. In the complementation stage, the target details and continuity information within the sub-regions are used to enhance the pseudo-labels. The experimental data are two SAR satellite seawater aquaculture datasets, which verify the effectiveness of the present invention. Brief Description of the Drawings

[0067] Figure 1 is a flowchart of a SAR seawater aquaculture information extraction method for a completely unsupervised Transformer network.

[0068] Figure 2 is a schematic diagram of the fine-grained discrimination and completion module.

[0069] Figure 3It is a result map for extracting seawater aquaculture from SAR images. Among them, column (a) is the SAR image, column (b) is the ground truth label, and column (c) is the network prediction result. The first row of each column is the selected GF-3 SAR satellite image sample 1, the second row of each column is the selected GF-3 SAR satellite image sample 2, the third row of each column is the selected Radarsat-2 SAR satellite image sample 1, and the fourth row of each column is the selected Radarsat-2 SAR satellite image sample 2.

[0070] Figure 4 It is a visualization result map of the pseudo-label changing with the number of network training rounds. Among them, column (a) is the initial pseudo-label; column (b) is the first-round optimized pseudo-label; column (c) is the final pseudo-label; column (d) is the ground truth label. The first row of each column is the selected GF-3 SAR satellite image sample 1, the second row of each column is the selected GF-3 SAR satellite image sample 2, the third row of each column is the selected GF-3 SAR satellite image sample 3, the fourth row of each column is the selected GF-3 SAR satellite image sample 4, and the fifth row of each column is the selected GF-3 SAR satellite image sample 5. Detailed implementation manners

[0071] To make the method problems solved by the present invention, the adopted method solutions and the achieved method effects clearer, the present invention will be further described in detail below with reference to the accompanying drawings and examples. It can be understood that the specific examples described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that, for the sake of description, only the parts related to the present invention are shown in the drawings rather than all the contents.

[0072] As Figure 1 shown, a method for extracting SAR seawater aquaculture information by a completely unsupervised Transformer network provided by an embodiment of the present invention includes:

[0073] Compile under python3.8, pytorch1.8.1 and cuda11.1 on the windows11 system, run using the GPU of NVIDIA GeForce RTX 3090, and the input data size is of SAR images.

[0074] The first step is to obtain SAR remote sensing image data on the satellite resource website, preprocess the remote sensing images, perform radiometric calibration, image enhancement, and terrain correction on the SAR data, and make a dataset within the research area.

[0075] Radiometric calibration: Convert the digital image data into actual physical quantities such as reflectance or radiance temperature, and eliminate the radiation bias of the SAR image caused by the sensor.

[0076] Image enhancement: The Refined Lee filtering method with a window size of is adopted to reduce the speckle noise in the SAR image.

[0077] Terrain correction: SAR remote sensing images have certain geometric distortions, so registration and correction are required.

[0078] Dataset production: For the filtered SAR remote sensing images, they are cropped into images of size, and more than 13,000 images are produced as the upstream self-supervised pre-training dataset, 369 images are used as the training set for downstream unsupervised pseudo-label training, and 160 images are used as the test set. The longitude, latitude and shooting time corresponding to each image are recorded, and finally the test set labels are produced.

[0079] In the second step, the upstream pre-training stage, a self-supervised Transformer network is used to train the data. As Figure 1 shown in the upper part, this network is composed of and combined. First, the unlabeled seawater aquaculture upstream self-supervised pre-training dataset produced in the first step is processed by random cropping. A global view is cropped according to the area ratio [0.4, 1], and a local view is cropped according to the area ratio [0.05, 0.4]. The global view is input into for training, and the local view is input into for training. Then, the output result of is operated by center bias (Center), and then the probability result is output through Softmax. At the same time, the output result of is output through Softmax as the probability result. Calculate the cross-entropy loss of and

[0080] (1);

[0081] (2);

[0082] where denotes the length of the output probability distribution vector of the upstream pre-trained model; denotes the predicted probability on the denotes the predicted probability on the th category; denotes the predicted probability on the th category in logarithm; denotes the gradient of the cross-entropy loss function with respect to the parameter for optimizing the parameter of ; is the gradient of the cross-entropy loss with respect to the output probability denotes the probability of and the probability of ; the error between them; is the partial derivative of the output probability with respect to the parameter ;

[0083] Meanwhile, the network parameters of are updated by using the Exponential Moving Average (EMA) method as follows: where,

[0084] (3)

[0085] where is the smoothing factor of EMA, used to control the influence degree of historical parameters;

[0086] In the third step, after the upstream self-supervised Transformer network is trained, the parameters of are frozen Figure 1 . Then, downstream unsupervised pseudo-label training is carried out. As shown in the lower part, the downstream unsupervised pseudo-label training network consists of a Teacher network and a Student network. The network parameters of pre-trained by the upstream self-supervised are used to initialize the network parameters of the Teacher network , ​​​. The Teacher network consists of 6 multi - head self - attention modules, which can extract features and perform global modeling on the input data. In the multi - head self - attention module, it contains Query ( ), Key ( ), and Value (Query, ). In the last - layer multi - head self - attention module, there is a special classification vector , which is used to mark the overall classification information. First, the training images of the downstream unsupervised pseudo - label training dataset made in the first step are defined as , segmented into 1024 image patches, input into the Teacher network, and passed through 12 multi - head self - attention modules for feature extraction to obtain the last - layer multi - head self - attention feature representation , , p = 1024, the image patch feature corresponds to the feature of the th image patch in the input image. , where and respectively represent the query and key of the th image patch in the th self - attention head. Then, the last - layer multi - head self - attention feature is input into the global - optimal pseudo - label generator.

[0087] The affinity vector between the classification vector and the image patch feature is calculated using formula (4) as . The affinity matrix between the image patch feature and the image patch feature is calculated using formula (5) as . Formulas (4) and (5) are as follows:

[0088] (4);

[0089] (5);

[0090] Among them, represents the affinity vector between the classification vector and the image patch feature under the self - attention head ; is the dimensional space of the calculation result, where is the number of image patches, ; represents the query vector of the classification vector in the self - attention head ; Indicates the th image patch in the key of the self-attention head and take the transpose; Indicates the currently computed self-attention head; Indicates the similarity matrix between image patch features under the self-attention head is the dimensional space of the calculation result; Indicates the th image patch in the key of the self-attention head and take its transpose.

[0091] Then, define the image patch with the highest affinity in the affinity vector as the key image patch , . Perform the calculation sequentially for each self-attention head to obtain the set of key image patches , . Among them, the key image patches contain the most discriminative features in the training target. Then, diffuse the key image patches of each self-attention head to the global image patch set , where the global image patch set represents all image patches with high similarity to the key image patches in the th self-attention head, as shown in the following formula:

[0092] (6)

[0093] Among them, represents the currently computed image patch index; represents the set of image patch indices, is the number of image patches; represents the th self-attention head, the similarity between the key image patch and the image patch , extracted from the similarity matrix .

[0094] Subsequently, use the -means clustering algorithm (K-Means Clustering) for computing semantic similarity to cluster the global image patch set of each self-attention head into two categories: aquaculture and seawater. Define the set of image patches with the category of aquaculture as , and define the set of image patches with the category of seawater as , and use formula (7) to calculate the pseudo-labels at the image patch level for each self-attention head, where Indicates image blocks, After each block-level pseudo-label is generated, it can be spliced ​​into a pseudo-label at the whole image level according to the position of the image block in the whole image. , For the The pseudo labels of the whole image are obtained by the self-attention heads. Then the initial pseudo label set is obtained , , formula (7) is as follows:

[0095] (7);

[0096] Then, the initial pseudo-label set Filter and use formula (8) to calculate the entire image-level pseudo label corresponding to each attention head The density value , and then use formula (9) to calculate the image-level pseudo-label with the largest density value as the final initial pseudo-label , formulas (8) and (9) are as follows:

[0097] (8);

[0098] (9);

[0099] in, is the number of image blocks; For image blocks The average distance of image patches within the class to which they belong; For image blocks The average distance to its different category image patches; Indicates taking and The maximum value in The value range is [-1, 1].

[0100] In the fourth step, the segmentation results of the Student network and the feature maps output by the Teacher network are input into the fine-grained discrimination and completion module to further optimize the pseudo labels. The Student network consists of an encoder and a decoder. The encoder is based on a multi-scale Transformer architecture to extract the features of the input data. The decoder adopts a Unified Perceptual Parsing Head (UPerHead) structure, including feature fusion and multi-scale information aggregation units. The initial parameters of the Student network are . The downstream unsupervised pseudo-label training dataset images obtained in the first step Data augmentation methods such as rotation, scaling, and linear stretching are used to obtain the enhanced image , and the enhanced image is input into the encoder of the Student network to extract multi-scale features. The extracted multi-scale features are input into the decoder, and after successive upsampling and feature fusion operations, a feature map is output . Subsequently, the feature map is binarized after convolution processing, and finally the output segmentation result is obtained . At the same time, the enhanced image is input into the Teacher network. The Teacher network learns the correlation between different spatial positions based on the multi-head self-attention module and generates a saliency feature map . The saliency feature map is convolved and binarized, and finally the saliency feature result is output . As Figure 2 shown in the specific architecture of fine-grained discrimination and completion, the segmentation result of the Student network and the saliency feature result output by the Teacher network are input into the fine-grained discrimination and completion module. First, the Canny edge detection operator is used to extract the boundary information of the target in the segmentation result . The boundary information is represented in the form of a set, and each element in the set represents an independent target boundary in the segmentation result . According to the coordinates of the four points of the upper left corner, upper right corner, lower left corner, and lower right corner of each boundary , a set of minimum bounding boxes is generated . Each minimum bounding box is shown in formula (10) as follows

[0101] (10);

[0102] where represents the minimum bounding box of the th target boundary; the bounding box contains all pixel points that meet the conditions ; represents a pixel point in the image, whose abscissa is , and the ordinate is ; represents the minimum value of the abscissas of all pixel points in the boundary ; represents the boundary The maximum value of the abscissa of all pixel points in Indicates the boundary The minimum value of the ordinate of all pixel points in Indicates the boundary The maximum value of the abscissa of all pixel points in Indicates the number of aquaculture areas in the image, which is also the number of boundaries.

[0103] The number of boundaries and their corresponding bounding boxes in each image is determined by the number of targets in the image and is represented in set form. Obtained using formula (11) In the bounding box The features within , obtained using formula (12) In the bounding box The features within , Formulas (11) and (12) are as follows:

[0104] (11);

[0105] (12);

[0106] Among them, Contains all pixels that satisfy The pixel values of ; Contains all pixels that satisfy The pixel values of ;

[0107] Then, the set of feature extraction results can be obtained from the above operations. Specifically: The set of features of the segmentation result output by the Student network in the bounding box And the set of features of the saliency result output by the Teacher network in the bounding box , where Is the number of bounding boxes obtained in step 4.2. In the discrimination stage of the fine-grained discrimination and completion module, by inputting And The features within the same bounding box are used to calculate the IoU (Intersection over Union), as shown in equation (13):

[0108] (13);

[0109] Among them, Indicates the IoU calculation for And ; Indicates the calculation of And The intersection between pixels; Denote the calculation and the union between pixels;

[0110] Finally, the fine-grained discrimination and completion module operates on the features within each bounding box independently. When and the IoU between them exceeds the IoU threshold then fuse and the features between them. Otherwise, retain the features until the completion operation is done for each box, and finally obtain the optimized pseudo-label , as shown in Formulas (14) and (15):

[0111] (14);

[0112] (15);

[0113] wherein, denotes the feature result obtained in the bounding box ; denotes the fused features between and ; denotes the threshold of IoU, with a value of 0.3.

[0114] Step 5, optimize the Student network by the two cross-entropy losses in the lower part of Figure 1 , and establish the objective function and network iterative optimization process of the downstream network. Calculate the cross-entropy loss between the initial pseudo-label generated by the global optimal pseudo-label and the segmentation result output by the Student network using Formula (16). Similarly, calculate the cross-entropy loss between the optimized pseudo-label and the segmentation result output by the Student network using Formula (17). Formulas (16) and (17) are as follows:

[0115] (16);

[0116] (17);

[0117] wherein, represents the total number of pixels of the image , =262144.

[0118] Therefore, the total objective function of the Student network is , and are coefficients. and are related to the number of training rounds of the downstream unsupervised pseudo-label training network When , ; when , . After that, the objective function is used as the loss of the Student network, and the weight parameters of the Student network are updated backward using formula (18) ; at the same time, using formula (19), the weight parameters of the Student network are used to update the weight parameters of the Teacher network through the method of Exponential Moving Average (EMA). Formulas (18) and (19) are as follows:

[0119] (18);

[0120] (19);

[0121] Among them, is the learning rate; is the parameter of the loss function with respect to the network; is the smoothing factor of EMA, used to control the influence degree of historical parameters.

[0122] Step 5: When steps 2 to 4 are completed, one round of training is completed. The training of the downstream unsupervised pseudo-label training network affects the Student network and the Teacher network through pseudo-labels, and at the same time, the pseudo-labels are optimized by the prediction results of the Student network and the Teacher network, so as to achieve iterative update of the pseudo-labels. When the number of training rounds is reached, the downstream unsupervised pseudo-label training stops training, and the segmentation output by the Student network is used as the final output of the network.

[0123] Figure 4This is a visualization result graph of the change of pseudo-labels with the number of network training rounds. Among them, column (a) is the initial pseudo-label, column (b) is the pseudo-label optimized in the first round, column (c) is the final pseudo-label, and column (d) is the ground truth label. Among them, the first row of each column is the selected high-score No. 3 SAR satellite image sample 1, the second row of each column is the selected high-score No. 3 SAR satellite image sample 2, the third row of each column is the selected high-score No. 3 SAR satellite image sample 3, the fourth row of each column is the selected high-score No. 3 SAR satellite image sample 4, and the fifth row of each column is the selected high-score No. 3 SAR satellite image sample 5. Analysis Figure 4 From columns (a) to (d) in it, it can be seen that: column (a) shows that the initial pseudo-label is generated by the model under unsupervised conditions and contains more noise and misclassifications; column (b) shows that after one round of training optimization, the quality of the pseudo-label has improved, and the misclassifications in some areas have been corrected, but there are still certain errors; column (c) shows that after multiple rounds of optimization, the pseudo-label is closer to the true label, the classification boundary is more accurate, and the overall prediction quality has improved; column (d) is the ground truth label, which is used as a reference standard to evaluate the effect of pseudo-label optimization and can intuitively compare the matching degree between the final pseudo-label and the ground truth. Finally, using the pseudo-label as the ground truth label to train the network, fully unsupervised training is achieved.

[0124] Figure 3 This is the result graph of seawater aquaculture extraction from SAR images. Among them, column (a) is the SAR image, which is the original image, column (b) is the ground truth label, and column (c) is the network prediction result. Among them, the first row of each column is the selected high-score No. 3 SAR satellite image sample 1, the second row of each column is the selected high-score No. 3 SAR satellite image sample 2, the third row of each column is the selected Radarsat-2 SAR satellite image sample 1, and the fourth row of each column is the selected Radarsat-2 SAR satellite image sample 2. Analysis Figure 3 From columns (a) to (c) in it, it can be seen that: the segmentation results obtained by using the pseudo-label as the ground truth to train the network show the segmentation effects of different SAR satellite image samples. The first and second rows of each column are high-resolution images, and the third and fourth rows of each column are lower-resolution images. Through pseudo-label optimization training, the model can more effectively extract the features of the target area and reduce misjudgments and missed detections. The optimized pseudo-label enhances the adaptability of the model to SAR images with different resolutions, making the segmentation results more accurate and the boundaries clearer.

[0125] Finally, it should be noted that the above examples are only used to illustrate the method solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that: modifying the method solutions recorded in the foregoing examples, or equivalently replacing some or all of the method features therein, does not make the essence of the corresponding method solutions deviate from the scope of the method solutions of the present invention in each example.

Claims

1. A method for extracting SAR aquaculture information from a fully unsupervised Transformer network, characterized in that, The SAR seawater aquaculture information extraction method includes two stages: upstream self-supervised pre-training and downstream unsupervised pseudo-label training. First, in the upstream stage, a large amount of unlabeled SAR image data is input into the self-supervised Transformer network for pre-training to obtain the trained network weights. The self-supervised transformer network is defined as the self-supervised Transformer network. Next, in the downstream stage, a globally optimal pseudo-label generator and a fine-grained discrimination and completion module are designed. Specifically: First, the network weights obtained from the upstream self-supervised pre-training are used to initialize the Teacher network, and the downstream training data is input into the Teacher network. The initial pseudo-labels of the training images are generated through the globally optimal pseudo-label generator. Second, the initial pseudo-labels are used as constraints to guide the training of the Student network. The image is segmented by the Student network, and its output result is combined with the features learned by the Teacher network and jointly input into the fine-grained discrimination and completion module. The initial pseudo-labels are finely optimized through the fine-grained discrimination and completion module to refine the detailed features of the aquaculture target and enhance its continuous information, generating optimized pseudo-labels. Finally, the optimized pseudo-labels are used to replace the initial pseudo-labels to constrain the segmentation network, and the output of the segmentation network in turn optimizes the pseudo-labels. Through the iterative optimization mechanism of the pseudo-labels, completely unsupervised SAR seawater aquaculture information extraction is achieved, including the following steps: Step 1: Construct an unlabeled seawater aquaculture dataset to obtain an upstream self-supervised pre-training dataset and a downstream unsupervised pseudo-label training dataset. Step 2: In the upstream pre-training stage, use the self-supervised Transformer network to train the upstream self-supervised pre-training dataset. Step 2.1, the self-supervised Transformer network consists of and combined, the network parameters of are the network parameters of are and the two output probabilities of and are respectively Step 2.2, update the parameters of ; Step 2.3, using network parameters of to update network parameters of ; Step 2.4, after the upstream self-supervised Transformer network training is completed, freeze the parameters of ; In the third step, downstream unsupervised pseudo-label training is carried out. The downstream training network consists of a Teacher network and a Student network. First, freeze the parameters from the second step as the initial weights of the Teacher network. The weights of the Teacher network , then use the Teacher network to extract features from the images in the downstream unsupervised pseudo-label training dataset. After that, the features output by the Teacher network are input into the global optimal pseudo-label generator to generate initial pseudo-labels, and use the initial pseudo-labels as constraints to train the Student network. In the fourth step, the segmentation result of the Student network and the feature map output by the Teacher network are input into the fine-grained discrimination and completion module to further optimize the pseudo labels. The Student network consists of an encoder and a decoder. The encoder part is based on a multi-scale Transformer architecture and is used to extract the features of the input data. The decoder adopts a UPerHead structure, including a feature fusion and a multi-scale information aggregation unit. The initial parameters of the Student network are ; Step 5: Establish the objective function and network iterative optimization process of the downstream unsupervised pseudo-label training network.

2. The method for extracting SAR seawater aquaculture information of a completely unsupervised Transformer network according to claim 1, characterized in that, The specific content of the first step is as follows: Step 1.1: Obtain SAR remote sensing image data from the satellite resource platform and preprocess the SAR remote sensing image. The preprocessing includes radiometric calibration, image enhancement, and terrain correction. Step 1.2: Filter the SAR remote sensing image processed in Step 1.

1. Step 1.3: Crop the filtered SAR remote sensing image into a fixed size to generate a SAR seawater aquaculture dataset, and divide it to obtain an upstream self-supervised pre-training dataset and a downstream unsupervised pseudo-label training dataset. Divide the downstream unsupervised pseudo-label training dataset into a training set and a test set, and make labels for the test set.

3. The method for extracting SAR aquaculture information of a completely unsupervised Transformer network according to claim 1, characterized in that The specific content of the second step is as follows: Step 2.1, specifically: First, for each image in the upstream self-supervised pre-training dataset perform random cropping to obtain a global view and a local view ; input the global view into for training, input the local view into for training, and then obtain and with the two output probabilities being and ; Step 2.2, calculate using Equation (1) and 's cross-entropy loss, and as shown in Equation (2), use the calculated cross-entropy loss as the 's loss function for backpropagation to update the parameters , Equations (1) and (2) are as follows: (1); (2); Among them, represents the length of the output probability distribution vector of the upstream pre-trained model; represents the prediction probability on categories; represents the prediction probability on the th category; represents the logarithm of the prediction probability on the th category; represents the gradient of the cross-entropy loss function with respect to the parameter for optimizing parameters; is the gradient of the cross-entropy loss with respect to the output probability, representing the error between the probability of and the probability of; Step 2.3, at the same time, using the method of exponentially weighted moving average to utilize the network parameters of to update the network parameters of , as shown in formula (3): (3); Among them, is the smoothing factor of EMA, which is used to control the influence degree of historical parameters.

4. A method for extracting SAR aquaculture information of a completely unsupervised Transformer network according to claim 3, characterized in that, The specific content of the third step is as follows: Step 3.1, using the parameters of after upstream pre-training and freezing to initialize the parameters of the Teacher network. The Teacher network consists of N multi-head self-attention modules, and each multi-head self-attention module consists of H self-attention heads to perform feature extraction and global modeling on the input data. In the multi-head self-attention module, there are queries , keys and values . In the last layer of the multi-head self-attention module, there is a special classification vector used to mark the overall classification information; Step 3.2, define the training images of the downstream unsupervised pseudo-label training dataset as , and divide into P image patches, input them into the Teacher network, and perform feature extraction through N multi-head self-attention modules to obtain the last-layer multi-head self-attention feature representation , , where the image patch feature corresponds to the feature of the -th image patch in the input image. , where and respectively represent the query and key of the -th image patch in the -th self-attention head; Step 3.3, input the last-layer multi-head self-attention features obtained in Step 3.2 into the global optimal pseudo-label generator, and calculate the classification vector using Equation (4) and the affinity vector between the classification vector and the patch features is . Then, calculate the affinity matrix between the patch features and the patch features using Equation (5). Equations (4) and (5) are as follows: (4); (5); Among them, represents the affinity vector between the classification vector and the image patch feature under the self-attention head ; is the dimensional space of the calculation result, where is the number of image patches, ; represents the query vector of the classification vector in the self-attention head ; represents the transpose of the key of the th image patch in the self-attention head ; represents the currently calculated self-attention head; represents the similarity matrix between image patch features under the self-attention head ; is the dimensional space of the calculation result; represents the transpose of the key of the th image patch in the self-attention head Step 3.4, take the affinity vector with the highest affinity in it and define it as the key image patch , ; perform the calculation on each self-attention head in turn to obtain the key image patch set , where is the number of self-attention heads; then, diffuse the key image patches of each self-attention head to the global image patch set , the global image patch set represents all the image patches with high similarity to the key image patches in the -th self-attention head, as shown in Equation (6): (6); Among them, represents the index of the currently calculated image patch; represents the set of image patch indices, which is the number of image patches; represents in the th self-attention head, the similarity between the key image patch and the image patch , extracted from the similarity matrix . Step 3.5, using the -means clustering algorithm, cluster the global image patch sets of each self-attention head obtained in Step 3.4 into two categories: aquaculture and sea water. Define the set of image patches with the category of aquaculture as and the set of image patches with the category of sea water as . Calculate the pseudo-labels at the image patch level for each self-attention head using formula (7) , where represents the th image patch, and the pseudo-label at the image patch level of the th self-attention head. After generating the pseudo-label for each image patch, the pseudo-label at the whole image level can be stitched together according to the position of the image patch in the whole image . is the pseudo-label at the whole image level obtained by the th self-attention head; thus, the initial pseudo-label set is obtained. , is the number of self-attention heads, and formula (7) is as follows: (7); Step 3.6, for the initial pseudo-label set obtained in Step 3.5 perform screening, and calculate the density value of the whole-image-level pseudo-label corresponding to each attention head using formula (8) Then, use formula (9) to calculate the image-level pseudo-label with the largest density value as the final initial pseudo-label , and formulas (8) and (9) are as follows: ​ (8); (9); Among them, is the number of image blocks; is the image block average distance within the category to which it belongs; is the image block average distance from image blocks of different categories; means taking and the maximum value of; the density value ranges from [-1, 1].

5. The method for extracting SAR seawater aquaculture information of a completely unsupervised Transformer network according to claim 4, wherein The specific content of the fourth step is as follows: Step 4.1, the images of the downstream unsupervised pseudo-label training dataset Through data augmentation methods such as rotation, scaling, and linear stretching, obtain the augmented images , and input the augmented images into the encoder of the Student network to extract multi-scale features; The extracted multi-scale features are input into the decoder. After successive upsampling and feature fusion operations, a feature map is output , and then, the feature map is binarized after convolution processing to obtain the output segmentation result ; Meanwhile, the enhanced image is input into the Teacher network. The Teacher network learns the correlation between different spatial positions based on the multi-head self-attention module to generate a saliency feature map . The saliency feature map is convolved and binarized, and finally, the saliency feature result is output ; Step 4.2, the segmentation result of the Student network and the saliency feature result output by the Teacher network are input into the fine-grained discrimination and completion module; First, use the Canny edge detection operator to extract the boundary information of the targets in the segmentation result , and the boundary information is represented in the form of a set, and each element in the set represents an independent target boundary in the segmentation result, ; According to the coordinates of the four points of the upper left corner, upper right corner, lower left corner, and lower right corner of each boundary , a set of minimum bounding boxes is generated , and each minimum bounding box is shown in formula (10): (10); Among them, represents the minimum bounding box of the th target boundary; the bounding box contains all pixel points that meet the conditions ; represents a pixel point in the image, whose abscissa is , and the ordinate is ; represents the minimum value of the abscissas of all pixel points in the boundary ; represents the maximum value of the abscissas of all pixel points in the boundary ; represents the minimum value of the ordinates of all pixel points in the boundary ; represents the maximum value of the abscissas of all pixel points in the boundary ; represents the number of aquaculture areas in the image, which is also the number of boundaries; Step 4.3, the number of boundaries and their corresponding bounding boxes in each image is determined by the number of targets in the image and is represented in a set form; the following is obtained using formula (11) the features within the bounding box are obtained using formula (12) the features within the bounding box are obtained using formula (12) The formulas (11) and (12) are as follows: ​ (11); (12); Among them, contains all pixel values that satisfy ; ; contains all pixel values that satisfy ; ; Step 4.4, from Step 4.3, the set of feature extraction results can be obtained, specifically: the feature set of the segmentation result output by the Student network within the bounding box and the feature set of the saliency result output by the Teacher network within the bounding box , where is the number of bounding boxes obtained in Step 4.2; in the discrimination stage of the fine-grained discrimination and completion module, by inputting and features within the same bounding box, IoU calculation is performed; Step 4.4, the fine-grained discrimination and completion module operates on the features within each bounding box independently. When and the IoU between them exceeds the IoU threshold , the features between and are fused; otherwise, the features of are retained until the completion operation is finished for each box, and finally the optimized pseudo-label is obtained, as shown in formulas (14) and (15): (14); (15); Among them, represents the feature result obtained from the bounding box; represents the fusion between features; represents the IoU threshold.

6. The method for extracting SAR seawater aquaculture information of a completely unsupervised Transformer network according to claim 5, characterized in that In the fourth step: In Step 4.3, the IoU calculation is shown in Equation (13): (13); Among them, represents performing IoU calculation on and ; represents calculating the intersection between and pixels; represents calculating the union between and pixels. In step 4.3, the threshold value ranges from [0, 1].

7. A method for extracting SAR seawater aquaculture information of a completely unsupervised Transformer network according to claim 5, characterized in that, The specific content of the fifth step is as follows: Step 5.1, calculate the initial pseudo-labels generated by the global optimal pseudo-labels using formula (16) and the segmentation result output by the Student network The cross-entropy loss between ; Similarly, use formula (17) to calculate the optimized pseudo-labels and the segmentation result output by the Student network The cross-entropy loss between , Formulas (16) and (17) are as follows: (16); (17); Among them, represents the total number of pixels of the downstream unsupervised pseudo-label training images ; Step 5.2, the overall objective function of the Student network is , and are coefficients; taking the objective function as the loss of the Student network, use Equation (18) to reversely update the weight parameters of the Student network ; at the same time, use Equation (19) to update the weight parameters of the Teacher network by means of exponential weighted moving average using the weight parameters of the Student network . Equation (18) and Equation (19) are as follows: (18); (19); wherein, is the learning rate; is the parameter of the loss function with respect to the network; is the smoothing factor of EMA, used to control the influence degree of historical parameters; Step 5.3, when steps 3 to 5.2 are completed, one round of training is finished; the training of the downstream unsupervised pseudo-label training network affects the Student network and the Teacher network through pseudo-labels, and at the same time, the pseudo-labels are optimized by the prediction results of the Student network and the Teacher network, so as to realize the iterative update of the pseudo-labels; when the number of training rounds is reached, the downstream unsupervised pseudo-label training stops further training, and the segmentation output by the Student network is used as the final output of the network.

8. A method for extracting SAR aquaculture information of a completely unsupervised Transformer network according to claim 7, characterized in that, In the said step 5.2, and are related to the number of training rounds of the downstream unsupervised pseudo-label training network When then ; when then .

Citation Information

Patent Citations

  • A SAR image semantic segmentation method based on depth convolution network and weak supervised learning

    CN109344837A

  • Mariculture area extraction method based on deep learning

    CN112766155A

  • Semantic enhancement feature fusion self-supervised transform-based whole-scene SAR mariculture multi-target extraction method

    CN117036934A

  • Transform-based efficient defogging semantic segmentation method and application thereof

    CN117058024A