An Unsupervised Deep Learning Method for Intelligent Extraction of Seawater Aquaculture in SAR Images
By building a dual-deep learning network, superpixel blocks are used to enhance semantic information and generate pseudo-labels, the problem of extracting semantic information in seawater farming without labels is solved, and high-precision intelligent extraction of seawater farming is achieved, avoiding coherent spot noise interference.
Patent Information
- Application Number
- CN202210567764.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-24
AI Technical Summary
The prior art is difficult to build an unsupervised deep learning network in the absence of completely unlabeled conditions, effectively extract semantic information of seawater aquaculture, and is susceptible to coherent spot noise in SAR images.
The dual-deep learning network structure is adopted to enhance semantic information through superpixel blocks and generate pseudo-labels, so that the dual-networks are constantly iterated alternately and updated pseudo-labels, and unsupervised deep learning high-precision seawater aquaculture extraction is realized.
The extraction accuracy of traditional unsupervised methods is improved, coherent spot noise interference in SAR images is avoided, and the high-precision effect of unsupervised deep learning in intelligent sea aquaculture extraction is achieved.
Smart Images

Figure CN115035293B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - technical field of marine remote sensing and artificial intelligence, and relates to a method for intelligent extraction of seawater aquaculture from SAR images based on unsupervised deep learning. Background Technique
[0002] China is a country with developed seawater aquaculture in the world, ranking first in the world in terms of both aquaculture area and total output. Driven by economic interests, in many areas, seawater aquaculture has developed blindly and disorderly. Large - scale reclamation has caused a reduction in the sea area and a decrease in the tidal prism, weakening the self - purification ability of the ocean and exacerbating the deterioration of the water environment. To reasonably plan the aquaculture area, satellite remote sensing has been widely used in the extraction of aquaculture information. Floating raft aquaculture has clear and obvious characteristics under visible light, but visible light is easily interfered by natural conditions such as clouds, rain, and snow, resulting in a large - scale loss of information. Based on Synthetic Aperture Radar (SAR) remote sensing images, they have the ability to observe the ground all - day and all - weather, and can overcome the interference of clouds on aquaculture extraction. Therefore, it is of great significance to intelligently extract aquaculture based on SAR images.
[0003] Most of the existing seawater aquaculture extraction algorithms are based on supervised methods. The general process is: select samples and make labels, train a classifier, and finally predict the results. This method requires a large number of labeled samples. It is difficult to obtain remote sensing data labels, and the sea conditions are complex and the targets change greatly, resulting in too high label costs. The traditional unsupervised methods that formulate classification rules from the data itself are easily affected by speckle noise in SAR images, with low accuracy, and cannot extract effective aquaculture semantic information. There has been no relevant research on the extraction of seawater aquaculture information from SAR images using unsupervised deep learning methods. Therefore, it is necessary to build an unsupervised deep learning model for seawater aquaculture in long - sequence SAR images to achieve intelligent extraction of seawater aquaculture in a deep learning network without any labeled samples. Summary of the Invention
[0004] The present invention mainly solves the problem of how to build an unsupervised deep learning network in the case of completely no labels, enable it to learn seawater aquaculture semantic information, and overcome the interference of speckle noise in SAR images. A method for intelligent extraction of seawater aquaculture from SAR images based on unsupervised deep learning is proposed. A dual deep learning network structure is adopted, semantic information is enhanced through super - pixel blocks, pseudo - labels are generated, and the dual networks are continuously iterated alternately to update the pseudo - labels, realizing high - precision extraction of seawater aquaculture by unsupervised deep learning.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] An intelligent extraction method for seawater aquaculture in SAR images based on unsupervised deep learning. This method aims to solve problems such as the difficulty of obtaining labeled samples for remote sensing images of seawater aquaculture, the inability of traditional unsupervised methods to avoid the interference of speckle noise in SAR images, and how to build an unsupervised deep learning network model and obtain semantic information of seawater aquaculture. It generates pseudo-labels containing semantic information of seawater aquaculture through block judgment and feature extraction networks, enables the fully convolutional semantic segmentation network to learn its semantic information, avoids the interference of speckle noise, and constructs an unsupervised deep learning network model through the alternating iteration of two deep learning networks. The method includes the following steps:
[0007] First step, process the original image using three image enhancement methods: linear truncation stretching, gamma transformation, and Gaussian filtering. All three image enhancement methods process the original image:
[0008] 1.1) Process the original image using linear truncation stretching. Linear truncation stretching is one of the most commonly used methods in remote sensing image enhancement processing. Enhance the image contrast by setting three different truncation values. As shown in formula (1), read the maximum and minimum values of the original image, then perform a histogram statistics on the original image, find the gray values of the original image corresponding to the truncation values. For example, if the truncation value is set to 2, find the gray values corresponding to 2% and 98%, and use these gray values as the maximum and minimum values of the output image. Finally, linearly stretch the original image into the output image. In this way, the gray values greater than 98% in the original image are replaced by the maximum value of the output image, and the gray values less than 2% in the original image are replaced by the minimum value of the output image. The aquaculture area has a stronger backscattering coefficient compared to the seawater area, but is mixed with speckle noise. Through linear truncation stretching, the speckle (outlier) noise problem can be improved.
[0009]
[0010] Where: A l is the output image of linear truncation stretching; O is the input image; d m and c m are the maximum and minimum values of the output image respectively, and b m and a m are the maximum and minimum values of the original image respectively.
[0011] 1.2) In SAR images, the microwave has a very weak penetration ability into seawater. When there are wind and waves, it may submerge the aquaculture facilities, resulting in a weakened backscattering of the submerged aquaculture area. Therefore, gamma transformation is introduced to enhance the dark details of the original image. As shown in formula (2), through non-linear transformation, the gray values of the darker areas in the image are enhanced.
[0012]
[0013] Among them, A y is the output value after gamma transformation; c is the gray-scale scaling coefficient; γ is the gamma factor size.
[0014] 1.3) In order to obtain a better image edge of the aquaculture area for subsequent extraction, Gaussian filtering is used to smooth the original image. The formula is shown in (3). By Gaussian filtering, Gaussian noise is eliminated.
[0015]
[0016] Among them, A g is the output image after Gaussian filtering, and σ is the standard deviation.
[0017] The original image is respectively processed by the above image enhancement methods. Among them, three different truncation values are set for linear truncated stretching, and a total of five different enhanced output images are obtained. And the output images after image enhancement are used as the inputs of the feature extraction network in the third step and the fully convolutional semantic segmentation network in the fourth step.
[0018] Second step, since aquaculture information is extracted under unsupervised conditions, some prior knowledge needs to be provided before the output image enters the network. The edge information of the aquaculture area and the pixel value difference between seawater and aquaculture are obtained by traditional unsupervised methods. The superpixel segmentation algorithm is used to divide the image into irregular pixel blocks with certain visual significance composed of adjacent pixels with similar texture, color, brightness and other characteristics. The superpixel algorithm used is Simple Linear Iterative Clustering (SLIC), which has a high comprehensive evaluation in terms of operation speed, object contour preservation and superpixel shape. The SLIC superpixel algorithm is as follows:
[0019] 2.1) Convert the original image to the color space of the International Commission on Illumination color model (Commission Internationale de L'Eclairage Lab, CIELAB), set the initial number of clustering centers k and the distance measurement D, and assign each pixel within the restricted area to the nearest clustering center. The formula is as follows:
[0020]
[0021] Among them, l i 、a i 、b i respectively represent the values of pixel point i in the L channel, a channel and b channel, and d c is the distance between the pixel point and the CIELAB value of the clustering center; x and y respectively represent the horizontal and vertical coordinates of the pixel point in the image, and ds is the distance of the spatial position between the pixel point and the clustering center; is the expected superpixel block size, determined by the picture size N and the number k of initial clustering centers; m is a constant that determines the maximum color distance.
[0022] 2.2) After assigning each pixel point to the nearest clustering center, take the average in the l, a, b, x, and y channels of the pixel points under the same clustering center to update the clustering center. Calculate the residual error of the new and previous clustering center positions according to the L2 norm, and the formula is as follows:
[0023]
[0024] 2.3) Iteratively repeat the pixel assignment and clustering center update until the error converges or a certain number of iterations is reached. Through practice, it is found that 10 iterations can obtain relatively ideal results for most pictures. In practice, generally, the iteration is fixed at 10 times.
[0025] After obtaining the result map of superpixel segmentation through the SLIC algorithm, the aquaculture contour is successfully detected, and the entire image is segmented into multiple small blocks. The edges of each small block are relatively consistent with the edges of the aquaculture area. However, there are also many misjudgments, the accuracy is low, and only the edges are detected, and no aquaculture extraction result is obtained. It is necessary to retain its edge information, that is, the positions of the pixel points included in each small block, and further process it in the third step to generate pseudo-labels.
[0026] In the third step, in a completely unsupervised situation, pseudo-labels are required as the target of the deep learning network. Use the five images obtained in the first step, perform channel transformation on them, and each image serves as an input channel. Finally, the input of the deep learning network is a five-channel enhanced image.
[0027] 3.1) Build a feature extraction network model, and obtain the depth features of the aquaculture area through the feature extraction network. The feature extraction network is mainly composed of a convolutional layer, a ReLU layer, and a BN layer. Take the convolutional layer, ReLU layer, and BN layer as an overall module. The three layers in the overall module are connected end to end, and the number of modules can be adjusted arbitrarily as the feature extraction network. For example, after the input image enters the convolutional layer, ReLU layer, and BN layer module, the output value enters a new convolutional layer, ReLU layer, and BN layer module again. Among them, the convolutional layer formula is as shown in (6). The output image in the first step serves as the input of the convolutional layer in the first module, and the input of the convolutional layer in the remaining modules is the output of the BN layer in the previous module.
[0028]
[0029] Among them, a c is the output of the convolutional layer; W gis the weight of the g-th convolution kernel; x i is the i-th input; b i is the g-th bias; I is the total number of inputs; G is the total number of convolution kernels.
[0030] In order to make the feature extraction network converge quickly, a ReLU layer is added after the convolution layer. The formula is shown in (7), that is, the output of the convolution layer is used as the input of the ReLU layer.
[0031] a r =max(0,a g )(7)
[0032] Among them, a r is the output of the ReLU layer.
[0033] In order to prevent gradient vanishing and overfitting during feature extraction network training, a BN layer is added after the ReLU layer. The formula is shown in (8), that is, the output of the ReLU layer is used as the input of the BN layer.
[0034]
[0035] Among them, ε is the minimum value automatically generated by the system.
[0036] 3.2) The output of the last layer of the feature extraction network model is connected to the softmax function. The formula is shown in (9), that is, the output of the Nth convolutional layer, ReLU layer and BN layer modules is used as the input of the softmax function.
[0037]
[0038] in, is the output of the Nth convolutional layer, ReLU layer, and BN layer module.
[0039] 3.3) The output value can be converted into a probability distribution in the range of [0,1] and with a sum of 1 through the softmax function. In order to distinguish between seawater and aquaculture, it is converted into a binary classification problem, and the number of output channels of the softmax function is set to 2, each channel represents the probability of the seawater area and the probability of the aquaculture area. In order to obtain a pseudo-label, the output probability of the 2 channels is converted into a single channel index value, and the softmax output is indexed by the channel maximum value, and the channel index value of the maximum value is returned. In this way, there are only 0 or 1 index values on the single-channel image, which can be used as a pseudo-label.
[0040] 3.4) Due to the speckle noise in SAR images, the pseudo-labels do not have good semantic information, and the aquaculture areas may be mixed with sea water labels. Therefore, the position information of the superpixel blocks obtained in the second step is used to perform block judgment operations. The detailed explanation is as follows. The superpixel segmentation blocks are corresponding to the pseudo-labels (channel maximum index) just obtained. In each small block of the pseudo-label map (the small blocks come from the superpixel block position information), the dominant proportion is judged, and each small block in the label map is replaced with its dominant label. This ensures that there is a unified pseudo-label in the small block, avoids the interference of speckle noise, enhances the spatial consistency, and generates pseudo-labels containing aquaculture semantic information.
[0041] Step 4: To make full use of the pseudo-labels containing aquaculture semantic information generated in the previous step, a fully convolutional semantic segmentation network model is built.
[0042] 4.1) The fully convolutional semantic segmentation network is used to obtain the extraction results of the seawater aquaculture areas. The pseudo-labels with aquaculture semantic information obtained in the third step are used as the targets for training, and the objective function L U-Net is shown as follows:
[0043]
[0044] where x is the input image, N is the number of input images, SLIC(·) is the superpixel segmentation function, PE is the block judgment method, f(·) is the feature extraction network model, and g(·) is the fully convolutional semantic segmentation network model.
[0045] In this way, through the fully convolutional semantic segmentation network, the aquaculture semantic information in the pseudo-labels containing aquaculture semantic information is further strengthened and learned, the extraction accuracy is improved, and the aquaculture areas ignored in the dominant proportion judgment of the block judgment are improved.
[0046] 4.2) To gradually improve the extraction accuracy of the dual network in the unsupervised case, the pseudo-labels need to be continuously updated and optimized. Therefore, the cross-entropy loss is calculated between the extraction results of the seawater aquaculture areas generated by the fully convolutional semantic segmentation network and the pseudo-labels output by the feature extraction network in the third step, and the generation process of the pseudo-labels is optimized by backpropagation (that is, the pseudo-labels are optimized through the extraction results of the seawater aquaculture areas to achieve alternating optimization, that is, the pseudo-labels optimize the extraction results, and the extraction results in turn optimize the pseudo-labels, and the cycle continues). The objective function L FEN is as follows:
[0047]
[0048] In this way, the pseudo-labels of the feature extraction network can be optimized by the fully convolutional semantic segmentation network, and then the optimized pseudo-labels are used as the targets of the fully convolutional semantic segmentation network again for the second round of iteration, completing the alternating iteration of the dual networks, that is, fixing the pseudo-labels of the feature extraction network to optimize the fully convolutional semantic segmentation network, or fixing the aquaculture extraction results of the fully convolutional semantic segmentation network to optimize the feature extraction network. The alternating update of the pseudo-labels and the seawater aquaculture area extraction results is realized, and a fully convolutional semantic network with stronger generalization ability and higher accuracy is gradually generated. When the pseudo-labels generated by the feature extraction network and the block operation no longer change, the alternating iteration of the dual networks stops, and only the fully convolutional semantic segmentation network is trained until the network converges.
[0049] The beneficial effects of the present invention are as follows:
[0050] The present invention improves the extraction accuracy of traditional unsupervised methods, avoids the interference of speckle noise in SAR images, and provides an unsupervised deep learning method for the intelligent extraction of seawater aquaculture areas in SAR images. In deep learning convolutional neural networks, an explicit loss target, that is, a true value label, is often required. The present invention uses a feature extraction network to generate simple pseudo-labels, and proposes a block operation for the problem of seawater aquaculture extraction. While avoiding the speckle noise in SAR images, the simple pseudo-labels also contain aquaculture semantic information. At the same time, in order to further strengthen this semantic information and improve the aquaculture extraction accuracy, the present patent proposes a dual network structure, adding a fully convolutional semantic segmentation network. While enhancing the aquaculture semantics, the results can also be used to reverse-optimize the feature extraction network and thus optimize the pseudo-labels. Moreover, the alternating update of the dual networks provides a direction for the optimization of the pseudo-labels, making the pseudo-labels closer and closer to the true value, and realizing high-precision aquaculture extraction of unsupervised deep learning. The method proposed by the present invention has high accuracy and meets the feasibility of long-sequence SAR image seawater aquaculture monitoring. Description of the Drawings
[0051] Figure 1 It is the overall block diagram of an intelligent extraction method for seawater aquaculture in SAR images based on unsupervised deep learning;
[0052] Figure 2 It is a schematic diagram of the result of the superpixel segmentation algorithm. (a), (b), and (c) are three different aquaculture areas respectively;
[0053] Figure 3 It is the pseudo-labels without semantic information obtained by the feature extraction network. (a), (b), and (c) are three different aquaculture areas respectively;
[0054] Figure 4 It is the pseudo-labels with semantic information obtained by the block operation method. (a), (b), and (c) are three different aquaculture areas respectively;
[0055] Figure 5 This is the intelligent extraction result of the method of this patent. (a), (b), and (c) are three different aquaculture areas respectively;
[0056] Figure 6 This is the detailed intelligent extraction result of the GF-3 satellite for raft aquaculture. (a) is the original image of Area 1, (b) is the ground truth image of Area 1, (c) is the result of the method of this patent, and the accuracy Accuracy = 90.97%;
[0057] Figure 7 This is the detailed extraction result of the RADARSAT-2 satellite for cage aquaculture. (a) is the original image, (b) is the ground truth image, (c) is the result of the method of this patent, and the accuracy Accuracy = 92.34%. Specific implementation manner
[0058] To make the method problems solved by the present invention, the adopted method solutions and the achieved method effects clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention are shown in the drawings, rather than all the content.
[0059] As Figure 1 shown, a method for intelligent extraction of seawater aquaculture from SAR images based on unsupervised deep learning provided by an embodiment of the present invention includes:
[0060] Compile in python3.6.12, pytorch1.7.1 and cuda11.0 under the windows10 system, run using the GPU of RTX 3080, and input SAR images with a size of 256×256.
[0061] In the first step, three image enhancement methods, namely linear truncation stretching, gamma transformation, and Gaussian filtering, are adopted. In addition, all three image enhancement methods are used to process the original image:
[0062] 1.1) The original image is processed by linear truncated stretching, which is one of the most commonly used methods in remote sensing image enhancement processing. By setting different truncation values to enhance the image contrast, as shown in formula (1), the maximum and minimum values of the original image are read, and then the histogram of the original image is statistically analyzed to find the gray values of the original image corresponding to the truncation values. For example, if the truncation value is set to 2, the gray values corresponding to 2% and 98% are found and used as the maximum and minimum values of the output image. Finally, the original image is linearly stretched into the output image. In this way, the gray values greater than 98% in the original image are replaced by the maximum value of the output image, and the gray values less than 2% in the original image are replaced by the minimum value of the output image. The aquaculture area has a stronger backscattering coefficient than the seawater area, but is mixed with speckle noise. Through linear truncated stretching, the speckle (outlier) noise problem can be improved.
[0063]
[0064] Where: A l is the output image of linear truncated stretching; O is the input image; d m and c m are the maximum and minimum values of the output image respectively, and b m and a m are the maximum and minimum values of the original image respectively.
[0065] Actually, for linear truncated stretching, the parameters that need to be set are only the truncation value and the output image. This patent provides three different truncation values of 2, 5, and 7 for seawater aquaculture extraction, and different intensities of speckle noise are improved through different truncation values.
[0066] 1.2) In SAR images, the microwave penetration ability into seawater is very weak. When there are wind and waves, it may submerge the aquaculture facilities, resulting in a weakened backscattering of the submerged aquaculture area. Gamma transformation is introduced to enhance the details of the dark part of the original image, as shown in formula (2). Through non-linear transformation, the gray values of the darker areas in the image are enhanced.
[0067]
[0068] Where, A y is the output value after gamma transformation; c is the gray scaling coefficient, usually taken as 1; γ is the size of the gamma factor, which controls the scaling degree of the whole transformation. The gamma factor selected in this patent is 0.5, and the gray scaling coefficient is 1.
[0069] 1.3) In order to obtain a better image edge of the aquaculture area for subsequent extraction, Gaussian filtering is used to smooth the original image, and the formula is as shown in (3). Through Gaussian filtering, Gaussian noise is eliminated.
[0070]
[0071] Among them, A g is the output image after Gaussian filtering, σ is the standard deviation, which controls the smoothing effect, and the selected standard deviation is 2.
[0072] The original image is respectively processed by the above image enhancement methods. Among them, three different truncation values (2, 5, 7) are set for linear truncated stretching, and a total of five different enhanced output images are obtained. And the output images after image enhancement are used as the inputs of the feature extraction network in the third step and the fully convolutional semantic segmentation network in the fourth step.
[0073] In the second step, since the aquaculture information is extracted under unsupervised conditions, some prior knowledge needs to be provided before the output image enters the network. The edge information of the aquaculture area and the pixel value difference between seawater and aquaculture are obtained through traditional unsupervised methods. The traditional unsupervised method uses the superpixel segmentation algorithm to divide the image into irregular pixel blocks with certain visual significance composed of adjacent pixels with similar texture, color, brightness and other characteristics. The superpixel algorithm used is Simple Linear Iterative Clustering (SLIC), which has a high comprehensive evaluation in terms of operation speed, object contour preservation and superpixel shape. The SLIC superpixel algorithm is as follows:
[0074] 2.1) Convert the original image to the color space of the Commission Internationale de l'Eclairage Lab (CIELAB) color model, set the initial number of clustering centers k and the distance measurement D, and assign each pixel within the restricted area to the nearest clustering center. The formula is as follows:
[0075]
[0076] Among them, l i 、a i 、b i respectively represent the values of pixel point i in the L channel, a channel and b channel, and d c is the distance between the pixel point and the CIELAB value of the clustering center; x and y respectively represent the horizontal and vertical coordinates of the pixel point in the image, and d s is the distance of the spatial position between the pixel point and the clustering center; is the expected size of the superpixel block, which is determined by the picture size N and the initial number of clustering centers k = 30; m = 5 is a constant that determines the maximum color distance.
[0077] 2.2) After each pixel is assigned to the nearest cluster center, the average values of the l, a, b, x, and y channels of the pixels under the same cluster center are taken to update the cluster center. Calculate the residual error of the position of the new and previous cluster centers according to the L2 norm, and the formula is as follows:
[0078]
[0079] 2.3) The assignment of pixels and the update of the cluster center are iterated repeatedly until the error converges or a certain number of iterations are reached. Through practice, it is found that 10 iterations can achieve relatively ideal results for most images. In practice, the number of iterations is generally fixed at 10 times. This patent stipulates that the SLIC superpixel algorithm stops after 10 iterations.
[0080] After obtaining the result map of superpixel segmentation through the SLIC algorithm, the result is as Figure 2 shown. The breeding contour is successfully detected, and the entire image is segmented into multiple small blocks. The edges of each small block are relatively consistent with the edges of the breeding area. However, there are also many misjudgments, the accuracy is low, and only the edges are detected, and the breeding extraction result is not obtained. It is necessary to retain its edge information, that is, the position of the pixels included in each small block, and further process it in the third step to generate pseudo-labels.
[0081] In the third step, in a completely unsupervised situation, pseudo-labels are required as the targets of the deep learning network. Using the five images obtained in the first step, perform channel transformation on them, and each image serves as an input channel. Finally, the input of the deep learning network is an enhanced image with five channels.
[0082] 3.1) Build a feature extraction network model. The depth features of the breeding area obtained through the feature extraction network provide prior knowledge for the generation of pseudo-labels. The feature extraction network is mainly composed of a convolutional layer, a ReLU layer, and a BN layer. Taking the convolutional layer, the ReLU layer, and the BN layer as an overall module T, the number of modules T = 5 can be adjusted as the feature extraction network. For example, after the input image enters the convolutional layer, ReLU layer, and BN layer module, the output value enters a new convolutional layer, ReLU layer, and BN layer module again. Among them, the convolutional layer formula is as shown in (6). The output image in the first step serves as the input of the convolutional layer in the first module, and the input of the convolutional layer in the remaining modules is the output of the BN layer in the previous module.
[0083]
[0084] Among them, a c is the output of the convolutional layer; W g is the weight of the g-th convolutional kernel; x i is the i-th input; b i is the g-th bias; I is the total number of inputs; G is the total number of convolutional kernels.
[0085] In order to make the feature extraction network converge quickly, a ReLU layer is added after the convolution layer. The formula is shown in (7), that is, the output of the convolution layer is used as the input of the ReLU layer.
[0086] a r =max(0,a g )(7)
[0087] Among them, a r is the output of the ReLU layer.
[0088] In order to prevent gradient vanishing and overfitting during feature extraction network training, a BN layer is added after the ReLU layer. The formula is shown in (8), that is, the output of the ReLU layer is used as the input of the BN layer.
[0089]
[0090] Among them, ε is the minimum value automatically generated by the system.
[0091] The weight size of each convolution kernel is W g is 3×3, the step size is 1, and the first module T=1 input x i The number of channels I = 5, the number of output channels G = 100, and then each input x i Output a from the previous module BN layer b It is found that the input I of the middle three modules is equal to the output channel number G, which is equal to 100. The input channel number of the fifth module is I = 100, the output channel number G = 2, and the offset Obtained by network training.
[0092] 3.2) The output of the last layer of the feature extraction network model is connected to the softmax function. The formula is shown in (9), that is, the output of the Nth convolutional layer, ReLU layer and BN layer modules is used as the input of the softmax function.
[0093]
[0094] in, is the output of the Nth convolutional layer, ReLU layer, and BN layer module.
[0095] 3.3) The output values can be converted into a probability distribution within the range of [0, 1] and with a sum of 1 through the softmax function. To distinguish seawater from aquaculture, it is converted into a binary classification problem, and the number of output channels of the softmax function is set to 2. Each channel represents the probability of the seawater area and the probability of the aquaculture area respectively. To obtain the pseudo-label, the output probabilities of the 2 channels are converted into single-channel index values. The maximum value index of the channels is performed on the softmax output, and the channel index value of the maximum value is returned. In this way, only 0 or 1 index values exist on the single-channel image, which can be used as the pseudo-label. The result is as Figure 3 shown.
[0096] 3.4) Due to the speckle noise in the SAR image, the pseudo-label does not have good semantic information, and the aquaculture area may be mixed with seawater labels. Then, using the position information of the superpixel blocks obtained in the second step, the block judgment operation is carried out. The detailed explanation is as follows. The superpixel segmentation blocks are corresponding to the pseudo-label (channel maximum value index) just obtained. In each small block of the pseudo-label map (the small block comes from the superpixel block position information), the dominant proportion judgment is carried out, and each small block in the label map is replaced with its dominant label. In this way, it is ensured that there is a unified pseudo-label in the small block, avoiding the interference of speckle noise, enhancing the spatial consistency, and generating a pseudo-label containing aquaculture semantic information. The result is as Figure 4 shown.
[0097] Fourth step, to make full use of the pseudo-label containing aquaculture semantic information generated in the previous step, a fully convolutional semantic segmentation network model is built. The selected semantic segmentation network is the U-Net network. The convolutional layer is used to extract depth information. It consists of four downsamplings and four upsamplings to form a U-shaped structure. The size of the convolutional kernel is 3×3, as shown in formula (7). The downsampling method selects max pooling to retain the largest value within the 2×2 area, and the upsampling uses interpolation for transposed convolution.
[0098] 4.1) The fully convolutional semantic segmentation network is used to obtain the extraction result of the seawater aquaculture area. Using the pseudo-label with aquaculture semantic information obtained in the third step as the target for training, the objective function L U-Net is as follows:
[0099]
[0100] where x is the input image, N is the number of input images, SLIC(·) is the superpixel segmentation function, PE is the block judgment method, f(·) is the feature extraction network model, and g(·) is the fully convolutional semantic segmentation network model.
[0101] In this way, through the fully convolutional semantic segmentation network, the seawater aquaculture semantic information in the pseudo-label containing aquaculture semantic information is further strengthened and learned, the extraction accuracy is improved, and the aquaculture area ignored in the dominant proportion judgment of the block judgment is improved. The result is asFigure 5 as shown
[0102] 4.2) To gradually improve the extraction accuracy of the dual network in an unsupervised situation, the pseudo-labels need to be continuously updated and optimized. Therefore, the cross-entropy loss is calculated between the extraction result of the mariculture area generated by the fully convolutional semantic segmentation network and the pseudo-labels output by the feature extraction network in the third step, and the generation process of the pseudo-labels is optimized by backpropagation. Its objective function L FEN is as follows
[0103]
[0104] In this way, the pseudo-labels of the feature extraction network can be optimized by the fully convolutional semantic segmentation network, and then the optimized pseudo-labels are used as the target of the fully convolutional semantic segmentation network again for the second round of iteration, completing the alternating iteration of the dual network, that is, fixing the pseudo-labels of the feature extraction network to optimize the fully convolutional semantic segmentation network, or fixing the mariculture extraction result of the fully convolutional semantic segmentation network to optimize the feature extraction network. The alternating update of the pseudo-labels and the mariculture area extraction result is realized, and a fully convolutional semantic network with stronger generalization ability and higher accuracy is gradually generated. When the pseudo-labels generated by the feature extraction network and the block operation no longer change, the alternating iteration of the dual network stops, and only the fully convolutional semantic segmentation network is trained until the network converges. Finally, the mariculture extraction results are as Figure 6 and Figure 7 shown, which are the mariculture extraction results on the GF-3 satellite and the RADARSAT-2 satellite respectively Figure 6 The mariculture extraction accuracy Accuracy = 90.97% Figure 7 The mariculture extraction accuracy Accuracy = 92.34%
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the method solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: modifying the method solutions recorded in the foregoing embodiments, or equivalently replacing some or all of the method features therein, does not make the essence of the corresponding method solutions deviate from the scope of the method solutions of the embodiments of the present invention
Claims
1. An intelligent extraction method for seawater aquaculture in SAR images using unsupervised deep learning, characterized in that: S1: 1.1) Enhance the image contrast by setting three different truncation values, read the maximum and minimum values of the original image, perform a histogram statistics on the original image, find the gray values of the original image corresponding to the truncation values, and use the gray values as the maximum and minimum values of the output image. Finally, linearly stretch the original image into the output image; Improve the speckle noise problem; Set three different truncation values to obtain three output images respectively; 1.2) In the SAR image, introduce gamma transformation to enhance the dark details of the original image. Through non-linear transformation, enhance the gray values of the darker regions in the image to obtain a fourth enhanced output image; 1.3) To obtain a better image edge of the aquaculture area, use Gaussian filtering to smooth the original image. Eliminate Gaussian noise through Gaussian filtering to obtain a fifth enhanced output image; The output image after image enhancement is used as the input of the feature extraction network in S3 and the fully convolutional semantic segmentation network in S4; S2. Before the output image enters the network, some prior knowledge needs to be provided. Obtain the edge information of the aquaculture area and the pixel value difference between seawater and aquaculture through traditional unsupervised methods. Adopt the superpixel segmentation algorithm to segment the image into irregular pixel blocks composed of adjacent pixels with similar texture, color, and brightness characteristics; be able to detect the aquaculture contour and segment the entire image into multiple small blocks. The edge of each small block is more consistent with the edge of the aquaculture area, and retain its edge information, that is, the position of the pixel points included in each small block, and further process it in S3 to generate pseudo-labels; S3. In a completely unsupervised situation, pseudo-labels are required as the target of the deep learning network; use the five images obtained in S1 for channel transformation, with each image as an input channel, and the input of the deep learning network is a five-channel enhanced image; 3.1) Build a feature extraction network model, and obtain the deep features of the aquaculture area through the feature extraction network. The feature extraction network is mainly composed of a convolutional layer, a ReLU layer, and a BN layer. Take these three layers as an overall module, and in the overall module, the convolutional layer, the ReLU layer, and the BN layer are connected end to end, and the number of modules can be adjusted arbitrarily as the feature extraction network; Among them, the output image in S1 is used as the input of the convolutional layer in the first module, and the input of the convolutional layer in the remaining modules is the output of the BN layer in the previous module; 3.2) Connect the output of the last layer of the feature extraction network model to the softmax function; 3.3) Convert the output value into a probability distribution with a range in [0,1] and a sum of 1 through the softmax function; Binary classification of seawater and aquaculture: Set the number of output channels of the softmax function to 2, and each channel represents the probability of the seawater area and the probability of the aquaculture area respectively; Convert the output probability of the 2 channels into a single-channel index value, perform a channel maximum index on the softmax output, and return the channel index value of the maximum value. In this way, only 0 or 1 index values exist on the single-channel image, and use it as the pseudo-label; 3.4) Use the position information of the superpixel blocks obtained in S2 to perform block judgment; correspond the superpixel segmentation blocks to the pseudo-labels obtained in 3.3), perform a dominant proportion judgment in each small block of the pseudo-label map, replace each small block in the label map with its dominant label, ensure that the small block has a unified pseudo-label, avoid speckle noise interference, and generate a pseudo-label containing aquaculture semantic information; S4. Build a fully convolutional semantic segmentation network model; 4.1) Use the fully convolutional semantic segmentation network to obtain the extraction result of the mariculture area, and use the pseudo-label with aquaculture semantic information obtained in S3 as the target for training; Strengthen and learn the mariculture semantic information in the pseudo-label with aquaculture semantic information through the fully convolutional semantic segmentation network; 4.2) Calculate the cross-entropy loss between the extraction result of the mariculture area generated by the fully convolutional semantic segmentation network and the pseudo-label output by the feature extraction network in S3, and perform backpropagation to optimize the generation process of the pseudo-label; The pseudo-label of the feature extraction network can be optimized by the fully convolutional semantic segmentation network, and then the optimized pseudo-label is used as the target of the fully convolutional semantic segmentation network for the second round of iteration, completing the alternating iteration of the double network, that is, fixing the pseudo-label of the feature extraction network to optimize the fully convolutional semantic segmentation network, and fixing the mariculture extraction result of the fully convolutional semantic segmentation network to optimize the feature extraction network; realize the alternating update of the pseudo-label and the mariculture area extraction result; when the pseudo-label generated by the feature extraction network and the block operation no longer changes, the alternating iteration of the double network stops, and only the fully convolutional semantic segmentation network is trained until the network converges.
Citation Information
Patent Citations
SAR image change detection method based on semi-supervised confrontation depth network
CN110263845A
Remote sensing image change detection method and device, electronic device and storage medium
CN113255451A