Multi-category culture area remote sensing extraction method and device based on improved UNet

By introducing DC Block and Stransformer modules into UNet networks, the downsampling and upsampling structure is optimized, and the problem that traditional remote sensing technology is difficult to accurately extract and classify multi-scale and multi-category breeding areas in complex offshore environments is solved, achieving higher classification accuracy and model stability.

CN120014401APending Publication Date: 2025-05-16FUJIAN SATELLITE DATA DEV CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411839877.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional remote sensing technology is difficult to accurately extract and classify multi-scale and multi-category aquaculture areas in complex offshore water environments, especially when water bodies are turbid and aquaculture floating rafts and cages are of different sizes.

Method used

The multi-category breeding area remote sensing extraction method based on improved UNet is adopted, and the UNet network is improved by introducing DC Block and Stransformer modules, downsampling and upsampling structures are optimized, and multi-scale analysis and feature perception capabilities of the model are improved.

Benefits of technology

The classification accuracy of the model in the multi-category aquaculture area in complex offshore water environments is significantly improved, and it can more accurately extract and classify raft and cage aquaculture areas, which is suitable for the extraction tasks of aquaculture area under different scopes and backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014401A_ABST
    Figure CN120014401A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-category culture area remote sensing extraction method and equipment based on an improved UNet. The method comprises the following steps: acquiring and preprocessing a remote sensing image of an offshore culture area; the method comprises the following steps: constructing a raft and cage culture sample database in a culture state, and constructing a multi-category culture area automatic classification model: optimizing a down-sampling structure of a UNet model based on DC Block, and optimizing an up-sampling structure of the UNet model based on a Strusformer module; model training: inputting the sample data in the training set into an automatic classification model for model training, and performing testing and precision evaluation on the automatic classification model; and preprocessing the remote sensing time sequence of the culture area, inputting the preprocessed remote sensing time sequence into the trained automatic classification model, and accurately classifying the raft culture area and the net cage culture area. The method has the advantages that the classic UNet network is improved by utilizing the DC Block module and the Straransform module respectively, and the classification precision of the model on multi-scale and multi-category culture areas in an offshore complex water body environment is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing information processing and application, and in particular to a remote sensing extraction method and equipment for multi-category aquaculture areas based on improved UNet. Background Art

[0002] In order to build a strong marine nation and develop the marine economy, it is beneficial to scientifically manage and plan marine aquaculture by quickly and accurately distinguishing the categories of aquaculture areas and obtaining the spatial distribution of aquaculture areas. Traditional field measurement methods are time-consuming and labor-intensive and can hardly meet actual needs. The development of remote sensing technology has provided powerful assistance for monitoring large offshore aquaculture areas and has been widely used in offshore aquaculture area classification tasks. The commonly used methods for extracting and classifying offshore aquaculture areas based on remote sensing images mainly include visual interpretation, object-oriented segmentation, classical machine learning and deep learning. However, offshore aquaculture areas have problems such as turbid water bodies and different sizes of aquaculture rafts and cages. How to use AI technology and high-resolution image data to efficiently extract multi-scale and multi-category offshore aquaculture areas is particularly important for effective governance of illegal aquaculture and reasonable planning. Summary of the invention

[0003] In order to solve the above problems, the purpose of the present invention is to provide a multi-category aquaculture area remote sensing extraction method based on improved UNet. The classic UNet network is improved by using DC Block and Stransformer modules based on remote sensing image data, which effectively improves the classification accuracy of the model for multi-scale and multi-category aquaculture areas in complex offshore water environments.

[0004] To achieve the above object, the present invention adopts the following technical solutions:

[0005] Technical Solution 1

[0006] The invention discloses a multi-category aquaculture area remote sensing extraction method based on improved UNet, comprising the following steps: obtaining and preprocessing the remote sensing images of offshore aquaculture areas; constructing a raft and cage aquaculture sample database under aquaculture status by using the obtained remote sensing images, and dividing the sample data into a training set and a test set data set; constructing a multi-category aquaculture area automatic classification model: optimizing the downsampling structure of the UNet model based on the DCBlock, and optimizing the upsampling structure of the UNet model based on the Stransformer module; model training: inputting the sample data in the training set into the automatic classification model for model training, and after completion, testing the automatic classification model based on the test set, and evaluating the accuracy of the extraction result; and inputting the trained automatic classification model after preprocessing the remote sensing time series of the aquaculture area to accurately classify the raft aquaculture area and the cage aquaculture area.

[0007] Preferably, the UNet downsampling structure is optimized based on the DC Block module, and the steps are as follows: the DCBlock module includes two parallel light convolution groups, each convolution group includes 1 3×3 convolution, 2 3×3 convolutions and 3 3×3 convolution calculations, and then each convolution group splices the calculation results of different convolutions, and finally the calculation results of the two parallel convolution groups are fused and transmitted to the next stage as the output of the current DC Block module to realize the function of extracting features of different scales.

[0008] Preferably, the upsampling structure of UNet optimized based on the Stransformer module comprises the following steps: the Stransformer module is composed of two Transformer modules with different self-attention mechanisms, the feature map of the Stransformer module is input, firstly subjected to feature extraction of the first Transformer module, and then subjected to feature extraction of the second Transformer module; wherein the first Transformer module is composed of two batch normalization layers, a window-based multi-head self-attention module, a neighboring window connection module, and a multi-layer perceptron, and the second Transformer module is composed of two batch normalization layers, a window-based multi-head self-attention module, a neighboring window connection module, and a multi-layer perceptron. The input feature map passes through the first normalization layer of the first Transformer module, the window-based multi-head self-attention module, the neighboring window connection module, the second normalization layer and the multi-layer perceptron, and then the output of the first Transformer module is used as the input of the second Transformer module, and then passes through the first normalization layer of the second Transformer module, the window-based multi-head self-attention module based on spatial domain replacement, the neighboring window connection module, the second normalization layer and the multi-layer perceptron, and finally the output is transmitted to the next stage.

[0009] More preferably, the neighboring window connection module is implemented by a depthwise separable convolution operation, and the depthwise separable convolution includes two steps: depthwise convolution and pointwise convolution; the depthwise convolution operation is to perform a convolution operation on each channel of the feature map using multiple convolution kernels, and the number of convolution kernels is equal to the number of channels of the input feature map. The number of feature maps after the depthwise convolution is the same as the number of channels of the input layer, and pointwise convolution is performed after the depthwise convolution. The convolution kernel size of the pointwise convolution is 1×1×M, where M is the number of channels of the previous layer. The convolution operation is to perform a weighted combination of the feature maps obtained by the depthwise convolution operation in the depth direction to generate a new feature map. The number of feature maps output after the pointwise convolution operation is the same as the number of convolution kernels.

[0010] Preferably, in the model training step, the step of precision evaluation is: inputting the test sample into the automatic classification model to obtain the extraction result, and constructing the confusion matrix of the extraction result; based on the confusion matrix, respectively calculating the precision, recall, and F1 score of the raft aquaculture area extraction result and the precision, recall, and F1 score of the cage aquaculture area extraction result; wherein, precision represents the proportion of correctly classified raft or cage aquaculture areas extracted; recall represents the proportion of correctly classified aquaculture areas to the corresponding actual aquaculture areas on the image, and F1 score is the harmonic mean of precision and recall, which takes both precision and recall into consideration and can more objectively evaluate the performance of the model; the calculation formula is as follows: precision =

[0011] TP / (TP+FP), recall=TP / (TP+FN), F1 score=

[0012] 2×(precision×recall) / (precision+recall); where TP represents the correctly classified positive class, FP represents the incorrectly classified positive class, and FN represents the incorrectly classified negative class.

[0013] Based on the same inventive concept, the present invention also provides a multi-category aquaculture area remote sensing extraction device based on improved UNet.

[0014] Technical Solution 2

[0015] A multi-category aquaculture area remote sensing extraction device based on improved UNet includes a memory storing an executable program and a processor, wherein the processor runs the program to execute the method steps described in technical solution one.

[0016] The present invention has the following beneficial effects:

[0017] 1. Aiming at the problem that traditional extraction methods are difficult to distinguish between raft aquaculture and turbid water bodies, cage aquaculture and fishing boats and reefs, and the classification accuracy of aquaculture areas of different scales in complex offshore background environments is large, the present invention proposes a new semantic segmentation model suitable for accurate classification of remote sensing time series of large-scale raft and cage aquaculture areas in offshore areas. The DC Block and Stransformer modules are combined with UNet to optimize the downsampling structure, global feature extraction and upsampling structure, so as to improve the model's aquaculture area extraction performance in different ranges and backgrounds.

[0018] 2. The automatic classification model of the present invention is also applicable to the classification of other targets in seawater, especially to the situation where the target and the surrounding environment are intertwined and the target sizes are different.

[0019] 3. The automatic classification model of the present invention has high classification accuracy for targets in seawater background, and the model runs stably, with small amount of calculation and low system resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the process of the present invention;

[0021] Figure 2 A schematic diagram of the framework of the automatic classification model of the present invention;

[0022] Figure 3 It is a schematic diagram of the DC Block module of the present invention;

[0023] Figure 4 It is a schematic diagram of the STransformer module of the present invention;

[0024] Figure 5 It is a schematic diagram of the spatial domain replacement operation of the present invention;

[0025] Figure 6 It is a schematic diagram of the adjacent window connection module of the present invention;

[0026] Figure 7 It is a graph showing the extraction results of the present invention. DETAILED DESCRIPTION

[0027] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0028] Embodiment 1

[0029] See also Figure 1 ,The multi-category aquaculture area remote sensing extraction method based on improved UNet includes the following steps:

[0030] Step 10: Obtain remote sensing images of offshore aquaculture areas and perform preprocessing.

[0031] The remote sensing image source comes from the GF-1 satellite sensor. The specific process is as follows: query and screen the GF-1 high-resolution remote sensing images of the offshore area of ​​a city with a cloud cover of less than 20% within a preset time period, and download the selected data. Preprocess the downloaded GF-1 remote sensing images, including radiometric calibration, atmospheric correction, and image fusion, and then perform true color band fusion on the images after the above operations.

[0032] Step 20, use the acquired remote sensing images to construct a raft and cage aquaculture sample database under aquaculture status, and divide the sample data into training set and test set data sets. Specifically, based on ArcGIS software tools, a raft and cage aquaculture sample database under aquaculture status is constructed by means of visual interpretation. The raft aquaculture area is marked as 1 (white), the cage aquaculture area is marked as 2 (gray), and the background is marked as 0 (black). The raft and cage aquaculture area labels extracted manually by visual inspection are rasterized, and they are cut with the corresponding remote sensing images into sample data pairs with a size of 512×512 pixels, and they are divided into training data sets and test data sets. The divided training sample data is subjected to image enhancement processing, and the sample data is expanded by methods such as rotation, mirroring, and Gaussian blurring to improve the model accuracy and generalization ability.

[0033] After completing the construction of sample data, the automatic classification model is constructed, trained, tested, and the accuracy is evaluated. The specific steps are as follows:

[0034] Step 30: Build a multi-category breeding area automatic classification model: Please refer to Figure 2 Based on DC Block, the downsampling structure of the UNet model is optimized to realize the classification of aquaculture areas of different scales, greatly improving the multi-resolution analysis capability of the model. Based on the Stransformer module, the upsampling structure of the UNet model is optimized to realize the extraction of global feature information, establish long-distance dependency relationships between feature points, and improve the feature perception ability of the model.

[0035] Step 31, please refer to Figure 3, the UNet downsampling structure is optimized based on the DC Block module. Due to the large difference in the target scales between raft aquaculture areas and cage aquaculture areas, in order to better distinguish between the two, it is necessary to improve the model's multi-scale analysis capabilities. Therefore, the DC Block module is used to replace the general convolution layer to expand the receptive field and realize the multi-scale analysis of the target. DC Block is an improvement on the Inception module. It cancels the 3×3, 5×5 and 7×7 ordinary convolution modules in the Inception module that consume more computational resources. Instead, it chooses to use two 3×3 convolutions and three 3×3 convolutions to replace one 5×5 convolution and 7×7 convolution respectively. Through a series of lightweight 3×3 convolution combinations, the calculation effects of convolution kernels of different sizes are achieved and the network complexity is reduced, the operation efficiency is improved, and then the calculation results of different convolutions are spliced ​​to complete the splicing and fusion of features of different scales. Specifically, the DC Block module includes two parallel light convolution groups, each convolution group includes the calculation of 1 3×3 convolution, 2 3×3 convolutions and 3 3×3 convolutions, and then each convolution group splices the calculation results of different convolutions. Finally, the splicing results of the two parallel convolution groups are fused and transmitted to the next stage as the output of the current DC Block module to realize the function of extracting features of different scales.

[0036] Step 32: Optimize the upsampling structure of UNet based on the Stransformer module. Figure 4, the global information of the feature map is captured through the self-attention mechanism and spatial permutation operation in the Stransformer module, the long-distance dependency between feature points is established, and the feature perception ability of the model is improved. The Stransformer module consists of two Transformers with different self-attention mechanisms. The feature map input to the Stransformer module first undergoes feature extraction in the first Transformer module and then in the second Transformer module; the first Transformer module consists of two batch normalization layers, a window-based multi-head self-attention module, a neighboring window connection module, and a multi-layer perceptron, and the second Transformer module consists of two batch normalization layers, a window-based multi-head self-attention module based on spatial permutation, a neighboring window connection module, and a multi-layer perceptron. The input feature map passes through the first batch of normalization layers, the window-based multi-head self-attention module, the neighboring window connection module, the second batch of normalization layers and the multi-layer perceptron of the first Transformer module. Then the output of the first Transformer module is used as the input of the second Transformer module, and then passes through the first batch of normalization layers, the window multi-head self-attention module based on spatial domain replacement, the neighboring window connection module, the second batch of normalization layers and the multi-layer perceptron of the second Transformer module, and finally the output is transmitted to the next stage. The difference between the second Transformer module and the first Transformer module is that the window multi-head self-attention module is replaced by the window multi-head self-attention module based on spatial domain replacement, and the other structures remain unchanged. The added spatial domain replacement operation can solve the problem that the feature receptive field is limited to the window range when multiple window self-attention modules are superimposed, thereby improving the feature perception ability of the model. The spatial domain replacement operation is as follows: Figure 5 As shown in the figure, assuming that the input is a one-dimensional sequence, the same number represents that in the same divided window, before performing self-attention calculations on each window, the contents of different windows are mixed based on the spatial domain permutation operation. The original window 1 will contain the contents of windows 2 and 3. Similarly, windows 2 and 3 will also contain the contents of other windows. This operation avoids window isolation and enables interaction between all windows to improve the receptive field. Introducing the spatial domain permutation operation into the window self-attention module can establish cross-window connections, especially for windows that are far apart. However, when the input image size is much larger than the window size, a "grid problem" may occur. Therefore, a neighboring window connection module is added after each self-attention module. Its structure is as follows: Figure 6As shown in the figure. The neighboring window connection module is implemented by the depthwise separable convolution operation. The depthwise separable convolution includes two steps: depthwise convolution and pointwise convolution. The depthwise convolution operation uses multiple convolution kernels to perform a convolution operation on each channel of the feature map. The number of convolution kernels is equal to the number of channels of the input feature map. The number of feature maps after the depthwise convolution is the same as the number of channels of the input layer. The feature map cannot be expanded, and the feature information of different channels at the same spatial position cannot be effectively utilized. Therefore, pointwise convolution is required to combine these feature maps to generate a new feature map. The pointwise convolution operation is very similar to the conventional convolution operation. The convolution kernel size is 1×1×M, where M is the number of channels of the previous layer. The convolution operation will perform a weighted combination of the feature maps obtained by the depthwise convolution operation in the depth direction to generate a new feature map. The number of feature maps output after the pointwise convolution operation is the same as the number of convolution kernels. The neighboring window connection module can improve the information flow of neighboring windows and alleviate the grid problem.

[0037] Step 40, model training: input the sample data in the training set into the automatic classification model for model training. After completion, test the automatic classification model based on the test set, and evaluate the accuracy of the extraction results. The steps of accuracy evaluation are: input the test samples into the automatic classification model to obtain the extraction results, and construct the confusion matrix of the extraction results; based on the confusion matrix, respectively calculate the precision, recall, and F1 score of the raft aquaculture area extraction results and the precision, recall, and F1 score of the cage aquaculture area extraction results; wherein, precision represents the proportion of correctly classified raft or cage aquaculture areas extracted; recall represents the proportion of correctly classified aquaculture areas to the corresponding actual aquaculture areas on the image, and F1score is the harmonic mean of precision and recall. Taking both precision and recall into account, the performance of the model can be evaluated more objectively; the calculation formula is as follows: precision = TP / (TP+FP), recall =

[0038] TP / (TP+FN), F1 score = 2×(precision×recall) / (precision+recall); where TP represents the correctly classified positive class, FP represents the incorrectly classified positive class, and FN represents the incorrectly classified negative class.

[0039] The present invention is evaluated by the accuracy of the verification results of the test set, as shown in Table 1.

[0040] Table 1. Accuracy verification results of the automatic classification model for raft and cage aquaculture areas.

[0041]

[0042] Step 50: Pre-process the remote sensing time series of the aquaculture area and input it into the trained automatic classification model to accurately classify the raft aquaculture area and the cage aquaculture area. Figure 7 As shown in the figure, it includes the original RGB images a, classification results b and true labels c of locations 1 and 2. By comparing the classification results b and the true labels c, it is found that the distribution range of white (gray) patches in the classification results b is basically similar to the distribution of white (gray) patches in the true labels. Combined with the evaluation index, it can be seen that the improved UNet model has a high classification accuracy for aquaculture areas of different scales in complex offshore background environments.

[0043] Embodiment 2

[0044] A multi-category aquaculture area remote sensing extraction device based on improved UNet includes a memory storing an executable program and a processor, wherein the processor runs the program to execute the method steps described in the first embodiment.

[0045] Since the device introduced in the second embodiment of the present invention is a hardware device for implementing the method of the first embodiment of the present invention, those skilled in the art can understand the specific implementation of the device based on the method introduced in the first embodiment of the present invention, so it is not described here in detail. All methods adopted in the first embodiment of the present invention belong to the scope of protection of the present invention.

[0046] The above description is only a specific implementation mode of the present invention, and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A multi-category aquaculture area remote sensing extraction method based on improved UNet, characterized by: The steps include: Obtain remote sensing images of offshore aquaculture areas and perform preprocessing; The acquired remote sensing images are used to construct a raft and cage aquaculture sample database under aquaculture conditions, and the sample data is divided into training set and test set data sets; Construct a multi-category aquaculture area automatic classification model: optimize the downsampling structure of the UNet model based on the DC Block, and optimize the upsampling structure of the UNet model based on the Stransformer module; Model training: input the sample data in the training set into the automatic classification model for model training. After completion, the automatic classification model is tested based on the test set, and the accuracy of the extraction results is evaluated; After pre-processing, the remote sensing time series of the aquaculture areas are input into the trained automatic classification model to accurately classify the raft aquaculture areas and cage aquaculture areas.

2. The multi-category aquaculture area remote sensing extraction method based on improved UNet according to claim 1 is characterized in that: The UNet downsampling structure is optimized based on the DC Block module, and the steps are as follows: the DC Block module includes two parallel light convolution groups, each convolution group includes 1 3×3 convolution, 2 3×3 convolutions and 3 3×3 convolution calculations, and then each convolution group splices the calculation results of different convolutions, and finally the calculation results of the two parallel convolution groups are fused and transmitted to the next stage as the output of the current DC Block module to realize the function of extracting features of different scales.

3. The multi-category aquaculture area remote sensing extraction method based on improved UNet according to claim 1 is characterized in that: The upsampling structure of UNet optimized based on the Stransformer module comprises the following steps: the Stransformer module is composed of two Transformer modules with different self-attention mechanisms, and the feature map of the Stransformer module is firstly subjected to feature extraction of the first Transformer module, and then subjected to feature extraction of the second Transformer module; wherein the first Transformer module is composed of two batch normalization layers, a window-based multi-head self-attention module, a neighboring window connection module and a multi-layer perceptron, and the second Transformer module is composed of two batch normalization layers, a window-based multi-head self-attention module based on spatial domain permutation, a neighboring window connection module and a multi-layer perceptron; the input feature map is subjected to the first batch normalization layers, the window-based multi-head self-attention module, the neighboring window connection module, the second batch normalization layers and the multi-layer perceptron of the first Transformer module, and then the output of the first Transformer module is used as the input of the second Transformer module, and then is subjected to the first batch normalization layers, the window-based multi-head self-attention module based on spatial domain permutation, the neighboring window connection module, the second batch normalization layers and the multi-layer perceptron of the second Transformer module, and finally the output is transmitted to the next stage.

4. The multi-category aquaculture area remote sensing extraction method based on improved UNet according to claim 3 is characterized by: The neighboring window connection module is implemented by a depthwise separable convolution operation, which includes two steps: depthwise convolution and pointwise convolution. The depthwise convolution operation uses multiple convolution kernels to perform a convolution operation on each channel of the feature map. The number of convolution kernels is equal to the number of channels of the input feature map. The number of feature maps after the depthwise convolution is the same as the number of channels of the input layer. After the depthwise convolution, pointwise convolution is performed again. The convolution kernel size of the pointwise convolution is 1×1×M, where M is the number of channels of the previous layer. The convolution operation is to perform weighted combination of the feature maps obtained by the depthwise convolution operation in the depth direction to generate a new feature map. The number of feature maps output after the pointwise convolution operation is the same as the number of convolution kernels.

5. The multi-category aquaculture area remote sensing extraction method based on improved UNet according to claim 1 is characterized in that: In the model training step, the accuracy evaluation step is as follows: Input the test sample into the automatic classification model to obtain the extraction result, and construct the extraction result confusion matrix; Based on the confusion matrix, the precision, recall, and F1 score of the raft aquaculture area extraction results and the precision, recall, and F1 score of the cage aquaculture area extraction results are calculated respectively; among them, precision represents the proportion of correctly classified aquaculture areas in the extracted raft or cage aquaculture areas; recall represents the proportion of correctly classified aquaculture areas to the corresponding actual aquaculture areas on the image; and F1 score is the harmonic mean of precision and recall. Considering both precision and recall at the same time, the performance of the model can be evaluated more objectively. The calculation formula is as follows: precision = TP / (TP+FP), recall = TP / (TP+FN), F1 score = 2×(precision×recall) / (precision+recall); where TP represents the correctly classified positive class, FP represents the incorrectly classified positive class, and FN represents the incorrectly classified negative class.

6. Multi-category aquaculture area remote sensing extraction equipment based on improved UNet, characterized by: It comprises a memory storing an executable program and a processor, wherein the processor runs the program to execute the method steps described in any one of claims 1 to 5.

Citation Information

Cited By

  • Multi-rule constrained offshore cage culture type discrimination method and system

    CN121725366A

  • A method and system for identifying a type of multi-rule-constrained offshore net cage culture

    CN121725366B