Marine ship target segmentation method and device based on deep learning algorithm
Through a deep learning method combining a fully convolutional neural network and feedback attention mechanism, the problems of low accuracy and poor generalization ability of traditional methods in complex environments are solved, and the accurate segmentation and robustness of ship targets of various scales and shapes are achieved.
Patent Information
- Application Number
- CN202411765451.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-09
AI Technical Summary
The traditional marine ship target segmentation method has low accuracy and poor generalization capabilities in complex environments, making it difficult to effectively deal with ship targets of various scales and shapes.
A deep learning-based method is adopted, combined with a full convolutional neural network and feedback attention mechanism, a UNet image segmentation neural network model is built, and the encoding capability of the encoding end is improved through the feedback attention module, and layer standardization and Gelu activation function are used in the feature extraction stage.
It improves the accuracy and robustness of offshore ship target segmentation, can effectively identify and segment ship targets of various scales and shapes in complex environments, and improves the training and inference speed of the model.
Smart Images

Figure CN119963571A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a method and device for segmenting marine ship targets based on a deep learning algorithm. Background Art
[0002] Maritime ship target segmentation is one of the important tasks in the field of ocean monitoring and navigation safety. A variety of technologies are used to accurately extract and identify maritime ship images. Traditional methods mainly include manual segmentation methods based on rules and feature engineering. However, these methods have the problems of low accuracy and poor model generalization ability. With the development of deep learning technology, neural network-based target segmentation models have gradually become mainstream, such as convolutional neural networks (CNNs) and semantic segmentation networks. These models can automatically learn image features and perform end-to-end training on large-scale data sets, thereby achieving better segmentation effects and generalization capabilities. In addition, combining sensor technology to obtain multi-source data and using advanced image processing and computer vision technologies also provide more comprehensive support for maritime ship target segmentation. However, the complexity of the marine environment and the instability of image quality are still challenges. Therefore, it is necessary to continue to explore stable, efficient, and generalized segmentation models and algorithms. Summary of the invention
[0003] The present invention provides a method and device for marine ship target segmentation based on deep learning. By combining a fully convolutional neural network and a feedback attention mechanism, the accuracy and robustness of marine ship target segmentation under complex environmental conditions are improved. The method can effectively solve the problems of complex marine environment and different ship sizes in traditional methods, and realize accurate segmentation of ship targets of various scales and shapes.
[0004] The deep learning-based marine ship target segmentation method provided in the present disclosure mainly includes the following steps:
[0005] S1, data acquisition: obtaining marine ship image data;
[0006] S2, data preprocessing, including: one or more of image denoising, size standardization, color space conversion, data enhancement, and data annotation;
[0007] S3, constructing and training an optimized deep learning model for segmenting marine ship targets; the deep learning model is constructed based on the UNet image segmentation neural network model, specifically including:
[0008] The input of the model first passes through two convolutional layers, and then undergoes four downsampling and four upsampling processes;
[0009] After the fourth upsampling, the decoder feature map does not pass through the output layer, but is calculated by the feedback attention module to obtain the new encoder feature.
[0010] The new encoded features replace the previous encoded features, and after four downsampling and four upsampling again, the final output result is obtained by passing through the output layer.
[0011] Furthermore, the convolutional layers are composed of a 3*3 convolution kernel, batch normalization, and a ReLU activation function.
[0012] Furthermore, the downsampling process is as follows:
[0013] Perform strided convolution on the input;
[0014] After flattening, the first layer is normalized;
[0015] First fully connected layer;
[0016] After reshaping, the second convolutional layer;
[0017] After flattening and using Gelu activation, use the second fully connected layer;
[0018] The output of the second fully connected layer is summed with the normalized output of the first layer;
[0019] The second layer normalization,
[0020] Output after reshaping.
[0021] Furthermore, the feedback attention module adopts a channel-based feedback attention module, and its calculation process is as follows:
[0022] For the features from the decoder, global average pooling is applied to convert them into a tensor of (b,c,1,1);
[0023] Among them, b is batch, which represents the batch size fed into the network model at one time, that is, how many images are fed into the network model at one time; c is channel, which represents the number of channels of the image fed into the network model;
[0024] After that, after flattening and fully connected layers, a new feature vector with dimension (b, c) is obtained, which is reshaped to (b, c, 1, 1) to obtain the feedback attention from the decoding end;
[0025] The encoding-side feature EFM is calculated with the obtained feedback attention according to the following formula to obtain the new encoding-side feature NEFM:
[0026] gama*EFM*FeedbackAttention+EFM=NEFM
[0027] Among them, gamma is initialized to 0 and is a learnable parameter;
[0028] The newly obtained encoding-end features are used to replace the original encoding-end features and re-enter the downsampling stage.
[0029] Furthermore, the method further comprises the following steps:
[0030] The segmentation result obtained in step S3 is optimized and screened by using one or more of the following processes: noise removal, hole filling, boundary smoothing, and morphological operation.
[0031] The deep learning-based marine ship target segmentation device using the above method mainly includes:
[0032] A data acquisition unit, used to acquire marine ship image data;
[0033] A data preprocessing unit, used for performing one or more of image denoising, size standardization, color space conversion, data enhancement, and data annotation;
[0034] Image segmentation unit based on deep learning model, used to segment marine ship targets;
[0035] Model training and optimization unit, used to train and optimize deep learning models;
[0036] The deep learning model is based on the structure of the image segmentation neural network model UNet, and includes: a convolutional layer, a downsampling part as an encoder, an upsampling part as a decoder, and a feedback attention module;
[0037] The input of the model first passes through two convolutional layers, and then undergoes four downsampling and four upsampling processes. After the fourth upsampling, the decoding-end feature map obtained does not pass through the output layer, but is calculated by the feedback attention module to obtain a new encoding-end feature; the new encoding feature replaces the previous encoding feature, and after four downsampling and four upsampling again, it passes through the output layer to obtain the final output result.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) based on the encoding end of SegFormer, a downsampling module based on fully connected and convolutional neural networks is designed, which removes the complex attention calculation mechanism and random inactivation layer, thereby improving the training and inference speed of the model;
[0039] (2) A channel-based feedback attention module (CFA) is used. The feedback mechanism of this module enables the encoder to obtain information from the decoder, thereby improving the encoding capability of the encoding segment. At the same time, the channel-based attention calculation greatly reduces the complexity of attention calculation.
[0040] (3) In the feature extraction stage, batch normalization and Relu activation function are not used. Instead, layer normalization and Gelu activation function are used. Layer normalization balances the features of each sample at different time steps or spatial positions, which helps the model maintain a stable distribution during training and can reduce training instability caused by changes in batch size. Secondly, the use of Gelu activation function can effectively retain negative information, thereby better capturing the complex features of the input image in the feature extraction stage and improving the expressiveness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.
[0042] Figure 1 is an overall flow chart of an exemplary embodiment according to the present disclosure;
[0043] Figure 2 A network structure model proposed for an exemplary embodiment;
[0044] Figure 3 This is an example of a downsampling module based on a fully connected and convolutional neural network;
[0045] Figure 4 An example of a channel-based feedback attention module. DETAILED DESCRIPTION
[0046] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0047] The present disclosure provides a method for segmenting marine ship targets based on a deep learning algorithm. The flowchart of the exemplary embodiment is shown in the attached figure. Figure 1 As shown, the process mainly includes the following steps.
[0048] 1. Data Collection
[0049] Collect image or video data of ships at sea from equipment such as satellite imagery, navigation monitoring systems and ship cameras.
[0050] 2. Data preprocessing
[0051] The collected data is processed, including image denoising, size standardization, color space conversion, data enhancement, data labeling, etc.
[0052] 3. Model construction and training optimization
[0053] Mainly include:
[0054] (1) Design a deep learning network model that combines feedback attention mechanism and UNet for object segmentation.
[0055] In this embodiment, based on the characteristics of the task of marine ship target segmentation, a deep learning network model based on convolutional neural network, fully connected neural network and feedback attention mechanism is constructed based on the Unet image segmentation neural network model. The characteristics of its network structure are mainly as follows:
[0056] 1) Downsampling module based on fully connected and convolutional neural network: Figure 1 The downsampling module based on fully connected and convolutional neural networks can not only improve the model's ability to extract features, but also improve the training and inference speed of the model by eliminating complex attention calculations, which is obviously helpful for real-time or near real-time maritime ship target segmentation tasks.
[0057] 2) Introduce an attention feedback mechanism, preferably using a channel-based feedback attention module (Channel Feedback Attention, abbreviated as CFA).
[0058] In ship target segmentation, the shape, size and position of the ship may vary greatly in different images. The feedback mechanism of the CFA module will enable the encoder to obtain information from the decoder, thereby improving the encoding capability of the encoding segment and helping to improve the adaptability to these changes.
[0059] At the same time, channel-based attention calculation greatly reduces the complexity of attention calculation.
[0060] 3) The upsampling module at the decoding end adopts the common upsampling method with skip links.
[0061] As a preferred embodiment, the specific structure of the deep learning model is as shown in the attached Figure 2 As shown:
[0062] The input of the model will first pass through two convolutional layers (both composed of 3*3 convolution kernels, batch normalization, and Relu activation functions), and then undergo four downsampling and four upsampling processes. After the fourth upsampling, the decoder feature map obtained will not pass through the output layer, but will be calculated by the feedback attention module to obtain the new encoding feature. The new encoding feature will replace the old encoding feature and undergo four downsampling and four upsampling again before passing through the output layer to obtain the final output result.
[0063] The main improvements include the following two aspects:
[0064] Ⅰ. Designed a downsampling module based on fully connected and convolutional neural networks: Based on the encoding end of SegFormer, the complex attention calculation mechanism and random inactivation layer were removed, and a new downsampling module was rebuilt in combination with a fully connected neural network.
[0065] The encoding end of the traditional SegFormer uses convolution operations to cleverly simplify the position encoding stage. However, in this embodiment, it is noted that during the network training process, the computational complexity of its attention layer is large, which is not conducive to model training and inference. At the same time, because it deals with natural image tasks and the training data set is large, SegFormer uses a large number of random deactivation layers to further improve the generalization ability of the model. These practices are not advantageous for processing marine ship image tasks, but cannot make good use of the random deactivation layer. Therefore, in this embodiment, the attention layer and random deactivation layer are removed, and the downsampling module is redesigned by combining the fully connected neural network and the convolutional neural network. The overall structure is shown in the attached figure. Figure 3 And as shown in Table 1.
[0066] Table 1 Network structure of downsampling module
[0067] layer parameter Output Dimensions enter (b,c,2h,2w) Convolution with stride 3×3 convolution kernel with a stride of 2 (b,c,h,w) Flattened layer standardization (b,h×w,c) Fully connected layer The input feature is c, and the output feature is 4c (b,h×w,4*c) Use convolutional layers after reshaping 3×3 convolution kernel (b,4c,h,w) Flattening uses Gelu activation and then uses a fully connected layer The input feature is 4c and the output feature is c (b,h×w,c) Sum with the normalized output of the first layer (b,h×w,c) Layer normalization followed by reshaping (b,c,h,w)
[0068] Compared with the SegFormer encoding end, this module does not use complex attention calculations and random inactivation layers; compared with the FCN and UNet encoding ends, this module integrates fully connected layers and convolutional layers, applies fully connected layers in feature extraction, and uses the advantages of full connection to improve the model's feature extraction capabilities. It should be noted that this module does not use batch normalization and Relu activation functions in the feature extraction stage, but uses layer normalization and Gelu activation functions.
[0069] II. Design of a channel-based feedback attention module
[0070] In previous deep learning methods, the role of the decoding end is often only to restore the size of the feature map to obtain a segmentation mask of a specified size. From the perspective of the segmentation model as a whole, the decoding-encoding structure itself is similar to a structure for calculating an attention feature map.
[0071] Based on this idea, the present disclosure extracts an attention feature map from the feature map calculated at the decoding end, and fuses the features corresponding to the encoding end with the attention feature map in a feedback mechanism, so as to improve the ability of the encoding end to extract semantic features; at the same time, in order to improve the model training and inference speed, inspired by feedback attention and channel attention, this embodiment provides a feedback attention module based on channels with fast calculation speed, as shown in the attached figure. Figure 4 As shown:
[0072] In this module, the features from the decoder are firstly subjected to global average pooling and converted into a tensor of (b, c, 1, 1). Then, a new feature vector with the dimension of (b, c) is obtained through flattening and a fully connected layer. After being reshaped into (b, c, 1, 1), the feedback attention from the decoder is obtained. Formula (1) is used to calculate the encoder feature (EFM) and the obtained feedback attention to obtain a new encoder feature (NEFM). The newly obtained encoder feature replaces the old encoder feature and re-enters the downsampling stage.
[0073] gama*EFM*FeedbackAttention+EFM=NEFM#(1)
[0074] Among them, gamma is initialized to 0 and is a learnable parameter.
[0075] (2) For the constructed deep learning model, the cross entropy loss function and Adam optimizer are used to train the model, and the model performance is optimized using techniques such as cross validation, batch processing, and regularization.
[0076] (3) Adjust model parameters such as learning rate, loss function, and optimizer to obtain the best training results.
[0077] (4) Use the validation set or test set to evaluate the performance of the trained model on unseen data, and adjust and improve the model based on the evaluation results.
[0078] Preferably, this embodiment further comprises the following steps:
[0079] 4. Post-processing:
[0080] (1) Use edge smoothing methods, such as pixel neighborhood-based filters or edge detection-based algorithms, to smooth the edges of the ship target, making the segmentation results clearer and more continuous.
[0081] (2) According to certain criteria and thresholds, adjacent small regions are removed or merged into a larger region by merging regions, thereby improving the overall quality of the segmentation results.
[0082] (3) Use morphological operations including corrosion, dilation, opening and closing operations to optimize and repair the segmentation results, making the target area more complete and connected.
[0083] 5. Intelligent Applications:
[0084] (1) Determine the specific needs and goals of intelligent applications;
[0085] (2) Deploy the trained model to actual application environments such as cloud servers, mobile devices, and embedded systems.
[0086] In summary, this embodiment provides a method for segmenting marine ship targets based on deep learning. By analyzing the image data of marine ships, the method can accurately identify and segment target ships, thereby improving the automation and intelligence level of the marine monitoring system. Its main features are:
[0087] 1. Downsampling module based on fully connected and convolutional neural network
[0088] Ⅰ. Reduce computational complexity: Removing complex attention calculation mechanisms and random inactivation layers, and using fully connected and convolutional neural network downsampling modules can significantly reduce the amount of computation and improve the training and inference speed of the model. This is especially important for real-time or near real-time marine ship target segmentation tasks;
[0089] II. Improve speed: The marine environment usually has complex backgrounds and dynamically changing targets, so efficient downsampling can speed up the processing of marine images and ensure rapid response, especially in monitoring and reconnaissance missions;
[0090] III. Maintaining feature expression capabilities: Full connectivity and convolutional downsampling can effectively extract local features and reduce redundant information, thereby ensuring that the model improves speed while maintaining accuracy.
[0091] 2. Channel-based Feedback Attention Module (CFA)
[0092] Ⅰ. Enhanced encoding capability: By feeding back information from the decoder to the encoder, the CFA module enables the encoder to capture key features more effectively. Because in ship target segmentation, the shape, size, and position of the ship may vary greatly in different images, the feedback mechanism helps improve the ability to adapt to these changes;
[0093] II. Reduce complexity: The traditional global attention mechanism has high computational complexity and memory requirements, while the CFA module uses channel-based attention calculation, which greatly reduces the consumption of computing resources. This is very critical for resource-constrained edge devices or real-time application scenarios.
[0094] III. Improve the model’s contextual understanding ability: Through the feedback mechanism, the model can better understand the contextual information of the target, thereby improving the accuracy of the segmentation of maritime ships.
[0095] 3. Use layer normalization and GELU activation function
[0096] Ⅰ. Adapt to small batch data: Layer normalization normalizes each sample and does not rely on batch statistics, which makes the model more stable when processing small batch data, especially in the offshore monitoring scenario, where the data samples may be small and unevenly distributed;
[0097] II. Improve nonlinear characteristics: Compared with ReLU, the GELU activation function can better retain negative information and enhance the nonlinear expression ability of the model; in complex marine environments, this feature helps to capture more detailed features and thus improve the segmentation accuracy of ship targets.
[0098] The above technical scheme is only an exemplary embodiment of the present invention. For those skilled in the art, it is easy to make various types of improvements or modifications based on the application methods and principles disclosed in the present invention, and it is not limited to the method described in the above specific embodiment of the present invention. Therefore, the method described above is only preferred and does not have a restrictive meaning.
Claims
1. A method for segmenting marine ship targets based on a deep learning algorithm, comprising the following steps: S1, data acquisition: obtaining marine ship image data; S2, data preprocessing, including: one or more of image denoising, size standardization, color space conversion, data enhancement, and data annotation; S3, constructing and training an optimized deep learning model for segmenting marine ship targets; the deep learning model is constructed based on the UNet image segmentation neural network model, specifically including: The input of the model first passes through two convolutional layers, and then undergoes four downsampling and four upsampling processes; After the fourth upsampling, the decoder feature map does not pass through the output layer, but is calculated by the feedback attention module to obtain the new encoder feature. The new encoded features replace the previous encoded features, and after four downsampling and four upsampling again, the final output result is obtained by passing through the output layer.
2. The method according to claim 1, characterized in that The convolutional layers are composed of a 3*3 convolution kernel, batch normalization, and a ReLU activation function.
3. The method according to claim 1, characterized in that The downsampling process is as follows: Perform strided convolution on the input; After flattening, the first layer is normalized; First fully connected layer; After reshaping, the second convolutional layer; After flattening and using Gelu activation, use the second fully connected layer; The output of the second fully connected layer is summed with the normalized output of the first layer; The second layer normalization, Output after reshaping.
4. The method according to claim 1, characterized in that: The feedback attention module adopts a channel-based feedback attention module, and its calculation process is as follows: For the features from the decoder, global average pooling is applied to convert them into a tensor of (b,c,1,1); Among them, b is batch, which represents the batch size fed into the network model at one time, that is, how many images are fed into the network model at one time; c is channel, which represents the number of channels of the image fed into the network model; After that, after flattening and fully connected layers, a new feature vector with dimension (b, c) is obtained, which is reshaped to (b, c, 1, 1) to obtain the feedback attention from the decoding end; The encoding-side feature EFM is calculated with the obtained feedback attention according to the following formula to obtain the new encoding-side feature NEFM: gama*EFM*FeedbackAttention+EFM=NEFM Among them, gamma is initialized to 0 and is a learnable parameter; The newly obtained encoding-end features are used to replace the original encoding-end features and re-enter the downsampling stage.
5. The method according to any one of claims 1 to 4, further comprising the following steps: The segmentation result obtained in step S3 is optimized and screened by using one or more of the following processes: noise removal, hole filling, boundary smoothing, and morphological operation.
6. A marine ship target segmentation device according to any one of the methods described in claims 1-5, characterized in that: include: A data acquisition unit, used to acquire marine ship image data; A data preprocessing unit, used for performing one or more of image denoising, size standardization, color space conversion, data enhancement, and data annotation; Image segmentation unit based on deep learning model, used to segment marine ship targets; Model training and optimization unit, used to train and optimize deep learning models; The deep learning model is based on the structure of the image segmentation neural network model UNet, and includes: Convolutional layer, the downsampling part as the encoder, the upsampling part as the decoder, and the feedback attention module; the input of the model first passes through two convolutional layers, and then undergoes four downsampling and four upsampling processes. After the fourth upsampling, the decoder feature map does not pass through the output layer, but is calculated by the feedback attention module to obtain a new encoding feature; the new encoding feature replaces the previous encoding feature, and after four downsampling and four upsampling again, it passes through the output layer to obtain the final output result.