Ship instance segmentation method based on multi-scale convolutional neural network

The MALCP-Net model based on a multi-scale convolutional neural network solves the problem of low segmentation accuracy of small and medium-sized ships in complex backgrounds in existing technologies, achieves high-precision ship instance segmentation, and improves detection and segmentation effects.

CN120673066APending Publication Date: 2025-09-19NANCHANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510834992.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing ship instance segmentation methods have low detection accuracy and poor segmentation effect when segmenting small ship instances densely distributed in complex offshore backgrounds.

Method used

The MALCP-Net model based on a multi-scale convolutional neural network is adopted. Through the improved Cascade-Mask-RCNN model, it combines the multi-scale atrous layered convolution module MALCM, the feature pyramid network FPN, the SE attention module, the Tri-Attention module and the proposal generation and imitation learning network CFINet for image enhancement and feature extraction. It uses a multi-task loss function for training to achieve high-precision instance segmentation.

Benefits of technology

It significantly improves the detection accuracy and segmentation effect of small ship targets in complex backgrounds, can accurately locate and identify ships in multi-scale environments, generate accurate masks, and improves detection accuracy and segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673066A_ABST
    Figure CN120673066A_ABST
Patent Text Reader

Abstract

The invention discloses a ship instance segmentation method based on a multi-scale convolutional neural network, and relates to the technical field of instance segmentation. The invention discloses a ship instance segmentation method based on a multi-scale convolutional neural network. The ship instance segmentation method comprises the following steps of image acquisition and processing, image enhancement, model construction, model training and instance segmentation. According to the method, the multi-proportion hole hierarchical convolution module MALCM is added to the output part of the backbone network, different scale features can be effectively captured and feature hierarchy can be enhanced by means of multi-scale hole convolution and a transverse addition connection structure in the MALCM module, and the multi-scale ship target detection accuracy of the MALCP-Net network model can be remarkably improved. Especially, the detection and segmentation capability of a small ship target is improved, the MALCP-Net network model is assisted to accurately position and identify the ship under a complex background, and the detection precision and segmentation effect of the ship instance segmentation method are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of instance segmentation, and in particular to a ship instance segmentation method based on a multi-scale convolutional neural network. Background Art

[0002] With the rapid development of modern science and technology, people's demand for effective monitoring and management of maritime ships is increasing. Synthetic Aperture Radar (SAR), as an active microwave imaging sensor, has the characteristics of all-weather and all-day imaging because its microwave signals can penetrate clouds and are not affected by weather and sunlight. It has become one of the important means of earth observation. In recent years, the use of SAR images for ship target detection and instance segmentation has received great attention in both military and civilian fields. In the military field, it is beneficial to tactical deployment and improve coastal defense warning capabilities, while in the civilian field, it is helpful for maritime monitoring and management. Instance segmentation is to perform pixel-level segmentation on the target based on the detection box to extract the precise boundary or mask of the target. Compared with object detection based on bounding boxes, instance segmentation not only needs to determine the location and type of the target, but also needs to generate a mask vector for each target to describe its precise boundary. Therefore, instance segmentation can obtain richer detailed information, thereby effectively separating overlapping targets and distinguishing between ships, and distinguishing ships from their surroundings.

[0003] Current advanced instance segmentation methods mainly include Cascade-Mask-RCNN, SOLOv2, YOLACT, etc. If models such as Cascade-Mask-RCNN are directly applied to the instance segmentation task of small ships densely distributed under complex backgrounds, the complex background will interfere with the model's extraction of ship features, and the extractable features of small ships themselves are limited. The dense distribution will cause the features to be confused with each other, making it difficult for the model to accurately distinguish the boundaries and features of each ship. As a result, the existing ship instance segmentation methods have low detection accuracy and poor segmentation effect when performing instance segmentation on small ships densely distributed under complex offshore backgrounds.

[0004] Based on the above situation, the present invention proposes a ship instance segmentation method based on a multi-scale convolutional neural network with high detection accuracy and good segmentation effect. Summary of the Invention

[0005] In order to overcome the shortcomings of existing ship instance segmentation methods, such as low detection accuracy and poor segmentation effect, when performing instance segmentation of small ships densely distributed in complex offshore backgrounds, the present invention proposes a ship instance segmentation method based on a multi-scale convolutional neural network with high detection accuracy and good segmentation effect.

[0006] A ship instance segmentation method based on a multi-scale convolutional neural network comprises the following steps: Image acquisition and processing: using synthetic aperture radar (SAR) to acquire an image dataset containing ship targets, and then labeling, cropping, and scaling the SAR images in the image dataset to obtain a pre-processed image dataset. Image enhancement: performing image enhancement and normalization processing on the SAR images in the preprocessed image dataset to obtain an enhanced image dataset; Model construction: construct the MALCP-Net network model based on the improved Cascade-Mask-RCNN model; Model training: The constructed MALCP-Net network model is trained using the acquired enhanced image dataset to obtain a ship instance segmentation model; Instance segmentation: Preprocess and enhance the SAR images acquired in real time. The processed SAR images are input into the ship instance segmentation model and instance segmentation is performed to obtain the ship instance segmentation results.

[0007] As a preferred aspect of the invention, the specific steps of performing image enhancement on the SAR image in the preprocessed image dataset are: The Lee filter method is selected to enhance the SAR image and suppress the coherent speckle noise in the image. The calculation formula of the Lee filter method is as follows:

[0008] in Indicates that the Lee filtered image is at position The pixel value of Indicates that the original image is at position The pixel value of represents the local mean in the neighborhood, represents the local variance within the neighborhood, Indicates the grayscale level of the image, represents the estimated speckle noise variance.

[0009] As a preferred aspect of the invention, the specific steps of constructing the MALCP-Net network model based on the improved Cascade-Mask-RCNN model are: Construct a backbone network, selecting ResNet50 as the backbone network to extract deep features of SAR images. Add a multi-scale atrous layered convolution module (MALCM) to the output of the backbone network to extract multi-scale features from the feature maps output by ResNet50 and enhance the hierarchical nature of features by fusing features at different scales. To build a neck network, we used a feature pyramid network (FPN) and embedded the SE attention module on top of the FPN to form the SE-FPN network. This network enhances the multi-scale feature fusion effect and highlights important feature information through the attention mechanism. We then added the triple attention mechanism module (Tri-Attention) after the SE-FPN network to capture cross-dimensional interaction information, improving the expressiveness of features and the efficiency of spatial information utilization. A detection head is constructed, using the proposal generation and imitation learning network model CFINet to detect ship targets and generate candidate boxes. An instance segmentation mask branch is then added to CFINet to further extract and upsample features within each candidate box through a series of convolutional and deconvolutional layers, generating pixel-level masks corresponding to the candidate boxes and achieving instance segmentation of ships. The backbone network, neck network and detection head constructed above are connected and integrated to obtain the MALCP-Net network model.

[0010] As a preferred aspect of the invention, the structure of the multi-scale hole layered convolution module MALCM is specifically composed of: The channel conversion convolution layer is used to convert the number of channels of the feature map output by ResNet50 into a number of channels suitable for subsequent processing; The branch processing layer is used to copy the feature map after channel conversion into four copies and input them into four branches respectively, where branch 1 is the standard convolution layer, i.e. Convolution layer, branch 2 is a dilated convolution layer with a dilation rate of 3, branch 3 is a dilated convolution layer with a dilation rate of 6, and branch 4 is a dilated convolution layer with a dilation rate of 9. Before the dilated convolution operation, branches 2 and 4 need to add the results of the dilated convolution operation of branches 1 and 3 respectively. The output calculation formula of each branch is as follows:

[0011] in Indicates the The input feature map of the branch, and , Express The expansion rate is The dilated convolution calculation, Indicates the Output feature map of each branch; The feature fusion layer is used to splice the output feature maps of the four branches.

[0012] As a preferred aspect of the invention, the SE-FPN network is mainly composed of a feature pyramid network FPN and an SE attention module embedded therein, wherein the feature pyramid network FPN part includes a bottom-up path and a top-down path as well as lateral connections, which are used to generate feature maps of different scales, and the SE attention module is embedded in each stage of the FPN, which is used to perform attention weighting on the channel dimension of the feature map, enhance the response of important feature channels and suppress unimportant channels.

[0013] As a preferred aspect of the invention, the triple attention mechanism module Tri-Attention consists of three parallel attention branches, corresponding to the attention calculations in the three dimensions of channel-height, channel-width and space, and each branch includes a rotation operation, a Z-pool layer, a convolution and normalization layer, and an activation and application operation, which is used to simultaneously capture the complex interaction relationship between the channel and spatial dimensions, and comprehensively model the dependencies between features.

[0014] As a preferred aspect of the invention, the proposal generation and imitation learning network model CFINet mainly includes a coarse-to-fine region proposal network CRPN and a feature imitation FI branch, wherein the coarse-to-fine region proposal network CRPN is used to generate high-quality candidate boxes for small ship targets through a dynamic anchor selection strategy and cascade regression; the feature imitation FI branch is mainly composed of an example feature set and a feature-to-embedding module, the example feature set is used to retain high-quality ROI features, and the feature-to-embedding module is used to map regional features to an embedding space.

[0015] As a preferred aspect of the invention, the specific steps of training the constructed MALCP-Net network model using the acquired enhanced image dataset are: Data preparation and preprocessing: Obtain an enhanced image dataset, and divide the obtained enhanced image dataset into a training set, a validation set, and a test set according to the ratio of 70%, 15%, and 15%, respectively; Forward propagation and loss calculation, input the sample data in the training set into the MALCP-Net network model to obtain the instance segmentation result of the ship, using the multi-task loss function Calculate the loss between the instance segmentation result and the true label; Back propagation and parameter update, using the chain rule to reversely calculate the gradient of each layer parameter from the loss, and using the stochastic gradient descent optimization algorithm SGD to adjust the parameters; Iterative optimization: divide the training set into multiple small batches, repeat the previous two steps batch by batch, and traverse the entire training set multiple times until the MALCP-Net network model converges or reaches the preset number of iterations; Verification and optimization: Periodically evaluate the MALCP-Net network model using sample data from the validation set, calculate various error metrics of the MALCP-Net network model on the validation set, and use this to evaluate the fitting effect and generalization ability of the MALCP-Net network model. Based on the evaluation results of the validation set, adjust the hyperparameters of the MALCP-Net network model until the segmentation performance of the MALCP-Net network model meets the preset standards. Model testing,After training is completed, the final performance evaluation of the MALCP-Net network model is performed,and various error indicators are calculated using the sample data in the test set.

[0016] As a preferred aspect of the invention, the multi-task loss function The calculation formula is:

[0017] in represents the bounding box regression loss, calculated using the Smooth-L1 loss function, represents the classification loss, which is calculated using the cross entropy loss function, and It represents the total detection loss, Represents the total segmentation loss, calculated using the Dice loss function. and Represent the total detection loss and the total loss of segmentation The weight coefficient of .

[0018] The present invention has the following advantages: 1. By adding a multi-scale atrous layered convolution module (MALCM) to the output part of the backbone network, the present invention can effectively capture features of different scales and enhance feature hierarchy with the help of the multi-scale atrous convolution and lateral additive connection structure in the MALCM module, significantly improving the detection and segmentation capabilities of the MALCP-Net network model for multi-scale ship targets, especially small ship targets, and helping the MALCP-Net network model to accurately locate and identify ships in complex backgrounds, thereby improving the detection accuracy and segmentation effect of this ship instance segmentation method.

[0019] 2. The present invention embeds the SE attention module on the basis of the feature pyramid network FPN and adds the triple attention mechanism module Tri-Attention after the SE-FPN network. It can not only use the channel attention mechanism to weight the channel dimension of the feature map, thereby enhancing the response of important feature channels and suppressing unimportant channels to improve feature expression ability and multi-scale feature fusion effect, but also use the Tri-Attention module to simultaneously capture the complex interaction relationship between channels and spatial dimensions, and comprehensively model the dependencies between features, thereby enhancing the expression ability of features and the utilization efficiency of spatial information, so as to improve the accuracy and robustness of the MALCP-Net network model. The combination of the two not only enhances the expression ability of features, but also makes the MALCP-Net network model utilize features more fully and accurately, further improving the performance of the MALCP-Net network model in tasks such as target detection and segmentation, and improving the detection accuracy and segmentation effect of this ship instance segmentation method.

[0020] 3. By adopting the proposal generation and imitation learning network model CFINet as the detection head, the present invention can not only effectively improve the detection accuracy and recall rate of the MALCP-Net network model for small ship targets in SAR images with the help of coarse-to-fine region proposal and feature imitation learning mechanism, but also generate high-quality candidate boxes with the help of dynamic anchor selection strategy and cascade regression to reduce the amount of subsequent processing calculations. On the basis of CFINet, an instance segmentation mask branch is added to further perform pixel-level segmentation on the detected ship targets and generate accurate masks to achieve accurate segmentation of the targets, thereby improving the detection accuracy and segmentation effect of this ship instance segmentation method.

[0021] 4. The present invention adopts a multi-task loss function in the training process of the MALCP-Net network model By calculating the loss between the instance segmentation result and the true label, both detection loss and segmentation loss can be considered simultaneously, enabling the MALCP-Net network model to simultaneously optimize both target detection and instance segmentation tasks. By introducing weight coefficients to balance the contributions of different tasks, the MALCP-Net network model can more comprehensively learn the characteristics of ship targets, thereby improving the MALCP-Net network model's ability to detect and segment multi-scale ship targets in complex backgrounds, and enhancing the detection accuracy and segmentation effect of this ship instance segmentation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A schematic flow chart of a ship instance segmentation method based on a multi-scale convolutional neural network adopted in an embodiment of the present invention.

[0023] Figure 2Schematic diagram of the architecture of the MALCP-Net network model adopted in an embodiment of the present invention.

[0024] Figure 3 This is a schematic diagram of the structure of the MALCM module used in an embodiment of the present invention.

[0025] Figure 4 This is a comparison chart of instance segmentation results on HRSID using the MALCP-Net network and the Cascade-Mask-RCNN model used in an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0027] Example 1, a ship instance segmentation method based on a multi-scale convolutional neural network, the overall implementation process is as follows Figure 1 As shown, the following steps are included: Image acquisition and processing: using synthetic aperture radar (SAR) to acquire an image dataset containing ship targets, and then labeling, cropping, and scaling the SAR images in the image dataset to obtain a pre-processed image dataset. Image enhancement: performing image enhancement and normalization processing on the SAR images in the preprocessed image dataset to obtain an enhanced image dataset; Model construction, based on the improved Cascade-Mask-RCNN model, the MALCP-Net network model is constructed. The specific architecture of the MALCP-Net network model is as follows: Figure 2 As shown; Model training: The constructed MALCP-Net network model is trained using the acquired enhanced image dataset to obtain a ship instance segmentation model; Instance segmentation: Preprocess and enhance the SAR images acquired in real time. The processed SAR images are input into the ship instance segmentation model and instance segmentation is performed to obtain the ship instance segmentation results.

[0028] The specific steps of performing image enhancement on the SAR image in the preprocessed image dataset are: The Lee filter method is selected to enhance the SAR image and suppress the coherent speckle noise in the image. The calculation formula of the Lee filter method is as follows:

[0029] in Indicates that the Lee filtered image is at position The pixel value of Indicates that the original image is at position The pixel value of represents the local mean in the neighborhood, represents the local variance within the neighborhood, Indicates the grayscale level of the image, represents the estimated speckle noise variance.

[0030] The specific steps of constructing the MALCP-Net network model based on the improved Cascade-Mask-RCNN model are as follows: Construct a backbone network and select ResNet50 as the backbone network to extract the depth features of SAR images. Add a multi-scale hole layered convolution module MALCM to the output of the backbone network. The specific structure of the MALCM module is as follows: Figure 3 As shown in Figure 2, it is used to extract multi-scale features from the feature map output by ResNet50 and enhance the hierarchical nature of features by fusing features of different scales. To build a neck network, we used a feature pyramid network (FPN) and embedded the SE attention module on top of the FPN to form the SE-FPN network. This network enhances the multi-scale feature fusion effect and highlights important feature information through the attention mechanism. We then added the triple attention mechanism module (Tri-Attention) after the SE-FPN network to capture cross-dimensional interaction information, improving the expressiveness of features and the efficiency of spatial information utilization. A detection head is constructed, using the proposal generation and imitation learning network model CFINet to detect ship targets and generate candidate boxes. The proposal generation and imitation learning network model CFINet can enhance the detection effect of small targets and improve detection accuracy through the coarse-to-fine region proposal network CRPN and feature imitation FI branch. Then, an instance segmentation mask branch is added to CFINet to further extract and upsample the features within each candidate box through a series of convolutional and deconvolution layers, generate pixel-level masks corresponding to the candidate box, and achieve instance segmentation of ships. The backbone network, neck network and detection head constructed above are connected and integrated to obtain the MALCP-Net network model.

[0031] The above steps adopt the proposal generation and imitation learning network model CFINet as the detection head. It can not only effectively improve the detection accuracy and recall rate of small ship targets in SAR images by the MALCP-Net network model with the help of coarse-to-fine region proposal and feature imitation learning mechanism, but also generate high-quality candidate boxes with the help of dynamic anchor selection strategy and cascade regression to reduce the amount of subsequent processing calculations. On the basis of CFINet, an instance segmentation mask branch is added to further perform pixel-level segmentation on the detected ship targets and generate accurate masks to achieve accurate segmentation of the targets, thereby improving the detection accuracy and segmentation effect of this ship instance segmentation method.

[0032] The structure of the multi-scale atrous layered convolution module MALCM is as follows: The channel conversion convolution layer is used to convert the number of channels of the feature map output by ResNet50 into a number of channels suitable for subsequent processing; The branch processing layer is used to copy the feature map after channel conversion into four copies and input them into four branches respectively, where branch 1 is the standard convolution layer, i.e. Convolution layer, branch 2 is a dilated convolution layer with a dilation rate of 3, branch 3 is a dilated convolution layer with a dilation rate of 6, and branch 4 is a dilated convolution layer with a dilation rate of 9. Before the dilated convolution operation, branches 2 and 4 need to add the results of the dilated convolution operation of branches 1 and 3 respectively. The output calculation formula of each branch is as follows:

[0033] in Indicates the The input feature map of the branch, and , Express The expansion rate is The dilated convolution calculation, Indicates the Output feature map of each branch; The feature fusion layer is used to splice the output feature maps of the four branches.

[0034] The above steps add a multi-scale atrous layered convolution module MALCM to the output part of the backbone network. With the help of the multi-scale atrous convolution and lateral additive connection structure in the MALCM module, it can effectively capture features of different scales and enhance the feature hierarchy, significantly improving the MALCP-Net network model's ability to detect and segment multi-scale ship targets, especially small ship targets, and helping the MALCP-Net network model to accurately locate and identify ships in complex backgrounds, thereby improving the detection accuracy and segmentation effect of this ship instance segmentation method.

[0035] The SE-FPN network is mainly composed of a feature pyramid network (FPN) and an SE attention module embedded therein. The feature pyramid network (FPN) includes bottom-up and top-down paths as well as lateral connections to generate feature maps of different scales, while the SE attention module is embedded in each stage of the FPN to perform attention weighting on the channel dimension of the feature map, thereby enhancing the response of important feature channels and suppressing unimportant channels.

[0036] The triple attention mechanism module Tri-Attention consists of three parallel attention branches, corresponding to the attention calculations in the channel-height, channel-width and spatial dimensions respectively. Each branch contains a rotation operation, a Z-pool layer, a convolution and normalization layer, as well as activation and application operations, which are used to simultaneously capture the complex interactions between the channel and spatial dimensions and comprehensively model the dependencies between features.

[0037] Specifically, for the input tensor, a rotation operation is first performed to adjust the dimensional order of the tensor, and then the maximum pooling and average pooling operations are performed through the Z-pool layer and the results are spliced. Then, the convolution and normalization layers are used to extract features and generate attention weights. Finally, the attention weights are applied to the original tensor through the activation function and rotated back to the original dimension.

[0038] The above steps embed the SE attention module on the basis of the feature pyramid network FPN and add the triple attention mechanism module Tri-Attention after the SE-FPN network. It can not only use the channel attention mechanism to weight the channel dimension of the feature map, thereby enhancing the response of important feature channels and suppressing unimportant channels to improve feature expression ability and multi-scale feature fusion effect, but also use the Tri-Attention module to simultaneously capture the complex interaction between channels and spatial dimensions, and comprehensively model the dependencies between features, thereby enhancing the expression ability of features and the utilization efficiency of spatial information, so as to improve the accuracy and robustness of the MALCP-Net network model. The combination of the two not only enhances the expression ability of features, but also makes the MALCP-Net network model use features more fully and accurately, further improving the performance of the MALCP-Net network model in tasks such as target detection and segmentation, and improving the detection accuracy and segmentation effect of this ship instance segmentation method.

[0039] The proposal generation and imitation learning network model CFINet mainly includes a coarse-to-fine region proposal network CRPN and a feature imitation FI branch, wherein the coarse-to-fine region proposal network CRPN is used to generate high-quality candidate boxes for small ship targets through a dynamic anchor selection strategy and cascade regression; the feature imitation FI branch is mainly composed of an example feature set and a feature-to-embedding module, the example feature set is used to retain high-quality ROI features, and the feature-to-embedding module is used to map regional features to the embedding space.

[0040] The specific steps of training the constructed MALCP-Net network model using the acquired enhanced image dataset are as follows: Data preparation and preprocessing: Obtain an enhanced image dataset, and divide the obtained enhanced image dataset into a training set, a validation set, and a test set according to the ratio of 70%, 15%, and 15%, respectively; Forward propagation and loss calculation, input the sample data in the training set into the MALCP-Net network model to obtain the instance segmentation result of the ship, using the multi-task loss function Calculate the loss between the instance segmentation result and the true label; Back propagation and parameter update, using the chain rule to reversely calculate the gradient of each layer parameter from the loss, and using the stochastic gradient descent optimization algorithm SGD to adjust the parameters to minimize the loss; Iterative optimization: divide the training set into multiple small batches, repeat the previous two steps batch by batch, and traverse the entire training set multiple times until the MALCP-Net network model converges or reaches the preset number of iterations; Verification and optimization: The MALCP-Net network model is periodically evaluated using sample data from the validation set. Various error metrics, such as average precision (AP) and average recall (AR), are calculated on the validation set to evaluate the model's fitting and generalization capabilities. Based on the validation set evaluation results, the model's hyperparameters, such as the learning rate and weight coefficient, are adjusted until the segmentation performance reaches the preset standard. Model testing: After training is completed, the MALCP-Net network model is finally evaluated using sample data in the test set, and various error indicators are calculated to ensure that the MALCP-Net network model still has good detection and segmentation performance on unseen data.

[0041] The multi-task loss function The calculation formula is:

[0042] in represents the bounding box regression loss, calculated using the Smooth-L1 loss function, represents the classification loss, which is calculated using the cross entropy loss function, and It represents the total detection loss, Represents the total segmentation loss, calculated using the Dice loss function. and Represent the total detection loss and the total loss of segmentation The weight coefficient can be set according to actual needs, and .

[0043] The above steps are achieved by adopting a multi-task loss function in the training process of the MALCP-Net network model. By calculating the loss between the instance segmentation result and the true label, both detection loss and segmentation loss can be considered simultaneously, enabling the MALCP-Net network model to simultaneously optimize both target detection and instance segmentation tasks. By introducing weight coefficients to balance the contributions of different tasks, the MALCP-Net network model can more comprehensively learn the characteristics of ship targets, thereby improving the MALCP-Net network model's ability to detect and segment multi-scale ship targets in complex backgrounds, and enhancing the detection accuracy and segmentation effect of this ship instance segmentation method.

[0044] It should be noted that the quantitative comparative experimental results of the MALCP-Net network model and other instance segmentation models on the HRSID dataset are shown in Table 1: Table 1 Comparison of quantitative experimental results of instance segmentation models on the HRSID dataset

[0045] in represents the average precision of the instance segmentation model, and They represent the average precision of the instance segmentation model when the IoU threshold is 0.5 and 0.75, respectively. 、 and Represents the average precision of instance segmentation model for small objects, medium objects and large objects respectively.

[0046] It should be noted that the comparison of instance segmentation results of the MALCP-Net network model and the Cascade-Mask-RCNN model on the HRSID dataset is shown in the figure below. Figure 4As shown in the figure, each row of pictures represents from left to right: the real ground of the correct instance segmentation result of the ship target in the figure, the instance segmentation result of the Cascade-Mask-RCNN model and the instance segmentation result of the MALCP-Net network model, and the part circled by the yellow ring in the figure represents the situation where the MALCP-Net network model can detect the ship target but the Cascade-Mask-RCNN model fails to detect the ship target, and the part circled by the blue ring represents the situation where the Cascade-Mask-RCNN model has a false detection, while the MALCP-Net network model has no false detection.

[0047] It should be understood that those skilled in the art may make improvements or modifications based on the above description, and all such improvements and modifications shall fall within the scope of protection of the appended claims. Any portion of this specification not described in detail is prior art known to those skilled in the art.

Claims

1. A ship instance segmentation method based on multi-scale convolutional neural network, characterized in that: The following steps are involved: Image acquisition and processing: using synthetic aperture radar (SAR) to acquire an image dataset containing ship targets, and then labeling, cropping, and scaling the SAR images in the image dataset to obtain a pre-processed image dataset. Image enhancement: performing image enhancement and normalization processing on the SAR images in the preprocessed image dataset to obtain an enhanced image dataset; Model construction: construct the MALCP-Net network model based on the improved Cascade-Mask-RCNN model; Model training: The constructed MALCP-Net network model is trained using the acquired enhanced image dataset to obtain a ship instance segmentation model; Instance segmentation: Preprocess and enhance the SAR images acquired in real time. The processed SAR images are input into the ship instance segmentation model and instance segmentation is performed to obtain the ship instance segmentation results.

2. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 1 is characterized in that: The specific steps of performing image enhancement on the SAR image in the preprocessed image dataset are: The Lee filter method is selected to enhance the SAR image and suppress the coherent speckle noise in the image. The calculation formula of the Lee filter method is as follows: in Indicates that the Lee filtered image is at position The pixel value of Indicates that the original image is at position The pixel value of represents the local mean in the neighborhood, represents the local variance within the neighborhood, Indicates the grayscale level of the image, represents the estimated speckle noise variance.

3. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 2 is characterized in that: The specific steps of constructing the MALCP-Net network model based on the improved Cascade-Mask-RCNN model are as follows: Construct a backbone network, selecting ResNet50 as the backbone network to extract deep features of SAR images. Add a multi-scale atrous layered convolution module (MALCM) to the output of the backbone network to extract multi-scale features from the feature maps output by ResNet50 and enhance the hierarchical nature of features by fusing features at different scales. To build a neck network, we used a feature pyramid network (FPN) and embedded the SE attention module on top of the FPN to form the SE-FPN network. This network enhances the multi-scale feature fusion effect and highlights important feature information through the attention mechanism. We then added the triple attention mechanism module (Tri-Attention) after the SE-FPN network to capture cross-dimensional interaction information, improving the expressiveness of features and the efficiency of spatial information utilization. A detection head is constructed, using the proposal generation and imitation learning network model CFINet to detect ship targets and generate candidate boxes. An instance segmentation mask branch is then added to CFINet to further extract and upsample features within each candidate box through a series of convolutional and deconvolutional layers, generating pixel-level masks corresponding to the candidate boxes and achieving instance segmentation of ships. The backbone network, neck network and detection head constructed above are connected and integrated to obtain the MALCP-Net network model.

4. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 3 is characterized in that: The structure of the multi-scale atrous layered convolution module MALCM is as follows: The channel conversion convolution layer is used to convert the number of channels of the feature map output by ResNet50 into a number of channels suitable for subsequent processing; The branch processing layer is used to copy the feature map after channel conversion into four copies and input them into four branches respectively, where branch 1 is the standard convolution layer, i.e. Convolution layer, branch 2 is a dilated convolution layer with a dilation rate of 3, branch 3 is a dilated convolution layer with a dilation rate of 6, and branch 4 is a dilated convolution layer with a dilation rate of 9. Before the dilated convolution operation, branches 2 and 4 need to add the results of the dilated convolution operation of branches 1 and 3 respectively. The output calculation formula of each branch is as follows: in Indicates the The input feature map of the branch, and , Express The expansion rate is The dilated convolution calculation, Indicates the Output feature map of each branch; The feature fusion layer is used to splice the output feature maps of the four branches.

5. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 4 is characterized in that: The SE-FPN network is mainly composed of a feature pyramid network (FPN) and an SE attention module embedded therein. The feature pyramid network (FPN) includes bottom-up and top-down paths as well as lateral connections to generate feature maps of different scales, while the SE attention module is embedded in each stage of the FPN to perform attention weighting on the channel dimension of the feature map, thereby enhancing the response of important feature channels and suppressing unimportant channels.

6. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 5 is characterized in that: The triple attention mechanism module Tri-Attention consists of three parallel attention branches, corresponding to the attention calculations in the channel-height, channel-width and spatial dimensions respectively. Each branch contains a rotation operation, a Z-pool layer, a convolution and normalization layer, as well as activation and application operations, which are used to simultaneously capture the complex interactions between the channel and spatial dimensions and comprehensively model the dependencies between features.

7. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 6 is characterized in that: The proposal generation and imitation learning network model CFINet mainly includes a coarse-to-fine region proposal network CRPN and a feature imitation FI branch, wherein the coarse-to-fine region proposal network CRPN is used to generate high-quality candidate boxes for small ship targets through a dynamic anchor selection strategy and cascade regression; the feature imitation FI branch is mainly composed of an example feature set and a feature-to-embedding module, the example feature set is used to retain high-quality ROI features, and the feature-to-embedding module is used to map regional features to the embedding space.

8. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 7 is characterized in that: The specific steps of training the constructed MALCP-Net network model using the acquired enhanced image dataset are as follows: Data preparation and preprocessing: Obtain an enhanced image dataset, and divide the obtained enhanced image dataset into a training set, a validation set, and a test set according to the ratio of 70%, 15%, and 15%, respectively; Forward propagation and loss calculation, input the sample data in the training set into the MALCP-Net network model to obtain the instance segmentation result of the ship, using the multi-task loss function Calculate the loss between the instance segmentation result and the true label; Back propagation and parameter update, using the chain rule to reversely calculate the gradient of each layer parameter from the loss, and using the stochastic gradient descent optimization algorithm SGD to adjust the parameters; Iterative optimization: divide the training set into multiple small batches, repeat the previous two steps batch by batch, and traverse the entire training set multiple times until the MALCP-Net network model converges or reaches the preset number of iterations; Verification and optimization: Periodically evaluate the MALCP-Net network model using sample data from the validation set, calculate various error metrics of the MALCP-Net network model on the validation set, and use this to evaluate the fitting effect and generalization ability of the MALCP-Net network model. Based on the evaluation results of the validation set, adjust the hyperparameters of the MALCP-Net network model until the segmentation performance of the MALCP-Net network model meets the preset standards. Model testing,After training is completed, the final performance evaluation of the MALCP-Net network model is performed,and various error indicators are calculated using the sample data in the test set.

9. The ship instance segmentation method based on a multi-scale convolutional neural network according to claim 8, characterized in that: The multi-task loss function The calculation formula is: in represents the bounding box regression loss, calculated using the Smooth-L1 loss function, represents the classification loss, which is calculated using the cross entropy loss function, and It represents the total detection loss, Represents the total segmentation loss, calculated using the Dice loss function. and Represent the total detection loss and the total loss of segmentation The weight coefficient of .

Citation Information

Cited By

  • Breast cancer ultrasonic image intelligent segmentation method and system based on weak supervised learning

    CN121999218A

  • Adaptive ship instance segmentation method based on feature decoupling

    CN122244452A

  • An adaptive ship instance segmentation method based on feature decoupling

    CN122244452B