Ship image segmentation method and system based on spatial pyramid feature enhancement

By adopting a deep neural network of spatial pyramids in SAR image segmentation, combining the backbone network and spatial pyramid feature enhancement network, the problem of excessive background information brought about by ship segmentation in SAR images is solved, the accuracy and efficiency of segmentation are improved, and better scheduling and planning guidance is provided for shipping companies.

CN120070473APending Publication Date: 2025-05-30SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226466.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing SAR image ship segmentation method is prone to bring too much useless background information in remote sensing images, resulting in complex segmentation and low accuracy. At the same time, it is not possible to effectively consider whether the image meets the segmentation conditions, resulting in segmentation errors and inefficiency.

Method used

The ship image segmentation method based on spatial pyramid feature enhancement is adopted to segment the ship image through the spatial pyramid deep neural network, and the backbone network and the spatial pyramid feature enhancement network are used to capture the global spatial background and enhance the features, so that only useful global information is fused during the segmentation process.

Benefits of technology

It improves the accuracy and efficiency of ship image segmentation, reduces the occurrence of segmentation errors, and provides better guidance for shipping companies to optimize ship scheduling and route planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070473A_ABST
    Figure CN120070473A_ABST
Patent Text Reader

Abstract

The invention provides a ship image segmentation method and system based on spatial pyramid feature enhancement, and belongs to the technical field of synthetic aperture radar (SAR) ship image segmentation. The method comprises the following steps: firstly, performing data compliance verification on a synthetic aperture radar image data set, and training a spatial pyramid deep neural network by using the image data set conforming to the data compliance verification; and finally, inputting a ship image to be segmented into the trained spatial pyramid deep neural network for image segmentation so as to segment a ship from the ship image. According to the method, the features can be further enhanced through the global space background in each layer of the network structure, so that only useful global information is fused into a local area, the difference between objects is not reduced, and guidance is provided for a shipping company to optimize ship scheduling and route planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of synthetic aperture radar (SAR) ship image segmentation, and particularly relates to a ship image segmentation method and system based on enhanced spatial pyramid features. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Synthetic aperture radar (SAR) is a microwave imaging sensor. It is based on the electromagnetic wave scattering characteristics, can be used in all weather conditions, and has a certain ability to penetrate clouds and the ground. With the continuous development of marine resources and the increasing emphasis on marine ship monitoring, it has special benefits in marine monitoring, mapping, military, and all these fields. Therefore, synthetic aperture radar ship detection technology is of great significance for protecting the marine ecosystem, marine law enforcement, and territorial sea security. Marine ship monitoring has received extensive attention. Synthetic aperture radar (SAR) is more suitable for monitoring ocean-going ships than optical sensors because it can work in all weather conditions. Ship monitoring is a key marine task and is crucial for marine surveillance, national defense security, fishery management identification, etc. Instance segmentation is an important, complex, and challenging field in machine vision research. It features solving both object recognition and semantic segmentation problems simultaneously. In parallel with semantic segmentation, it has both pixel-level classification and object recognition attributes. Even if different instances belong to the same type, different instances must be located. For synthetic aperture radar (SAR) images, ship detection is a challenging task. Ships in SAR images have arbitrary directions and multi-scale features. Therefore, it is crucial to develop a SAR image ship instance segmentation system, which can assist in the detection and identification of ocean-going ships, and can help monitor shipping activities, optimize route planning, predict traffic congestion, etc., thereby improving the efficiency and safety of shipping management.

[0004] However, there are not many existing methods and systems specifically for and adapted to SAR image ship segmentation; moreover, the following technical problems generally exist:

[0005] (1) Existing SAR image ship segmentation methods usually integrate multiple spatial or channel attention blocks into the backbone network based on ordinary neural networks to achieve a global receptive field. However, in remote sensing images, objects actually only cover a very small area (especially ship images in this field). Therefore, this kind of segmentation may bring too much useless background information to the objects, making the segmentation of ships in the image more complex and the segmentation accuracy not high.

[0006] (2) Only considering the immediacy of ship segmentation in SAR images, that is, directly segmenting the acquired images to be segmented without considering whether the SAR ship images to be segmented meet the segmentation conditions, resulting in frequent segmentation errors and low segmentation efficiency. Summary of the Invention

[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a ship image segmentation method and system based on enhanced spatial pyramid features, which can further enhance features through the global spatial background in each layer of the network structure to ensure that only useful global information is fused into the local area without reducing the differences between objects, thereby providing guidance for shipping companies to optimize ship scheduling and route planning.

[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0009] The first aspect of the present invention provides a ship image segmentation method based on enhanced spatial pyramid features.

[0010] A ship image segmentation method based on enhanced spatial pyramid features includes:

[0011] Obtain a synthetic aperture radar image dataset;

[0012] Perform data compliance verification on the obtained image dataset, and use the image dataset that meets the data compliance verification to train a spatial pyramid deep neural network; wherein, the spatial pyramid deep neural network includes a backbone network and a spatial pyramid feature enhancement network;

[0013] Input the ship image to be segmented into the trained spatial pyramid deep neural network for image segmentation to segment the ship from the ship image.

[0014] Further, the data compliance verification includes size verification, binary map verification, class label verification, bounding box verification, and input image existence verification.

[0015] Further, each operation of the data compliance verification has different verification passing criteria. Specifically, the verification passing criteria for size verification, binary map verification, class label verification, bounding box verification, and input image existence verification are that the sizes of the image and the mask are consistent, the corresponding mask is a binary image, the class label is within the standard range, the bounding box is within the image range, and the input image is not empty.

[0016] Further, data compliance verification is performed by calling OpenCV functions, NumPy functions, and OS functions.

[0017] Further, the backbone network of the spatial pyramid deep neural network is a Resnet50 network, and the spatial pyramid feature enhancement network is a network with a multi-layer pyramid structure; specifically, each layer of the spatial pyramid feature enhancement network consists of a context aggregation block.

[0018] Further, the pixel spatial context in the context aggregation block is aggregated in the following manner, that is:

[0019]

[0020] where P m and Q m respectively represent the input feature map and the output feature map of the m-th layer of the spatial pyramid feature enhancement network, a m represents the reweighting matrix, B m represents the number of pixels, W k represents the linear projection matrix for information comparison, and W e represents the linear projection matrix for aggregation.

[0021] The second aspect of the present invention provides a ship image segmentation system based on spatial pyramid feature enhancement.

[0022] A ship image segmentation system based on spatial pyramid feature enhancement, adopting the ship image segmentation method described in the first aspect, includes:

[0023] A segmentation preparation module, configured to: build a user interface and integrate the user interface with VTK; create a reader to read the ship image to be segmented;

[0024] An image segmentation module, configured to: input the ship image to be segmented into the trained spatial pyramid deep neural network for image segmentation to segment the ship from the ship image;

[0025] An image processing module, configured to: create a mapper and a VTK renderer; convert different data types into image data based on the mapper, and perform a rendering operation on the image data based on the VTK renderer;

[0026] An image display module, configured to: call a painter and a painting window in VTK based on an interactor, and display the original image and the segmented image of the ship image to be segmented after the rendering operation on the user interface of the computer platform simultaneously.

[0027] Further, a plurality of function controls are provided on the built user interface, and the function controls include importing an image, performing segmentation, and clearing the interface.

[0028] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in a ship image segmentation method based on enhanced spatial pyramid features as described in the first aspect of the present invention are implemented.

[0029] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in a ship image segmentation method based on enhanced spatial pyramid features as described in the first aspect of the present invention are implemented.

[0030] The above one or more technical solutions have the following beneficial effects:

[0031] (1) The present invention segments ship images through a spatial pyramid deep neural network. Among them, the spatial pyramid deep neural network includes a backbone network and a spatial pyramid feature enhancement network (SCP network); the spatial pyramid feature enhancement network is a multi-layer pyramid-structured network; specifically, each layer of the spatial pyramid feature enhancement network consists of a context aggregation block. The present invention can capture the global spatial background at each feature pyramid level through the SCP network. This network learns and aggregates features from the entire feature map and combines them into each pixel using adaptive weights. Therefore, this strategy of the present invention can ensure that only useful global information is fused into the local area without reducing the differences between objects. Therefore, the segmentation of ships in the image becomes simpler and the segmentation accuracy is high, which can provide better guidance for shipping companies to optimize ship scheduling and route planning.

[0032] (2) Before segmenting the image, the present invention will first perform data compliance verification on the obtained image dataset, and then use the image dataset that meets the data compliance verification to train the spatial pyramid deep neural network. This well ensures that the SAR ship images on which the segmentation operation is performed meet the segmentation conditions. Therefore, compared with the prior art, the situation of segmentation errors in the present invention will be greatly reduced, and the segmentation efficiency will be significantly improved.

[0033] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0034] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0035] Figure 1 It is a flowchart of a ship image segmentation method based on enhanced spatial pyramid features in Embodiment 1 of the present invention.

[0036] Figure 2 This is a schematic diagram of the connection structure of the spatial pyramid depth neural network in the first embodiment of the present invention.

[0037] Figure 3 This is a schematic diagram of the structure of the SCP network in the first embodiment of the present invention.

[0038] Figure 4 This is a schematic diagram of the structure of the context aggregation block in the first embodiment of the present invention.

[0039] Figure 5 This is a schematic diagram of the relationship between the functional modules of a ship image segmentation system based on spatial pyramid feature enhancement in the second embodiment of the present invention.

[0040] Figure 6 This is a schematic diagram of the user interface before importing the ship image to be segmented in the second embodiment of the present invention.

[0041] Figure 7 This is a schematic diagram of the user interface after importing the ship image to be segmented and before performing the segmentation operation in the second embodiment of the present invention.

[0042] Figure 8 This is a schematic diagram of the user interface after performing the segmentation operation in the second embodiment of the present invention. Detailed implementation manners

[0043] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0044] It should be noted that the terms used herein are only for describing the specific implementation manners and are not intended to limit the exemplary embodiments of the present invention.

[0045] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0046] Embodiment 1

[0047] This embodiment discloses a ship image segmentation method based on spatial pyramid feature enhancement.

[0048] As Figure 1 shown, a ship image segmentation method based on spatial pyramid feature enhancement includes:

[0049] Step S1, obtaining a synthetic aperture radar image data set;

[0050] Step S2: Perform data compliance verification on the obtained image dataset, and use the image dataset that meets the data compliance verification to train the spatial pyramid deep neural network; wherein, the spatial pyramid deep neural network includes a backbone network and a spatial pyramid feature enhancement network;

[0051] Step S3: Input the ship image to be segmented into the trained spatial pyramid deep neural network for image segmentation, so as to segment the ship from the ship image.

[0052] Based on the above process, the present invention can further enhance features through the global spatial background in each layer of the network structure, so as to ensure that only useful global information is fused into the local area without reducing the differences between objects, thereby providing guidance for shipping companies to optimize ship scheduling and route planning. For the convenience of understanding the technical solution of the present invention, the following further explains and illustrates the specific implementation steps in the technical solution of the present invention.

[0053] Step S1: Obtain a synthetic aperture radar image dataset.

[0054] The synthetic aperture radar image dataset used in this embodiment is a public dataset, namely: SatelliteShip Detection Dataset (SSDD). The SSDD dataset is a public dataset dedicated to ship target detection in SAR images; it can be used to train and test detection algorithms, enabling researchers to compare algorithm performance under the same conditions.

[0055] The SSDD dataset has a total of 1160 images and 2456 ships. Its pictures are in jpg format, with an average of 2.12 ships per image, and the resolution varies between 1 and 15 meters. This dataset assigns two categories, namely background (category 0) and ship (category 1); it adopts the same format as the Microsoft Common Objects in Context (MS COCO) dataset and is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1, that is, it includes three JavaScript Object Notation (json) files for training, validation, and testing.

[0056] Step S2: Perform data compliance verification on the obtained image dataset, and use the image dataset that meets the data compliance verification to train the spatial pyramid deep neural network; wherein, the spatial pyramid deep neural network includes a backbone network and a spatial pyramid feature enhancement network.

[0057] Step S2-1: Perform data compliance verification on the obtained image dataset.

[0058] In the prior art, generally only the immediacy of ship segmentation in SAR images is considered, that is, the to-be-segmented images taken are directly segmented, without considering whether the SAR ship images to be segmented meet the segmentation conditions, resulting in frequent segmentation errors and low segmentation efficiency. In response to this, the present invention chooses to add a data compliance verification link to ensure that the SAR ship images to be segmented meet the segmentation conditions, thereby reducing the occurrence of segmentation errors.

[0059] Further, in this embodiment, the data compliance verification operation mainly includes five types of verifications, namely: size verification, binary map verification, class label verification, bounding box verification, and input image existence verification. Among them, each operation of data compliance verification has different verification passing criteria. Specifically, the verification passing criteria for size verification, binary map verification, class label verification, bounding box verification, and input image existence verification are that the sizes of the image and the mask are consistent, the corresponding mask is a binary image, the class label is within the standard range, the bounding box is within the image range, and the input image is not empty. In the actual implementation process, these verification links are used to perform data compliance verification by calling OpenCV functions, NumPy functions, and OS functions, and can be specifically implemented through the following process.

[0060] Regarding checking whether the sizes of the image and the mask match, specifically: First, read the image and the mask using the OpenCV library of Python, obtain the sizes of the image and the mask, and compare them; if they do not match, an exception is thrown or an error is recorded, and finally, it is ensured that the sizes of the image and the corresponding mask are consistent.

[0061] Regarding checking whether the mask is a binary map, specifically: Use the NumPy library to check whether the pixel values of the mask are only two; if the mask is not a binary image, an exception is thrown or an error is recorded, and finally, it is ensured that the mask is a binary image.

[0062] Regarding checking whether the class label is within a reasonable range, specifically: The range of the class label is [0, 2], and use the NumPy library to check whether the label is within this range; if the label exceeds the range, an exception is thrown or an error is recorded, and finally, it is ensured that the class label is within the predefined range.

[0063] Regarding checking the bounding box, specifically, check whether the bounding box (if any) is within the image range to ensure that the coordinates of the bounding box are within the size range of the image. Regarding checking whether the input image is empty, specifically: Use the OS library of Python to check the file size; if the file size is 0, an exception is thrown or an error is recorded, and finally, it is ensured that the image file is not empty.

[0064] Finally, traverse each sample in the image dataset and sequentially execute the above verification steps. If the verification of a certain sample fails, record the error and skip that sample. Ultimately, ensure the integrity and accuracy of the data, thereby avoiding errors caused by data problems during model training.

[0065] Step S2-2: Train the spatial pyramid deep neural network using the image dataset that complies with data compliance verification.

[0066] After step S2-1, it can be determined whether the data in the image dataset complies with data compliance verification; when it complies with data compliance verification, input the ship images included in the training set, validation set, and test set in the image dataset into the spatial pyramid deep neural network respectively, and perform model training, validation, and testing in sequence.

[0067] Existing SAR image ship segmentation methods usually integrate multiple spatial or channel attention blocks into the backbone network based on ordinary neural networks to achieve a global receptive field. However, in remote sensing images, objects actually only cover a very small area (especially ship images within the scope of this field). Therefore, this kind of segmentation may bring too much useless background information to the objects, making the segmentation of ships in the image more complex and the segmentation accuracy not high. In response to this, the present invention provides a new neural network model to improve this problem existing in the existing model. Therefore, the present invention optimizes the existing deep neural network architecture for ship image segmentation, but the training, validation, and testing of the model based on the dataset still adopt the existing technology. Therefore, this embodiment does not elaborate and explain too much on the training, validation, and testing processes of the model here.

[0068] Furthermore, the present invention makes improvements on the basis of the basic neural network model, and the improved network structure diagram is Figure 2 the spatial pyramid deep neural network shown. As Figure 2As shown in the figure, ROI is a method for solving the problems that occur in ROIPooling. It can extract features in a more precise way, avoid the loss of position accuracy caused by pooling operations, and thus improve the accuracy. The specific implementation method is as follows: First, each ROI is divided into a fixed number of grids, and the grid coordinates are kept as more precise decimal coordinates. Then, for each position in each grid, bilinear interpolation is used to obtain the corresponding value from the feature map. Finally, all the interpolated feature values are integrated into an output of a fixed size as the feature representation of the ROI. Boxes represent the bounding boxes of the candidate regions, which are used to frame the possible target regions in the image and will be further refined later. Class represents the class label of each predicted region. Mask generation and refinement are divided into mask generation and mask refinement, which are used to generate accurate object segmentation results. Specifically, mask generation is to generate a pixel-level binary mask for each region candidate box (ROI) through a convolutional generation network. Mask refinement is to further improve the mask quality after mask generation to make it more accurate. The implementation method of mask refinement is as follows: Through more convolutional layers, the resolution of the mask is gradually improved.

[0069] Previously, attempts in this field usually integrated multiple spatial or channel attention blocks into the backbone network to achieve a global receptive field. Since the feature pyramid still contains spatially local information after aggregating feature maps at different levels. Therefore, the present invention proposes an improved deep neural network model, namely, a spatial pyramid deep neural network. The spatial pyramid deep neural network includes a backbone network and a spatial pyramid feature enhancement network.

[0070] Specifically, the backbone network of the spatial pyramid deep neural network is the Resnet50 network, and a Spatial Context Pyramid (SCP) network, that is, a spatial pyramid feature enhancement network, is introduced after the backbone network Resnet50. The structure of the SCP network is as Figure 3 shown. This network module is also a pyramid structure, so it can be easily inserted after the backbone or the neck. As Figure 3 shown, where CABlock is a context aggregation block; P and Q are respectively the input and output of the feature pyramid, and their subscripts are the input or output of the nth layer; the input P m , passes through the CABlock to obtain Q m , in order to obtain high-precision features.

[0071] The spatial pyramid feature enhancement network is a multi-layer pyramid-structured network. Specifically, each layer of the spatial pyramid feature enhancement network is composed of a context aggregation block (CABlock) with a residual connection. The detailed design of this network module is as Figure 4 shown. As Figure 4As shown, Sigmoid is used in the output layer for binary classification tasks. It is an activation function that can compress the output value of the model into the range of 0 to 1, representing the probability that each pixel belongs to a certain class. Conv represents the convolutional layer, and its convolutional kernel is 1×1 (the convolutional kernel is a small matrix) to operate on the input feature map, making the accuracy higher. Softmax is also an activation function that can convert the output of the network into the probability distribution of each class, ensuring that the sum of the probabilities that each pixel belongs to all classes is 1. A×B×C represents that the number of channels of the matrix is A, the height is B, and the width is C. Denotes batch matrix multiplication, and ⊙ denotes broadcast Hadamard product.

[0072] In each block, pixel spatial context is aggregated in the following way, that is:

[0073]

[0074] Among them, P m and Q m respectively represent the input feature map and the output feature map of the m-th layer of the spatial pyramid feature enhancement network, a m represents the reweighting matrix, B m represents the number of pixels, W k represents the linear projection matrix used for information comparison, W e represents the linear projection matrix used for aggregation.

[0075] The SCP network further enhances features by learning the global spatial background in each layer. It is found that in remote sensing images, objects only cover a very small area, and this design may bring too much useless background information to the objects. To solve this problem, it is proposed to add an additional path on this structure to learn the information volume of each pixel. The core idea of the present invention is that if the feature information volume of a pixel point is large enough, then it is less necessary to summarize the feature information of other spatial positions. Specifically, the SCP network can capture the global spatial background at each feature pyramid level. This module learns the aggregated features from the entire feature map and combines them into each pixel using adaptive weights. This strategy ensures that only useful global information is fused into the local area without reducing the differences between objects.

[0076] Step S3: Input the ship image to be segmented into the trained spatial pyramid deep neural network for image segmentation to segment the ship from the ship image.

[0077] Embodiment 2

[0078] This embodiment discloses a ship image segmentation system based on spatial pyramid feature enhancement.

[0079] AsFigure 5 As shown in Figure 5 , a ship image segmentation system based on enhanced spatial pyramid features adopts the ship image segmentation method described in Embodiment 1, including:

[0080] A segmentation preparation module, configured to: build a user interface and integrate the user interface with VTK; create a reader to read the ship image to be segmented;

[0081] An image segmentation module, configured to: input the ship image to be segmented into a trained spatial pyramid deep neural network for image segmentation to segment the ship from the ship image;

[0082] An image processing module, configured to: create a mapper and a VTK renderer; convert different data types into image data based on the mapper, and perform a rendering operation on the image data based on the VTK renderer;

[0083] An image display module, configured to: call a painter and a painting window in VTK based on an interactor, and simultaneously display the original image and the segmented image of the ship image to be segmented after the rendering operation on the user interface of the computer platform.

[0084] Based on the module design in the above system, the present invention can not only achieve accurate segmentation of the ship in the ship image, but also display the segmented ship image on the user interface of the computer to enhance visibility, so as to better facilitate providing guidance for shipping companies to optimize ship scheduling and route planning.

[0085] Further, a plurality of function controls are provided on the built user interface, and the function controls include importing an image, executing segmentation, and clearing the interface. The functions and implementations of each module in this embodiment are further explained and described below. Specifically:

[0086] Regarding the segmentation preparation module, its core function is to construct a user interaction interface, integrate the image reading function, and provide a data input and operation entry for subsequent processing. In specific implementation, the segmentation preparation module is mainly implemented through the user interface construction and the image reader process:

[0087] Further, the user interface construction is as follows: First, use the PyQt5 framework to create a main window, and the layout adopts a vertical arrangement. Set three function buttons at the top of the interface: "Import Image" is used to load the file to be processed, "Execute Segmentation" triggers the segmentation algorithm, and "Clear Interface" resets all displayed contents. Subsequently, embed a VTK rendering window component in the middle of the interface to achieve seamless integration with PyQt. Finally, bind the button click event and the corresponding processing function through the signal-slot mechanism.

[0088] Furthermore, the image reader is implemented to support common image formats (PNG, JPEG) and medical imaging formats (DICOM). First, the OpenCV library is used to read the image file, preserving the original pixel data and meta-information. Subsequently, the pixel data in the form of a Numpy array is converted into a vtkImageData object specific to VTK. The specific steps include: using the numpy_to_vtk function to convert the data format, setting the image dimensions (width, height, number of channels), and specifying the pixel data type (such as unsigned 8-bit integer).

[0089] Regarding the image segmentation module, its core function is to accurately segment the ship target through a pre-trained deep learning model and generate a binary mask. In terms of specific implementation, the image segmentation module is mainly achieved through network architecture design and model inference process:

[0090] Furthermore, the network architecture is designed as follows: It is improved based on the DeepLabv3+ model, and the backbone network uses ResNet-50 to extract multi-level features. A new Spatial Pyramid Pooling (SPP) module is added, which includes the following components: ① Adaptive average pooling layer: Downsample the feature map to a size of 4x4; ② 1x1 convolutional layer: Compress the number of channels to 512 dimensions; ③ Bilinear upsampling layer: Restore the feature map to the original input size. Finally, the output of the basic network and the output of the SPP module are concatenated in the channel dimension, and the segmentation result is generated through the final convolutional layer.

[0091] Furthermore, the model inference process is as follows: First, input preprocessing, for example: Normalize the image to the range [0,1], adjust the size to 512x512 pixels and convert it to the PyTorch tensor format (channel first). Immediately afterwards, load the pre-trained model, that is: Load the saved model weights (.pth file) from disk and set it to the evaluation mode (turn off training-specific layers such as Dropout). Subsequently, perform segmentation, that is: Obtain the class probabilities of each pixel through forward propagation, and generate a binary mask (0 - background, 1 - ship) by taking the class corresponding to the maximum probability. Finally, perform post-processing, that is: Apply morphological opening operation to eliminate small noises and improve the edge smoothness through boundary optimization processing.

[0092] Regarding the image processing module, its core function is to complete data format conversion and visualization parameter configuration to prepare for rendering and display. In terms of specific implementation, the image processing module is mainly achieved through data type conversion, color mapping configuration, and volume rendering parameter setting:

[0093] Furthermore, the data type conversion is as follows: Convert the segmentation mask from a Numpy array to a vtkImageData object of VTK and preserve the spatial coordinate information to ensure alignment with the original image.

[0094] Further, the color mapping is configured as follows: First, a vtkColorTransferFunction object is created to define the color mapping rules. Among them, the background (value 0) is set to transparent black, and the ship area (value 1) is set to semi-transparent green. Subsequently, the transparency curve is set to achieve a fade-out effect.

[0095] Further, the volume rendering parameters are set as follows: First, a GPU-accelerated ray casting algorithm is selected. Subsequently, the lighting parameters are configured, that is, a parallel light source is added to simulate natural light, and the ambient light intensity is adjusted to avoid overly dark areas.

[0096] Regarding the image display module, its core function is to realize the side-by-side comparison display of the original image and the segmentation result, and support interactive operations. In terms of specific implementation, the image processing module mainly realizes it through dual-viewport layout, interactive function implementation, and real-time rendering optimization:

[0097] Further, the dual-viewport layout is as follows: Two independent renderers are created, the left viewport is set to display the original image (0%-50% width), and the right viewport is set to display the segmentation result (50%-100% width).

[0098] Further, the implementation of the interactive function is as follows: The following operations are enabled through vtkRenderWindowInteractor: Mouse wheel zoom: Synchronously adjust the display ratio of the dual viewports; Left mouse button drag and pan: Keep the spatial positions of the two views consistent; Right mouse button menu: Provide window width / window level adjustment and measurement tools.

[0099] Further, the real-time rendering optimization is as follows: Hardware acceleration (OpenGL 4.5) is enabled, and the frame rate limit is set to 60 FPS to prevent over-rendering.

[0100] Based on the designs of the above-mentioned various modules, the various modules cooperate with each other. As Figure 6 、 Figure 7 、 Figure 8 shown are respectively the schematic diagrams of the user interface before importing the ship image to be segmented, after importing the ship image to be segmented and before performing the segmentation operation, and after performing the segmentation operation. From Figure 6 、 Figure 7 、 Figure 8 it is not difficult to see that the ship image segmentation system provided in this embodiment has a good segmentation effect, and thus can provide good guidance for shipping companies to optimize ship scheduling and route planning.

[0101] Embodiment 3

[0102] The purpose of this embodiment is to provide a computer-readable storage medium.

[0103] A computer-readable storage medium stores a computer program thereon. When the program is executed by a processor, it implements the steps in a ship image segmentation method based on spatial pyramid feature enhancement as described in Embodiment 1 of the present disclosure.

[0104] Embodiment 4

[0105] The purpose of this embodiment is to provide an electronic device.

[0106] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a ship image segmentation method based on spatial pyramid feature enhancement as described in Embodiment 1 of the present disclosure.

[0107] The steps involved in the devices in the above Embodiments 2, 3, and 4 correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0108] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0109] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they do not limit the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A ship image segmentation method based on spatial pyramid feature enhancement, characterized in that: include: Get synthetic aperture radar image dataset; Performing data compliance verification on the obtained image data set, and using the image data set that complies with the data compliance verification to train the spatial pyramid deep neural network; wherein the spatial pyramid deep neural network includes a backbone network and a spatial pyramid feature enhancement network; The ship image to be segmented is input into the trained spatial pyramid deep neural network for image segmentation to segment the ship from the ship image.

2. A ship image segmentation method based on spatial pyramid feature enhancement as claimed in claim 1, characterized in that: The data compliance verification includes size verification, binary image verification, category label verification, bounding box verification, and input image existence verification.

3. A ship image segmentation method based on spatial pyramid feature enhancement as claimed in claim 2, characterized in that: Each operation of the data compliance verification has different verification pass standards. Specifically, the verification pass standards for size verification, binary image verification, category label verification, bounding box verification, and input image existence verification are that the size of the image and the mask are consistent, the corresponding mask is a binary image, the category label is within the standard range, the bounding box is within the image range, and the input image is not empty.

4. A ship image segmentation method based on spatial pyramid feature enhancement as described in any one of claims 2-3, characterized in that: Data compliance verification is performed by calling OpenCV functions, NumPy functions, and OS functions.

5. The ship image segmentation method based on spatial pyramid feature enhancement according to claim 1, characterized in that: The backbone network of the spatial pyramid deep neural network is the Resnet50 network, and the spatial pyramid feature enhancement network is a multi-layer pyramid structure network; specifically, each layer of the spatial pyramid feature enhancement network consists of a context aggregation block.

6. A ship image segmentation method based on spatial pyramid feature enhancement as claimed in claim 5, characterized in that: The pixel spatial context in the context aggregation block is aggregated in the following manner, namely: Among them, P m and Q m They represent the input feature map and output feature map of the mth layer of the spatial pyramid feature enhancement network, respectively. m represents the reweighting matrix, B m Indicates the number of pixels, W k represents the linear projection matrix used for information comparison, W e Represents the linear projection matrix used for aggregation.

7. A ship image segmentation system based on spatial pyramid feature enhancement, using the ship image segmentation method according to any one of claims 1 to 6, characterized in that: include: The segmentation preparation module is configured to: build the user interface and integrate the user interface with VTK; Create a reader to read the ship image to be segmented; The image segmentation module is configured to: input the ship image to be segmented into the trained spatial pyramid deep neural network for image segmentation, so as to segment the ship from the ship image; The image processing module is configured to: create a mapper and a VTK renderer; convert different data types into image data based on the mapper, and perform rendering operations on the image data based on the VTK renderer; The image display module is configured to: call the renderer and the drawing window in VTK based on the interactor, and display the original image of the ship image to be segmented and the segmented image after the rendering operation on the user interface of the computer platform at the same time.

8. A ship image segmentation system based on spatial pyramid feature enhancement as claimed in claim 7, characterized in that: The constructed user interface is provided with a plurality of functional controls, including importing an image, performing segmentation, and clearing the interface.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the ship image segmentation method based on spatial pyramid feature enhancement as described in any one of claims 1 to 6 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the ship image segmentation method based on spatial pyramid feature enhancement as described in any one of claims 1 to 6 are implemented.