Ship segmentation method and device based on staggered adaptive sensing module remote sensing image
By adopting an interlaced adaptive sensing module method in remote sensing image processing, integrating multi-scale feature information, the problem of difficulty in taking into account local details and global information in the prior art is solved, and the accuracy of ship segmentation is significantly improved.
Patent Information
- Application Number
- CN202510137097.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing remote sensing image ship segmentation method is difficult to take into account local details and global information, resulting in a reduction in segmentation accuracy in complex scenarios.
Using the method based on the interleaved adaptive perception module, the depth separation convolution, point-by-point convolution, multi-scale convolution kernel and branch fusion operations are integrated to integrate local details and global context features.
It significantly improves the accuracy of ship target segmentation, and can capture the boundary details and structural information of ship targets more accurately, especially in complex background scenarios.
Smart Images

Figure CN120070893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image processing, and more specifically, to a method and device for remote sensing image ship segmentation based on an interleaved adaptive perception module. Background Art
[0002] With the rapid development of remote sensing technology, high-resolution remote sensing images have been widely used in fields such as marine monitoring, ship management, and disaster prevention and mitigation. Among them, ship segmentation, as an important task in remote sensing image processing, is of great significance for improving the accuracy of target detection and classification. However, ship targets in remote sensing images usually have characteristics such as small targets, complex backgrounds, and diverse shapes, which pose great challenges to traditional segmentation algorithms. In recent years, the rise of deep learning technology has provided new solutions for the ship segmentation task. By constructing feature extraction, feature aggregation, and segmentation decoding modules, achieving high-precision segmentation of ship targets has become a research hotspot.
[0003] In existing segmentation methods, the encoder-decoder structure based on convolutional neural networks (CNNs) has been widely used. The encoder is responsible for extracting multi-level features and gradually reducing the resolution to capture local and global information of the target; the decoder then restores the resolution through gradual upsampling to generate pixel-level segmentation results. However, in practical applications, due to the usually small and complex shapes of ship targets, single-scale feature extraction is difficult to balance local details and global context information. This limitation makes existing methods vulnerable to background noise interference in complex scenarios, resulting in a decrease in segmentation accuracy. Therefore, how to construct an effective feature aggregation module to balance the extraction and integration of multi-scale information while enhancing the feature expression ability is a key issue in remote sensing image ship segmentation. Summary of the Invention
[0004] In view of this, the present invention provides a method and device for remote sensing image ship segmentation based on an interleaved adaptive perception module, which can solve the limitation problem that it is difficult to balance local details and global information in existing ship segmentation methods; this module effectively integrates feature information of different scales by combining various operation methods such as depthwise separable convolution, pointwise convolution, and multi-scale convolution kernels, and significantly improves the accuracy of ship target segmentation.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An embodiment of the present invention provides a method for remote sensing image ship segmentation based on an interleaved adaptive perception module, including the following steps:
[0007] Step 1: Construct a feature extraction encoder; the encoder extracts multi-level features from the remote sensing image to be segmented through convolution and pooling operations, gradually reducing the resolution and recording the pooling indices to capture the feature information of the ship target;
[0008] Step 2: Construct an interleaved adaptive perception module; input the output feature map of the encoder into the interleaved adaptive perception module, and integrate local details and global context features through depthwise separable convolution, pointwise convolution, multi-scale convolution kernels, and branch fusion operations;
[0009] Step 3: Construct a decoder; use the feature information provided by the interleaved adaptive perception module as the input of the decoder, gradually restore the resolution of the feature map, and accurately reconstruct the shape and boundary of the ship target, and finally generate a pixel-level segmentation result.
[0010] Furthermore, the step S1 includes:
[0011] Step 1.1: Obtain the remote sensing image to be segmented;
[0012] Step 1.2: Apply a convolutional layer to extract the edge and texture features of the ship through a local receptive field, and use the ReLU activation function to enhance the non-linear expression ability;
[0013] Step 1.3: Apply a spiking neuron layer to process the time-dependent information of the impulse activation transmission of the analog neuron;
[0014] Step 1.4: Use a pooling layer for downsampling and record the pooling indices;
[0015] Step 1.5: Repeat the structures of steps 1.2, 1.3, and 1.4. Through multi-level feature extraction, the encoder captures the multi-scale information of the ship target from the remote sensing image;
[0016] Step 1.6: The encoder outputs the final feature map f; the output feature map and pooling indices of the encoder provide rich feature information and spatial location information for the decoder.
[0017] Furthermore, the step S2 includes:
[0018] Step 2.1: Perform batch normalization on the feature map f output by the encoder to obtain f n ; apply depthwise separable convolution to reduce the number of model parameters and computational complexity, and extract feature information to obtain f DW ; perform batch normalization and ReLU activation again to enhance the feature expression ability and obtain f α ;
[0019] Step 2.2: For the feature map f nPerform global average pooling to reduce the dimension and capture the overall information of each channel; apply a spiking neuron layer to simulate neuron spike activities and model temporal relationships; apply an activation function to obtain the weight w α ; Multiply w α by f α , and then add it to f α to obtain the updated feature
[0020] Step 2.3: Apply 1x1 pointwise convolution to perform information integration in the channel dimension; perform batch normalization again to ensure the stability of the data distribution and obtain f β ;
[0021] Step 2.4: Flatten the feature map f of Step 2.1 DW , and perform linear transformation through a fully connected layer, apply an activation function, and apply full-dimensional dynamic convolution to extract a new feature representation. Apply ReLU activation again to obtain w β ; Multiply the feature map f of Step 2.3 β by w β to obtain
[0022] Step 2.5: Apply an activation function to the feature map to enhance the network's non-linear feature learning ability; introduce parallel operations of multi-scale convolutional kernels to aggregate information from different scales and extract local and global multi-scale context information to obtain the feature f M ;
[0023] Step 2.6: Construct multiple feature extraction branches, use the feature f M as the input, and respectively adopt depthwise separable convolution, pointwise convolution, and standard convolution operations to extract multi-dimensional feature information; perform element-wise addition on the feature maps of the three branches to fuse multi-scale features; perform non-linear activation on the fused feature map to obtain the final feature map
[0024] Furthermore, the step S3 includes:
[0025] Step 3.1: The decoder receives the feature map from Step S2 Based on the high-level features extracted by the encoder, the decoder gradually restores the resolution to ensure the accurate reconstruction of the shape and boundary information of the ship target;
[0026] Step 3.2: Use the pooling indices recorded by the encoder for upsampling to expand the resolution of the feature map to 2 times the original;
[0027] Step 3.3: Each upsampling layer is followed by a convolutional layer, and then the ReLU activation function is used.
[0028] Step 3.4: Repeat the structures of Step 3.2 and 3.3. Through multi-level upsampling and convolutional operations, the decoder gradually restores the resolution of the feature map and combines the feature information of the encoder to accurately reconstruct the shape and boundary of the ship target.
[0029] Step 3.5: Use the Softmax classifier to generate pixel-level segmentation results to accurately distinguish the ship target and the background in the remote sensing image.
[0030] Furthermore, the decoder consists of 5 stages. Each stage contains an upsampling layer, a convolutional layer, and an activation function; the resolution of each stage is doubled and the number of channels is halved.
[0031] In a second aspect, an embodiment of the present invention further provides a remote sensing image ship segmentation device based on an interleaved adaptive perception module, including:
[0032] A feature extraction encoder extracts multi-level features from the remote sensing image to be segmented through convolutional and pooling operations, gradually reducing the resolution and recording the pooling index to capture the feature information of the ship target.
[0033] Based on the interleaved adaptive perception module, the output feature map of the encoder is input into the interleaved adaptive perception module, and through depthwise separable convolution, pointwise convolution, multi-scale convolutional kernels, and branch fusion operations, local details and global context features are integrated.
[0034] A decoder uses the feature information provided by the interleaved adaptive perception module as the input of the decoder, gradually restores the resolution of the feature map, and accurately reconstructs the shape and boundary of the ship target, and finally generates pixel-level segmentation results.
[0035] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:
[0036] The remote sensing image ship segmentation method based on the interleaved adaptive perception module not only effectively integrates local and global information in the feature extraction stage, but also reduces the computational resource requirements of the model through parameter-efficient optimization design, improving the adaptability of the segmentation model in various computing environments. Compared with traditional methods, the segmentation technology of the present invention can capture the boundary details and structural information of the ship target more accurately, perform well in complex background scenarios, and has broad practical application prospects. Description of the Drawings
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0038] Figure 1 Flowchart of the remote sensing image ship segmentation method based on the interleaved adaptive perception module provided by the present invention.
[0039] Figure 2 Complete schematic diagram of the remote sensing image ship segmentation method based on the interleaved adaptive perception module provided by the present invention.
[0040] Figure 3 Architecture diagram of the interleaved adaptive perception module provided by the present invention.
[0041] Figure 4 Block diagram of the remote sensing image ship segmentation device based on the interleaved adaptive perception module provided by the present invention. Specific embodiments
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0043] Refer to Figure 1 As shown, the embodiments of the present invention disclose a remote sensing image ship segmentation method based on an interleaved adaptive perception module, including the following steps:
[0044] Step 1, construct a feature extraction encoder; the encoder extracts multi-level features from the remote sensing image to be segmented through convolution and pooling operations, gradually reducing the resolution and recording the pooling index to capture the feature information of the ship target;
[0045] Step 2, construct an interleaved adaptive perception module; input the output feature map of the encoder into the interleaved adaptive perception module, and effectively integrate local details and global context features through depthwise separable convolution, pointwise convolution, multi-scale convolution kernels, and branch fusion operations; thereby significantly enhancing the feature expression ability;
[0046] Step 3: Construct a decoder; use the feature information provided by the interleaved adaptive perception module as the input of the decoder, gradually restore the resolution of the feature map through upsampling and convolution operations, combine the pooling indices and multi-level features of the encoder, accurately reconstruct the shape and boundary of the ship target, and finally generate a pixel-level segmentation result.
[0047] In the embodiment of the present invention, the ship segmentation method for remote sensing images based on the interleaved adaptive perception module not only improves the segmentation accuracy, but also reduces the resource requirements for model training and deployment, making it applicable to remote sensing image processing tasks in various computing environments.
[0048] To further enhance the ship segmentation ability of remote sensing images, in the design of the interleaved adaptive perception module, first, the feature map is efficiently processed through depthwise separable convolution, which reduces the computational complexity while retaining key feature information. Depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution, significantly reducing the number of network parameters. Subsequently, the pointwise convolution operation is used to further enhance the feature expression ability in the channel dimension, effectively improving the network's perception ability of the global context. In addition, to better capture the local details and global features of the target, multi-scale convolutional kernels (such as 3×3, 5×5, etc.) are introduced into the module, and information is extracted from different scales through parallel operations. This design enables the network to simultaneously focus on the local edge details and global structure information of the ship target, significantly improving the feature identification ability.
[0049] To fully integrate multi-scale features, the module also realizes the fusion of multi-dimensional features by constructing multiple feature extraction branches. Each branch uses different convolution operations, such as depthwise separable convolution, pointwise convolution, and standard convolution, and enhances the feature expression ability through normalization layers and activation functions. Finally, the features of all branches are fused by element-wise addition and processed through a non-linear activation function to generate an aggregated feature map with rich information. This modular design not only improves the efficiency of feature extraction but also enhances the segmentation ability for small targets in complex scenes.
[0050] The above steps are described in detail below:
[0051] Refer to Figure 2 As shown, where Step 1: Construct a feature extraction encoder. The role of the encoder is to extract multi-level features from the input image, gradually reduce the resolution through pooling operations, and record the pooling indices for the decoder to restore the spatial information. Specifically, it includes:
[0052] Step 1.1: The input image is a remote sensing image, usually a 3-channel (RGB) or multi-spectral image. The input size is H×W×C, where H is the height, W is the width, and C is the number of channels. Ship targets in remote sensing images are usually small and the background is complex. The encoder extracts features through convolution and can effectively capture local and global information of the ships.
[0053] Step 1.2: First, apply the convolutional layer. Each convolutional layer uses a 3×3 convolutional kernel, with a stride of 1 and padding of 1 (to keep the size of the feature map unchanged). Then, use the ReLU activation function. The convolutional layer extracts features such as the edges and textures of the ships through local receptive fields, and the ReLU activation function enhances the non-linear expression ability.
[0054] Step 1.3: Apply the spiking neuron layer. After the traditional convolutional layer, add the spiking neuron layer to simulate the spiking activity of neurons. The spiking neuron layer can process information in a more biologically realistic way, especially for time-series related feature extraction. It can simulate how neurons respond to input signals and trigger spikes according to time-domain information. Each spiking neuron generates a spike based on the input stimulus (such as the convolutional output) and transmits it to the next layer. The spiking neuron layer can model time-series relationships during the feature extraction process and is suitable for tasks containing time information (such as the dynamic information of ships in remote sensing images). The introduction of the spiking neuron layer can not only enhance the capture of time-series features but also work in cooperation with the traditional convolutional layer to enhance the perception ability of spatial information.
[0055] Step 1.4: Use the pooling layer. Each pooling layer uses a 2×2 pooling window with a stride of 2 for downsampling. Record the pooling indices, that is, the position of the maximum value in each pooling window. The pooling layer reduces the resolution of the feature map, reduces the computational amount, and at the same time enhances the spatial invariance of the features. The pooling indices are used in the decoder to accurately restore the boundary information of the ship target.
[0056] Step 1.5: Repeat the structures of Steps 1.2, 1.3, and 1.4. The encoder consists of 5 stages, and each stage contains: several convolutional layers (usually 2 - 3 layers), a spiking neuron layer, and a max pooling layer. The resolution of each stage is halved and the number of channels is doubled. Through multi-level feature extraction, the encoder can capture multi-scale information of the ship target from the remote sensing image, including local details and global context.
[0057] Step 1.6: Encoder output. The final feature map f output by the encoder has a size of H / 32×W / 32×512. The feature map and pooling indices output by the encoder provide rich feature information and spatial position information for the decoder, which helps to accurately restore the shape and boundary of the ship target in the decoder.
[0058] Ship targets in remote sensing images are usually small and densely distributed. Through multi-level feature extraction, the encoder can effectively distinguish ships from the background. The pooling index is used for upsampling in the decoder, which can accurately restore the boundaries of ship targets and avoid information loss. The multi-level features extracted by the encoder can capture the local details and global context of ship targets, improving the segmentation accuracy. The feature extraction process of the encoder is completely automated and applicable to ship detection and segmentation tasks of large-scale remote sensing images.
[0059] Step 2 described: To improve the segmentation accuracy of ships in complex scenes in remote sensing images, a staggered adaptive perception module is constructed. This module fully extracts and aggregates multi-scale information in the feature map through various methods such as depthwise separable convolution, pointwise convolution, multi-scale convolutional kernels, and branch fusion, effectively capturing the local details and global features of ship targets, as Figure 3 shown. The specific steps are as follows:
[0060] Step 2.1: Perform batch normalization on the feature map f obtained in Step 1 to standardize the data distribution, eliminate the distribution differences between different samples, accelerate network convergence, and obtain f n . Then, through depthwise separable convolution operations, reduce the number of model parameters and computational complexity, and extract richer feature information to obtain f DW . Next, perform batch normalization on the feature map after depthwise separable convolution to further standardize the data distribution. Finally, use a non-linear activation function (such as ReLU) to increase the non-linear expression ability of the model and obtain f α .
[0061] Step 2.2: For f in Step 2.1 n , first perform global average pooling, which reduces the dimension of each channel of the input feature map by calculating the global mean. It generates a single average value for each channel of the feature map, thus compressing the spatial dimension to the scale of 1x1. This usually reduces the number of parameters and helps the network focus on global features. After global average pooling, the feature map becomes a single scalar value for each channel, which can be understood as the importance weight for each channel. It captures the overall information of each channel in the feature map but no longer focuses on the local spatial structure. Next, apply a spiking neuron layer. The output weights usually contain parameters related to the sensitivity or spike response of each neuron to the input. These weights control the firing frequency, time window, and spike triggering threshold of the neurons, affecting the final output pattern. In such a layer, the weights are usually related to the response pattern of the neurons, time delay, and the timing of the input signal. Finally, apply an activation function to obtain the weight w α :
[0062] w α = Sigmoid(SNL(GAP(fn )))
[0063] Among them, Sigmoid represents the sigmoid activation function, SNL represents the spiking neuron layer, and GAP represents the global average pooling layer. Next, multiply w α by f α , and then add it to f α to improve the network's attention to the target features and obtain the updated features
[0064]
[0065] Among them, represents matrix addition, represents matrix multiplication.
[0066] Step 2.3: Further enhance the information expression ability of the channel dimension and improve the network's perception ability in the global context. Through 1x1 pointwise convolution operation, perform information integration on the feature map in the channel dimension, further reduce the computational complexity, and enhance the feature expression. Then, perform batch normalization again to ensure the stability of the data distribution and obtain f β .
[0067] Step 2.4: Flatten the f DW feature map in Step 2.1 (usually flatten the spatial dimension into a one-dimensional vector), and then perform a linear transformation through the weight matrix and the bias term. The output of this layer is a new feature representation, which is usually used for classification, regression, or as the input of other layers. In the fully connected layer, the weight matrix is the key parameter of this layer. For each channel of the input feature map, there is a weight between each input node and the output node. The weights of the fully connected layer are a matrix, and its size is related to the size of the input feature map and the number of output nodes. This weight determines how the feature map is mapped to the new space. Then, apply the activation function. Next, apply the full-dimensional dynamic convolution, which is a convolution method that introduces a dynamic mechanism based on the traditional convolution operation. The filters of the traditional convolution are static, while in the dynamic convolution, the filters will change according to the input data or other conditions. For example, the dynamic convolution can adjust the parameters of the convolution kernel according to the input features, so that the convolution kernel has different forms and weights on different inputs. The weights of the dynamic convolution are dynamically adjusted and not fixed. During the training process, the weights of the convolution kernel will be updated as the input features change, so that the convolution kernel can adaptively adjust according to different inputs. This means that the weights of the dynamic convolution include not only the parameters of each convolution kernel (filter weights), but also the mechanism for dynamically generating or adjusting these parameters according to the input features. Finally, apply the activation function again to obtain w β :
[0068] w β = Sigmoid(OConv(ReLU(FC(f DW ))))
[0069] where FC represents the fully connected layer, OConv represents the full-dimensional dynamic convolution layer, and ReLU represents the ReLU activation function. Next, w β is multiplied by f obtained in step 2.3 β to obtain which further adaptively enhances the network's attention to the target.
[0070] Step 2.5: Activate the feature map to enhance the network's non-linear feature learning ability. To introduce parallel operations of multi-scale convolution kernels (such as 3x3, 5x5, etc.) and aggregate information from different scales, local and global multi-scale context information is extracted from the spatial dimension to obtain the feature f M .
[0071] Step 2.6: Combine different methods such as depthwise separable convolution, multi-scale convolution, and pointwise convolution to construct multiple feature extraction branches and aggregate multi-dimensional feature information. First, construct Branch 1, which passes through the normalization layer to obtain Next, construct Branch 2, which extracts features through depthwise separable convolution and is followed by a normalization layer:
[0072]
[0073] where DWConv represents depthwise separable convolution and BN represents the normalization layer. Next, construct Branch 3, which extracts features through pointwise convolution and is followed by a normalization layer:
[0074]
[0075] where PWConv represents the pointwise convolution operation. Finally, the feature maps of the three branches are added element-wise to fuse multi-scale features, and the fused feature map is non-linearly activated to enhance the network's expression ability. The final feature map
[0076]
[0077] where GeLU represents the GeLU activation function, represents element-wise addition.
[0078] Step 3 described: Construct the feature decoder and output the segmentation result. The role of the decoder is to gradually upsample the feature map extracted by the encoder, restore the resolution, and precisely restore the boundary information of the ship target in combination with the pooling indices recorded by the encoder, and finally generate a pixel-level segmentation result, as Figure 2 shown.
[0079] Step 3.1: Input the feature map obtained in Step 2 with a size of H / 32×W / 32×512. Based on the high-level features extracted by the encoder, the decoder gradually restores the resolution to ensure that the shape and boundary information of the ship target are precisely reconstructed.
[0080] Step 3.2: Apply the upsampling layer. Use Pooling Indices for upsampling. After upsampling, the resolution of the feature map is doubled. Upsampling through the pooling indices can precisely restore the boundary information of the ship target and avoid information loss.
[0081] Step 3.3: Apply the convolutional layer. Each upsampling layer is followed by a convolutional layer with a 3×3 convolutional kernel, a stride of 1, and a padding of 1. Then use the ReLU activation function. The convolutional layer further refines the upsampled feature map, enhances the feature expression ability, and ensures that the detailed information of the ship target is retained.
[0082] Step 3.4: Repeat the structure of Step 3.2 and Step 3.3. The decoder consists of 5 stages, and each stage contains: an upsampling layer, a convolutional layer, and an activation function. The resolution of each stage is doubled, and the number of channels is halved. Through multi-level upsampling and convolutional operations, the decoder gradually restores the resolution of the feature map and combines the feature information of the encoder to precisely reconstruct the shape and boundary of the ship target.
[0083] Step 3.5: Use the output layer. The last layer is a Softmax classifier for classifying each pixel. The output is the segmentation result with a size of H×W×K, where K is the number of classes (such as ships and background). The Softmax classifier generates a pixel-level segmentation result that can precisely distinguish ship targets and backgrounds in remote sensing images. The decoder upsamples through the pooling indices, can precisely restore the boundary information of the ship target, and avoid information loss. The decoder combines multi-level features of the encoder, can capture local details and global context of the ship target, and improve the segmentation accuracy. The upsampling and convolutional operations of the decoder are fully automated and applicable to ship detection and segmentation tasks of large-scale remote sensing images.
[0084] The present invention provides a method for segmenting ships in remote sensing images based on an interleaved adaptive perception module. Through the interleaved adaptive perception module, multi-scale feature information in remote sensing images is effectively integrated, significantly improving the segmentation accuracy of ship targets of different sizes, especially showing outstanding performance in complex backgrounds and small target detection. In addition, lightweight technologies such as depthwise separable convolution and pointwise convolution are adopted to greatly reduce the number of model parameters and computational complexity, reduce the demand for computing resources and inference time, and are more suitable for large-scale remote sensing data processing scenarios.
[0085] Based on the same inventive concept, an embodiment of the present invention further provides a device for segmenting ships in remote sensing images based on an interleaved adaptive perception module. Referring to Figure 4 as shown, it includes:
[0086] A feature extraction encoder extracts multi-level features from the remote sensing image to be segmented through convolution and pooling operations, gradually reducing the resolution and recording the pooling index to capture the feature information of ship targets;
[0087] Based on the interleaved adaptive perception module, the output feature map of the encoder is input into the interleaved adaptive perception module, and local details and global context features are integrated through depthwise separable convolution, pointwise convolution, multi-scale convolution kernels, and branch fusion operations;
[0088] A decoder uses the feature information provided by the interleaved adaptive perception module as the input of the decoder, gradually restores the resolution of the feature map, and precisely reconstructs the shape and boundary of the ship target, and finally generates a pixel-level segmentation result.
[0089] Hypothetical scenario: We need to segment ship targets from high-resolution remote sensing images for ship management and ocean monitoring.
[0090] 1. Feature extraction encoder:
[0091] Input: A remote sensing image containing ships and backgrounds.
[0092] Processing process:
[0093] Use convolutional layers to extract features such as edges and textures of the image.
[0094] Use a spiking neuron layer to simulate neuron spike activities, capture temporal information, and enhance the perception of spatial information.
[0095] Use pooling layers to reduce the resolution of the feature map, enhance the spatial invariance of features, and record the pooling index.
[0096] Repeat the above steps to construct a multi-stage feature extraction network, gradually reducing the resolution and extracting multi-level features.
[0097] Output: A feature map containing multi-scale information of ship targets and pooling indices.
[0098] 2. Based on the interleaved adaptive perception module:
[0099] Input: The output feature map of the feature extraction encoder.
[0100] Processing procedure:
[0101] Use depthwise separable convolution to reduce the number of model parameters and computational complexity, and extract rich feature information.
[0102] Use pointwise convolution to enhance the feature expression ability in the channel dimension and improve the network's perception ability of the global context.
[0103] Use global average pooling, spiking neuron layer and activation function to extract global features and generate weights.
[0104] Use full-dimensional dynamic convolution to further adaptively adjust the feature weights and improve the network's attention to the target.
[0105] Use multi-scale convolutional kernels for parallel operations, aggregate information from different scales, and extract local and global multi-scale context information.
[0106] Construct multiple feature extraction branches, respectively use depthwise separable convolution, pointwise convolution and standard convolution to extract features, and perform fusion to obtain an aggregated feature map with rich information.
[0107] Output: An aggregated feature map containing local details and global features of ship targets.
[0108] 3. Feature decoder, output the segmentation result:
[0109] Input: The aggregated feature map.
[0110] Processing procedure:
[0111] Use the pooling index for upsampling to restore the resolution of the feature map and accurately restore the boundary information of the ship target.
[0112] Use the convolutional layer to refine the upsampled feature map and enhance the feature expression ability.
[0113] Repeat the above steps to gradually restore the resolution of the feature map, and combine the encoder feature information to accurately reconstruct the shape and boundary of the ship target.
[0114] Output: Pixel-level segmentation result, distinguishing ship targets from the background.
[0115] In this embodiment, the staggered adaptive perception module effectively integrates multi-scale feature information, enabling accurate segmentation of ship targets of different sizes, especially showing outstanding performance in complex backgrounds and small target detection. By adopting lightweight technologies such as depthwise separable convolution and pointwise convolution, the number of model parameters and computational complexity are reduced, the computational resource requirements and inference time are decreased, making it more suitable for large-scale remote sensing data processing scenarios.
[0116] It can be applied to marine monitoring: monitoring the number, type, and activity trajectories of ships. Ship management: identifying and tracking ships for safety management and scheduling. Disaster prevention and mitigation: identifying damaged ships for rescue and emergency handling.
[0117] A remote sensing image ship segmentation device based on a staggered adaptive perception module provided by the present invention can effectively segment ship targets, and has advantages such as high precision and lightweight, and has broad application prospects.
[0118] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the method part for the relevant parts.
[0119] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A ship segmentation method for remote sensing images based on interlaced adaptive perception modules, characterized in that: The following steps are involved: Step 1: construct a feature extraction encoder; The encoder extracts multi-level features from the remote sensing image to be segmented through convolution and pooling operations, gradually reduces the resolution and records the pooling index to capture the characteristics of the ship target information; Step 2: construct an interlaced adaptive perception module; input the output feature map of the encoder into the interlaced adaptive perception module, and integrate local details and global context features through depth-separable convolution, point-by-point convolution, multi-scale convolution kernel and branch fusion operations; Step 3, construct a decoder; use the feature information provided by the interlaced adaptive perception module as the input of the decoder, gradually restore the resolution of the feature map, and accurately reconstruct the shape and boundary of the ship target, and finally generate a pixel-level segmentation result.
2. According to claim 1, a remote sensing image ship segmentation method based on staggered adaptive perception module is characterized in that: The step S1 comprises: Step 1.1: Obtain the remote sensing image to be segmented; Step 1.2: Apply the convolutional layer to extract the edge and texture features of the ship through the local receptive field, and use the ReLU activation function to enhance the nonlinear expression ability; Step 1.3: Apply the spiking neuron layer to simulate the time-dependent information transmitted by the spiking activation of neurons; Step 1.4: Use the pooling layer to downsample and record the pooling index; Step 1.5: Repeat the structures of steps 1.2, 1.3 and 1.4, and through multi-level feature extraction, the encoder captures multi-scale information of the ship target from the remote sensing image; Step 1.6: The encoder outputs the final feature map f; the encoder output feature map and pooling index provide the decoder with rich feature information and spatial position information.
3. The method for ship segmentation in remote sensing images based on interlaced adaptive perception modules according to claim 1 is characterized in that: The step S2 comprises: Step 2.1: Batch normalize the feature map f output by the encoder to obtain f n ; Apply depthwise separable convolution to reduce the number of model parameters and calculations, extract feature information to obtain f DW ; Perform batch normalization and ReLU activation again to enhance the feature expression ability and obtain f α ; Step 2.2: For the feature map f n Perform global average pooling to reduce the dimension and capture the overall information of each channel; apply the spike neuron layer to simulate the neuron spike activity and model the timing relationship; apply the activation function to get the weight w α ; α With f α Multiply them and add them to f α Add together to get the updated features Step 2.3: Apply 1x1 point-by-point convolution to integrate information in the channel dimension; perform batch normalization again to ensure the stability of data distribution, and get f β ; Step 2.4: The feature map f in step 2.1 DW Flatten, pass through the fully connected layer, apply the activation function, and apply full-dimensional dynamic convolution for linear transformation to extract new feature representations, and apply ReLU activation again to get w β ; The feature map f in step 2.3 β With w β Multiply them together and you get Step 2.5: Feature map Apply activation functions to improve the network's nonlinear feature learning ability; introduce multi-scale convolution kernels for parallel operation, aggregate information from different scales, extract local and global multi-scale context information, and obtain feature f M ; Step 2.6: Construct multiple feature extraction branches and extract the feature f M As input, depthwise separable convolution, pointwise convolution and standard convolution operations are used to extract multi-dimensional feature information; the feature maps of the three branches are added element by element to fuse multi-scale features; the fused feature maps are nonlinearly activated to obtain the final feature map 4. The method for ship segmentation in remote sensing images based on interlaced adaptive perception modules according to claim 3 is characterized in that: The step S3 comprises: Step 3.1: The decoder receives the feature map from step S2 The decoder gradually restores the resolution based on the high-level features extracted by the encoder, ensuring that the shape and boundary information of the ship target can be accurately reconstructed; Step 3.2: Use the pooling index recorded by the encoder to upsample and increase the resolution of the feature map to twice the original resolution; Step 3.3: Each upsampling layer is followed by a convolutional layer, and then a ReLU activation function is used; Step 3.4: Repeat the structure of steps 3.2 and 3.
3. Through multi-level upsampling and convolution operations, the decoder gradually restores the resolution of the feature map and combines the feature information of the encoder to accurately reconstruct the shape and boundary of the ship target. Step 3.5: Use the Softmax classifier to generate pixel-level segmentation results to accurately distinguish between ship targets and background in remote sensing images.
5. The method for ship segmentation in remote sensing images based on interlaced adaptive perception modules according to claim 1, characterized in that: The decoder consists of 5 stages, each of which includes an upsampling layer, a convolution layer and an activation function; the resolution of each stage is doubled and the number of channels is halved.
6. A remote sensing image ship segmentation device based on interlaced adaptive perception module, characterized in that: include: The feature extraction encoder extracts multi-level features from the remote sensing image to be segmented through convolution and pooling operations, gradually reduces the resolution and records the pooling index to capture the characteristics of the ship target information; Based on the interlaced adaptive perception module, the output feature map of the encoder is input into the interlaced adaptive perception module, and local details and global context features are integrated through depthwise separable convolution, point-by-point convolution, multi-scale convolution kernel and branch fusion operations; The decoder,uses the feature information provided by the interlaced adaptive perception module as,the input of the decoder, gradually restores the resolution of the feature map,and accurately reconstructs the shape and boundary of the ship target,and finally generates a pixel-level segmentation result.
Citation Information
Patent Citations
Burn area segmentation system based on pulse neural network U-shaped model
CN114187306A
Three-dimensional image automatic segmentation method, system, equipment and medium
CN115880312A
Three-dimensional indoor scene semantic segmentation method based on residual pulse neural network
CN116958557A
Semantic segmentation method and device for ocean remote sensing image and electronic equipment
CN117522884A
Remote sensing image segmentation method and system based on adaptive enhancement and fine granularity guidance
CN118096784A
Cited By
Remote sensing image segmentation method and system fusing stage perception and multi-dimensional orientation mechanism
CN121053541A