Remote sensing image target segmentation method based on multistage feature aggregation and segmentation

By adopting multi-level feature aggregation and segmentation methods in remote sensing image processing, a ship segmentation model containing multiple modules is constructed, which solves the detection problems of ship targets in the existing technology under multi-scale and complex backgrounds, and achieves efficient and accurate ship target segmentation and real-time detection.

CN119919658AActive Publication Date: 2025-05-02耕宇牧星(北京)空间科技有限公司

Patent Information

Application Number
CN202411978616.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The prior art has shortcomings in multi-scale processing, complex background adaptation and computing efficiency of ship targets in remote sensing images, especially in resource-constrained environments, where real-time or near-real-time ship detection is difficult to achieve.

Method used

Using a remote sensing image target segmentation method based on multi-level feature aggregation and segmentation, a ship segmentation model including feature extraction module, aggregation-segment module, attention fusion module, reparameterized convolution module and output module is used to achieve efficient and accurate segmentation of ship targets in remote sensing images.

Benefits of technology

The detection and segmentation accuracy of ship targets in remote sensing images is improved, the calculation complexity of the model is reduced, and it is suitable for various computing environments, real-time or near-real-time ship detection is realized, and the efficiency of marine monitoring is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919658A_ABST
    Figure CN119919658A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image target segmentation method based on multistage feature aggregation and segmentation, and relates to the technical field of remote sensing image processing, and the method comprises the steps: 1, collecting a remote sensing ship image, and constructing a training data set; 2, constructing a remote sensing image vessel segmentation model, and training the remote sensing image vessel segmentation model by adopting the training data set; the remote sensing image vessel segmentation model comprises a feature extraction module, an aggregation-segmentation module, an attention fusion module, a re-parameterization convolution module and an output module which are connected in sequence; and step 3, collecting a remote sensing image to be segmented, and inputting the remote sensing image to the trained remote sensing image vessel segmentation model to obtain a vessel target segmentation result. The method can improve the detection and segmentation precision of the vessel target in the remote sensing image, and guarantees the calculation efficiency and practicality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and more particularly to a remote sensing image target segmentation method based on multi-level feature aggregation and segmentation. Background Art

[0002] In the field of remote sensing image processing, ship target segmentation technology is the key to realizing applications such as ocean monitoring, navigation safety and military reconnaissance. This technology can identify and locate ship targets from a large amount of remote sensing data, providing important information for decision-making. Traditional ship segmentation methods mainly rely on image processing techniques such as edge detection, region growing and threshold segmentation, which often require a lot of manual intervention and have limited effect on ship detection under complex backgrounds and variable lighting conditions. For example, edge detection may fail when the image is noisy, and the region growing method may be affected by the selection of the initial seed point, resulting in unstable segmentation results.

[0003] With the development of deep learning technology, especially the breakthrough of convolutional neural network (CNN) in the field of image recognition and segmentation, deep learning-based methods have gradually become mainstream due to their powerful feature extraction and learning capabilities. These methods reduce the dependence on manual features and improve the accuracy of segmentation by automatically learning image features. However, existing deep learning-based ship segmentation methods still have shortcomings in dealing with multi-scale ship targets, complex ocean backgrounds, and real-time requirements. These methods usually require a lot of computing resources, which limits their widespread deployment in practical applications. For example, some high-performance deep learning models have high accuracy but high computational complexity, which is not suitable for resource-constrained embedded systems or mobile devices.

[0004] In addition, the adaptability and generalization ability of existing methods need to be improved for remote sensing images of different resolutions and complexities. The performance of existing models may be significantly reduced when faced with different sea conditions, different lighting conditions, and different ship types.

[0005] Therefore, how to provide a remote sensing image ship segmentation method that can efficiently process multi-scale targets, adapt to complex backgrounds and has high computational efficiency is a problem that technicians in this field urgently need to solve. Summary of the invention

[0006] In view of this, the present invention provides a remote sensing image target segmentation method based on multi-level feature aggregation and segmentation, which can improve the detection and segmentation accuracy of ship targets in remote sensing images while ensuring the computational efficiency and practicability of the model.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] A remote sensing image target segmentation method based on multi-level feature aggregation and segmentation comprises the following steps:

[0009] Step 1: Collect remote sensing ship images and build a training data set;

[0010] Step 2: Construct a remote sensing image ship segmentation model and use the training data set to train the remote sensing image ship segmentation model; the remote sensing image ship segmentation model includes a feature extraction module, an aggregation-segmentation module, an attention fusion module, a re-parameterized convolution module and an output module connected in sequence;

[0011] Step 3: Collect the remote sensing image to be segmented and input it into the trained remote sensing image ship segmentation model to obtain the ship target segmentation result.

[0012] Preferably, the feature extraction module includes multiple groups of encoders, each group of encoders includes sequentially connected convolution layers, nonlinear activation layers and pooling layers, each convolution layer performs convolution operations to extract spatial features in remote sensing ship images to obtain feature maps, and the number of channels of the feature maps increases after each convolution operation to extract more high-level features, and the nonlinear activation layer is applied to perform nonlinear transformation on the feature maps after the convolution operation, and then downsampled through the pooling operation of the pooling layer to reduce the spatial resolution of the feature maps. The input image passes through the convolution layer, the nonlinear activation layer and the pooling layer, and the high-level semantic features are gradually extracted through the step-by-step downsampling process; the nonlinear activation layer uses the ReLU activation function; and the pooling layer performs the maximum pooling operation.

[0013] Preferably, the convolution operation of the convolutional layer in the encoder is expressed as:

[0014] F i =Conv(X i-1 )

[0015] Among them, X i-1 represents the input feature map of the i-1th layer in the remote sensing ship image, Conv represents the 2D convolution operation, and F i Represents the feature map obtained after convolution, and its number of channels is generally more than the number of input channels; the remote sensing ship image is represented as Where H0, W0, and C0 represent the height, width, and number of channels of the image, respectively;

[0016] The nonlinear activation layer is used to introduce nonlinear transformation to the feature map, allowing the model to learn more complex features. The expression is:

[0017] F′ i =ReLU(F i )

[0018] Among them, F′ iRepresents the feature map after activation by the ReLU activation function; ReLU represents the ReLU activation function;

[0019] The expression of the pooling operation of the pooling layer is:

[0020] P i =MaxPool(F′ i )

[0021] Among them, P i Represents the feature map after pooling; MaxPool represents the maximum pooling operation.

[0022] Preferably, the aggregation-segmentation module includes a feature alignment layer, a fused convolution layer and a segmentation layer connected in sequence; the attention fusion module includes a weight fusion layer and a sum fusion layer using an attention mechanism; the feature map extracted by the feature extraction module is downsampled in the feature alignment layer through average pooling and bilinear interpolation operations to obtain an aligned feature map; the fused convolution layer performs a convolution operation on the aligned feature map to achieve reparameterization to obtain a fused feature; the segmentation layer segments the fused feature in the channel dimension to obtain a number of local feature maps; the weight fusion layer uses an attention mechanism to perform weight fusion on the local feature map to obtain an enhanced feature map; the sum fusion layer performs a sum operation on the local feature map and the enhanced feature map to obtain a comprehensive feature map; the reparameterized convolution module performs a convolution operation on the comprehensive feature map to obtain a fused enhanced feature. The aggregation-segmentation module is used to fuse feature maps to obtain high-resolution features that retain minimum target information.

[0023] Preferably, the feature map in the aggregation-segmentation module is downsampled in the feature alignment layer through average pooling and bilinear interpolation operations to obtain an aligned feature map, which is expressed as:

[0024]

[0025] Among them, AvgPool represents the average pooling operation; Bilinear represents the bilinear interpolation operation; represents the splicing and fusion operation; P1, P2, P3, and P4 respectively represent the feature maps output by a set of encoders of the feature extraction module;

[0026] The fused convolution layer includes the first convolution layer, the re-parameterized convolution module and the second convolution layer connected in sequence; the re-parameterized convolution module includes a 1*1 convolution layer, a fusion layer and a 3*3 convolution layer; the alignment feature map G align Input to the first convolutional layer for convolution operation to obtain the feature map F′ align ; Re-parameterized convolution module for feature map F′ align After performing 1*1 convolution, fusion and 3*3 convolution in sequence, the fusion feature F is obtained. fuse ; The second convolutional layer combines the features Ffuse Convolution is performed to obtain convolution fusion features, convolution fusion features and fusion features F fuse The number of channels is the same as that of the convolution module; the expression of the reparameterized convolution module is:

[0027]

[0028] Among them, F′ align It represents the feature map obtained after the convolution operation of the convolution layer; F fuse Indicates fusion features; Conv3 indicates a convolution operation with a convolution kernel of 3; Conv1 indicates a convolution operation with a convolution kernel of 1;

[0029] The segmentation layer calculates the number of channels for the convolutional fusion features and performs segmentation based on the number of channels to obtain the first feature map F. α and the second feature map F β .

[0030] Preferably, the process of feature processing by the attention fusion module is:

[0031] The weight fusion layer performs convolution and average pooling operations on the first feature map output by the segmentation layer to obtain the first local feature map, which is expressed as:

[0032] F α+ =AvgPool(Conv1(F α ))

[0033] Among them, F α+ represents the first local feature map; F α Represents the first feature map of the segmentation layer output; AvgPool represents the average pooling operation; Conv1 represents the convolution operation with a convolution kernel of 1;

[0034] The weight fusion layer performs convolution operation, activation function activation and average pooling operation on the first feature map output by the segmentation layer in sequence to obtain the attention weight, which is expressed as:

[0035] A α =AvgPool(σ(Conv1(F α )))

[0036] Among them, A α represents the attention weight; σ represents the Sigmoid activation function;

[0037] The weight fusion layer performs convolution and average pooling operations on the second feature map output by the segmentation layer in sequence to obtain the second local feature map, which is expressed as:

[0038] F β+ =AvgPool(Conv1(F β )) Among them, Fβ Represents the second feature map of the segmentation layer output; F β+ represents the second local feature map;

[0039] The weight fusion layer uses the attention mechanism to transform the second local feature map F β+ and attention weight A α Fusion is performed to obtain an enhanced feature map, the expression is:

[0040]

[0041] Among them, F αβ+ represents the enhanced feature map; Represents a pixel-by-pixel multiplication operation;

[0042] The addition and fusion layer adds the first local feature map and the enhanced feature map to obtain a comprehensive feature map, which is expressed as:

[0043]

[0044] in, Represents a comprehensive feature map.

[0045] Preferably, the re-parameterized convolution module is used to integrate the feature map After performing 1*1 convolution, fusion and 3*3 convolution in sequence, the fusion enhancement feature is obtained. The re-parameterized convolution module includes a 1*1 convolution layer, a fusion layer, and a 3*3 convolution layer.

[0046] Preferably, the output module includes an upsampling layer and a loss calculation layer; the upsampling layer includes a 1*1 convolution layer, a transposed convolution layer and an activation function layer connected in sequence, and the fusion enhancement feature After the convolution operation of the 1*1 convolution layer, the input is sent to the transposed convolution layer. The transposed convolution layer upsamples the fused enhanced features after convolution to the same resolution as the remote sensing ship image. The activation function layer converts the upsampled feature map into a probability distribution to obtain the ship target segmentation result. The loss calculation layer uses a mixed loss function to calculate the loss value based on the ship target segmentation result output by the activation function layer and the ship target segmentation result marked in the remote sensing ship image to optimize the model training and improve the accuracy and robustness of the model segmentation. The activation function layer uses the Softmax activation function; the mixed loss function used in the loss calculation layer includes the cross entropy loss function and the Dice loss function.

[0047] Preferably, the mixed loss function is expressed as:

[0048] L=ηL c +γL δ

[0049] Among them, η and γ represent the hyperparameters used to balance the weights of different loss functions; L c represents the cross entropy loss; L δ represents Dice loss;

[0050] Cross entropy loss L c It is expressed as:

[0051]

[0052] Among them, i represents the category index, Y i represents the segmentation result of the real ship target annotated in the image, P i It represents the ship target segmentation result predicted and output by the remote sensing image ship segmentation model;

[0053] Dice loss L δ It is expressed as:

[0054]

[0055] Among them, λ represents a smoothing term to prevent the denominator from being zero; P represents the distribution of ship target segmentation results predicted and output by the remote sensing image ship segmentation model; Y represents the distribution of the actual ship target segmentation results marked in the image.

[0056] Preferably, the Adam optimizer is used in the remote sensing image ship segmentation model training process to minimize the mixed loss function and optimize the model.

[0057] Preferably, step 1 also includes performing data enhancement processing on the collected remote sensing ship images, expanding the images through random rotation, flipping and scaling of the images, and constructing a training data set.

[0058] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a remote sensing image target segmentation method based on multi-level feature aggregation and segmentation, which performs multi-level feature aggregation and re-segmentation through an aggregation-segmentation model, and combines the attention mechanism to achieve efficient and accurate segmentation of ship targets in remote sensing images. The aggregation-segmentation model effectively captures and aggregates feature information at different levels, enhances the model's detection ability for small targets and complex backgrounds, and retains high-resolution feature details through re-segmentation, thereby improving the accuracy of segmentation. In addition, the use of the attention mechanism reduces unnecessary computational burdens, enables the model to automatically focus on key areas in the image, and further improves the ability to identify ship targets in complex backgrounds. The model has a fast processing speed while maintaining high accuracy, and is suitable for real-time or near real-time remote sensing image analysis needs. In addition, in terms of enhancing generalization ability, adaptability and reducing resource consumption, the present invention introduces a combination of data enhancement and multiple loss functions, which improves the robustness of the model under different remote sensing image conditions, enabling the model to work stably under changing environments and conditions, and reduces dependence on high-performance computing resources by optimizing the model structure and calculation process, making the model easier to deploy on various computing platforms, including edge computing devices, thereby expanding the application scope of the model.

[0059] The present invention can significantly reduce the computational complexity of the model while maintaining high accuracy, making it suitable for remote sensing image processing tasks in various computing environments. It is of great significance for improving the practicality and effectiveness of remote sensing image processing. It can not only improve the efficiency of ocean monitoring, but also realize real-time or near real-time ship detection in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0061] Figure 1 A schematic diagram of the structure of a remote sensing image ship segmentation model provided by the present invention;

[0062] Figure 2 This is a schematic diagram of the aggregation-segmentation module structure provided by the present invention. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] The embodiment of the present invention discloses a remote sensing image target segmentation method based on multi-level feature aggregation and segmentation, which is characterized by comprising the following steps:

[0065] Step 1: Collect remote sensing ship images and build a training data set;

[0066] Step 2: Construct a remote sensing image ship segmentation model and use the training data set to train the remote sensing image ship segmentation model; the remote sensing image ship segmentation model includes a feature extraction module, an aggregation-segmentation module, an attention fusion module, a re-parameterized convolution module and an output module connected in sequence;

[0067] Step 3: Collect the remote sensing image to be segmented and input it into the trained remote sensing image ship segmentation model to obtain the ship target segmentation result.

[0068] The remote sensing image of the present invention not only improves the accuracy of segmentation based on multi-level feature aggregation and re-segmentation, but also reduces the resource requirements for model training and deployment through efficient parameter fine-tuning and the use of adapters, making it suitable for remote sensing image processing tasks in various computing environments.

[0069] Furthermore, the feature extraction module includes multiple groups of encoders, each group of encoders includes a convolution layer, a nonlinear activation layer and a pooling layer connected in sequence. The convolution operation of each convolution layer extracts the spatial features in the remote sensing ship image to obtain a feature map. After the convolution operation, the nonlinear activation layer is applied to perform a nonlinear transformation on the feature map, and then down-sampled through the pooling operation of the pooling layer to reduce the spatial resolution of the feature map.

[0070] Furthermore, the aggregation-segmentation module includes a feature alignment layer, a fusion convolution layer and a segmentation layer connected in sequence; the feature map extracted by the feature extraction module is down-sampled in the feature alignment layer in sequence through an average pooling operation and a bilinear interpolation operation to obtain an aligned feature map; the fusion convolution layer of the aggregation-segmentation module performs a convolution operation on the aligned feature map to obtain a fusion feature; the segmentation layer segments the fusion feature in the channel dimension to obtain a number of local feature maps;

[0071] Furthermore, the attention fusion module includes a weight fusion layer and a sum fusion layer that adopt an attention mechanism; the weight fusion layer adopts the attention mechanism to perform weight fusion on the local feature map to obtain an enhanced feature map; the sum fusion layer performs a sum operation on the local feature map and the enhanced feature map to obtain a comprehensive feature map.

[0072] Furthermore, the re-parameterized convolution module performs convolution operations on the comprehensive feature map to obtain fused enhanced features.

[0073] Furthermore, the output module includes an upsampling layer and a loss calculation layer; the upsampling layer includes a 1*1 convolution layer, a transposed convolution layer, and an activation function layer connected in sequence, and the fusion enhancement feature After the convolution operation of the 1*1 convolutional layer, the image is input into the transposed convolutional layer. The transposed convolutional layer upsamples the fused enhanced features after convolution to the same resolution as the remote sensing ship image. The activation function layer converts the upsampled feature map into a probability distribution to obtain the ship target segmentation result. The loss calculation layer uses a hybrid loss function to calculate the loss value according to the ship target segmentation result output by the activation function layer and the ship target segmentation result marked in the remote sensing ship image to optimize the model training and improve the accuracy and robustness of the model segmentation.

[0074] In a specific embodiment, the process of processing the image by the remote sensing image ship segmentation model is as follows:

[0075] S1: Input and feature extraction of remote sensing ship images;

[0076] The input remote sensing ship image is Among them, H0, W0, and C0 are the height, width, and number of channels of the input image respectively; the purpose of feature extraction is to extract high-level semantic features of the image through a step-by-step downsampling process; the remote sensing image ship segmentation model includes 4 sets of encoders. In each layer of the encoder, the image undergoes a series of convolution and pooling operations, the spatial resolution gradually decreases, and the number of channels increases, thereby capturing richer context information, and finally obtaining 4 sets of feature maps; the specific process is as follows:

[0077] S11: In each layer of the encoder, the input image is processed by a convolution operation. The purpose of the convolution is to extract the spatial features in the feature map. The convolution operation for each layer can be expressed as:

[0078] F i =COnv(X i-1 )

[0079] Among them, X i-1 is the input feature map of the i-1th layer, Conv is the convolution operation (2D convolution), F iIt is the feature map obtained after convolution; the number of channels of the feature map will be increased after each convolution layer. This is to extract more advanced features and help the network capture more complex semantic information. The increment of the number of channels is usually set according to the design requirements, that is, Among them, H i and W i are the height and width of the feature map, C i is the number of channels of the feature map, which is usually more than the number of channels in the previous layer;

[0080] S12: After convolution, a nonlinear activation function is usually applied to the feature map to introduce nonlinear transformation. This is to enable the network to learn more complex features. The nonlinear activation function is expressed as:

[0081] F′ i =ReLU(F i )

[0082] Among them, F′ i It is the feature map after ReLU activation; the ReLU activation function is usually applied to the output of the convolutional layer, and the calculation formula is:

[0083] ReLU(x)=max(0,x);

[0084] S13: Perform pooling operations, especially maximum pooling, which is usually used for downsampling. In the encoder part, the feature map of each layer will reduce the spatial resolution through pooling operations to further compress the information. The role of pooling is to reduce the amount of calculation while maintaining important spatial features. The formula for the pooling operation can be expressed as:

[0085] P i =MaxPool(F′ i )

[0086] Among them, P i It is the feature map after pooling. MaxPool is the maximum pooling operation. The pooling window is 2×2 and the step size is 2. The pooling operation reduces the width and height of the feature map, but retains the most significant features of each local area.

[0087] S2: Output from S1 The features are input to the aggregation-segmentation module for fusion to obtain high-resolution features that retain small target information. The structure of the aggregation-segmentation module is as follows: Figure 2 As shown, the specific processing process is:

[0088] S21: First, the input feature map is downsampled by average pooling and bilinear interpolation to ensure the uniformity of the feature map size; the aligned feature map F is obtained by adjusting the feature map to the smallest feature size in the group. align , expressed as:

[0089]

[0090] Among them, AvgPool represents the average pooling operation, Bilinear represents the bilinear interpolation operation, represents the concatenation and fusion operation; the average pooling operation reduces the amount of computation by reducing the resolution of the feature map while maintaining important low-level features; bilinear interpolation helps to smooth the image and prevent excessive information loss; during the alignment process, the size of the feature map is ensured to be consistent so that subsequent modules can process it effectively. In order to control the computational delay and strike a balance between speed and accuracy, the feature map F4 output by the fourth encoder is selected as the target size of the feature alignment, thereby ensuring effective computational control and information retention; the above operations not only ensure efficient aggregation of information, but also minimize the computational complexity of subsequent modules;

[0091] S22: The fused convolution layer includes a first convolution layer, a re-parameterized convolution module, and a second convolution layer connected in sequence; the re-parameterized convolution module includes a 1*1 convolution layer, a fusion layer, and a 3*3 convolution layer;

[0092] Align feature map F align Input to the first convolutional layer for convolution operation to obtain the feature map F′ align ;

[0093] Re-parameterized convolution module for feature map F′ align After performing 1*1 convolution, fusion and 3*3 convolution in sequence, the fusion feature F is obtained. fuse , the expression is:

[0094]

[0095] The number of intermediate channels is an adjustable parameter to adapt to models of different scales. The reparameterized convolution module optimizes the representation capability of the feature map by flexibly adjusting the parameters of the convolution kernel and the number of intermediate channels, and enables the model to adapt to input features of different scales. In this process, feature extraction and aggregation are performed through convolution operations.

[0096] The second convolutional layer combines the features F fuse Convolution is performed to obtain convolution fusion features, convolution fusion features and fusion features F fuse The number of channels is the same;

[0097] S23: The convolution fusion features generated by the fusion convolution module are input to the segmentation layer to calculate the number of channels. The channel dimension is segmented to obtain F α and F β , so the input of the aggregation-segmentation module is P i , the output after the segmentation layer is Fα and F β ;

[0098] S3: The attention fusion module uses the attention mechanism to fuse the two sets of features output by the segmentation layer;

[0099] S31: The weight fusion layer applies the attention mechanism to further improve the fusion effect of information, such as Figure 1 As shown;

[0100] S311: First feature map F α Use the convolution operation, and then apply the average pooling operation to obtain the feature map F α+ :

[0101] F α+ =AvgPool(Conv1(F α ))

[0102] The attention weight A is obtained based on the feature map through the activation function α :

[0103] A α =AvgPool(σ(Conv1(F α )))

[0104] Among them, σ represents the Sigmoid activation function; the Sigmoid activation function can generate a normalized weight value to reflect the importance of the input feature map;

[0105] S312: feature map F β Perform convolution operation and obtain F through average pooling β+ :

[0106] F β+ =AvgPool(Conv1)(F β ))

[0107] S313: Through the attention mechanism, F β+ With attention weight A α Fusion is performed to obtain the final enhanced feature map F αβ+ :

[0108]

[0109] in, Represents an element-by-element multiplication operation; the attention weight gives a greater weight to the more important parts of the feature map. Through this attention mechanism, the model can automatically focus on the key areas, effectively enhance the information of small targets, and thus strengthen the information of the key areas;

[0110] S32: The summing and fusion layer combines the feature map F through the summing operation.α+ and the enhanced feature map F αβ+ , further improve the feature expression ability and ensure the comprehensiveness and accuracy of information:

[0111]

[0112] S4: The fused and enhanced features are obtained through the re-parameterized convolution module. That is, the final output of this step is The re-parameterized convolution module includes a 1*1 convolution layer, a fusion layer, and a 3*3 convolution layer. After performing 1*1 convolution, fusion and 3*3 convolution in sequence, the fusion enhancement feature is obtained.

[0113] S5: Fusion of these enhanced features Convert to the final segmentation result;

[0114] S51: Perform upsampling operation; since convolution operation usually reduces the spatial dimension of feature map, it is necessary to Upsample to the same resolution as the input image;

[0115] Fusion Enhancement Features After the convolution operation of the 1*1 convolution layer, the image is input to the transposed convolution layer, which is implemented through transposed convolution (also called deconvolution). Let the upsampled feature map be U, and then a softmax activation function is used to convert U into a probability distribution, so as to obtain the final ship target segmentation result P.

[0116] Furthermore, during the training process, a loss function is used to measure the difference between the segmentation result predicted by the model and the true label; a variety of loss functions are combined to train the model to improve the accuracy and robustness of the segmentation; the mixed loss function used in the training model in this embodiment is expressed as:

[0117] L=ηL c +γL δ

[0118] Among them, η and γ are hyperparameters used to balance the weights of different loss functions; by minimizing these loss functions, the model can be trained to improve its performance on the segmentation task;

[0119] Cross quotient loss L c as follows:

[0120]

[0121] Among them, i represents the category index, Y i and P iRepresent the true label and predicted probability respectively;

[0122] Dice loss is a loss function that measures the similarity between two samples. δ It is expressed as:

[0123]

[0124] Among them, λ is a smoothing term to prevent the denominator from being zero; P is the probability distribution predicted by the model, and Y is the one-hot encoding representation of the true label of the input.

[0125] Furthermore, during the training process, the Adam optimizer is used to minimize the total loss function L. Through iterative optimization, the model learns how to more accurately predict the category label of each pixel, thereby improving the accuracy and robustness of the segmentation.

[0126] Furthermore, in order to further improve the performance of the model, data augmentation techniques such as random rotation, flipping, and scaling are introduced to increase the diversity of training data, prevent model overfitting, and improve the model's adaptability to different image transformations.

[0127] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0128] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image target segmentation method based on multi-level feature aggregation and segmentation, characterized in that: The following steps are involved: Step 1: Collect remote sensing ship images and build a training data set; Step 2: Construct a remote sensing image ship segmentation model and use the training data set to train the remote sensing image ship segmentation model; the remote sensing image ship segmentation model includes a feature extraction module, an aggregation-segmentation module, an attention fusion module, a re-parameterized convolution module and an output module connected in sequence; Step 3: Collect the remote sensing image to be segmented and input it into the trained remote sensing image ship segmentation model to obtain the ship target segmentation result.

2. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 1, characterized in that: The feature extraction module includes multiple groups of encoders, each group of encoders includes convolutional layers, nonlinear activation layers and pooling layers connected in sequence. Each convolutional layer performs a convolution operation to extract the spatial features in the remote sensing ship image to obtain a feature map. After the convolution operation, a nonlinear activation layer is applied to perform a nonlinear transformation on the feature map, and then down-sampled through the pooling operation of the pooling layer to reduce the spatial resolution of the feature map.

3. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 2, characterized in that: The convolution operation of the convolutional layer in the encoder is expressed as: F i =Conv(X i-1 ) Among them, X i-1 represents the input feature map of the i-1th layer in the remote sensing ship image, Conv represents the 2D convolution operation, and F i Represents the feature map obtained after convolution; The expression of the nonlinear transformation of the feature map by the nonlinear activation layer is: F′ i =ReLU(F i ) Among them, F′ i Represents the feature map after activation by the ReLU activation function; ReLU represents the ReLU activation function; The expression of the pooling operation of the pooling layer is: P i =MaxPool(F′ i ) Among them, P i Represents the feature map after pooling; MaxPool represents the maximum pooling operation.

4. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 1, characterized in that: The aggregation-segmentation module includes a feature alignment layer, a fusion convolution layer, and a segmentation layer connected in sequence; the attention fusion module includes a weight fusion layer and an addition fusion layer using an attention mechanism; The feature map extracted by the feature extraction module is downsampled in the feature alignment layer through average pooling and bilinear interpolation operations to obtain an aligned feature map. The fusion convolution layer performs a convolution operation on the aligned feature map to obtain a fusion feature. The segmentation layer segments the fused features in the channel dimension to obtain several local feature maps; The weight fusion layer uses the attention mechanism to perform weight fusion on the local feature map to obtain the enhanced feature map; the sum fusion layer performs the sum operation on the local feature map and the enhanced feature map to obtain the comprehensive feature map; The re-parameterized convolution module performs convolution operations on the comprehensive feature map to obtain fused enhanced features.

5. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 4, characterized in that: The feature map in the aggregation-segmentation module is downsampled in the feature alignment layer through average pooling and bilinear interpolation to obtain the aligned feature map, which is expressed as: Among them, AvgPool represents the average pooling operation; Bilinear represents the bilinear interpolation operation; represents the splicing and fusion operation; P1, P2, P3, and P4 respectively represent the feature maps output by a set of encoders of the feature extraction module; The fused convolution layer includes the first convolution layer, the re-parameterized convolution module and the second convolution layer connected in sequence; the re-parameterized convolution module includes a 1*1 convolution layer, a fusion layer and a 3*3 convolution layer; the alignment feature map F align Input to the first convolutional layer for convolution operation to obtain the feature map F′ align ; Re-parameterized convolution module for feature map F′ align After performing 1*1 convolution, fusion and 3*3 convolution in sequence, the fusion feature F is obtained. fuse ; The second convolutional layer combines the features F fuse Convolution is performed to obtain convolution fusion features, convolution fusion features and fusion features F fuse The number of channels is the same as that of the convolution module; the expression of the reparameterized convolution module is: Among them, F′ align Represents the feature map obtained after the convolution operation of the convolution layer; F fuse Indicates fusion features; Conv3 indicates a convolution operation with a convolution kernel of 3; Conv1 indicates a convolution operation with a convolution kernel of 1; The segmentation layer calculates the number of channels for the convolutional fusion features and performs segmentation based on the number of channels to obtain the first feature map F. α and the second feature map F β .

6. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 4, characterized in that: The process of feature processing by the attention fusion module and the reparameterized convolution module is as follows: The weight fusion layer performs convolution and average pooling operations on the first feature map output by the segmentation layer to obtain the first local feature map, which is expressed as: F α+ =AvgPool(Conv1(F α )) Among them, F α+ represents the first local feature map; F α Represents the first feature map of the segmentation layer output; AvgPool represents the average pooling operation; Conv1 represents the convolution operation with a convolution kernel of 1; The weight fusion layer performs convolution operation, activation function activation and average pooling operation on the first feature map output by the segmentation layer in sequence to obtain the attention weight, which is expressed as: A α =AvgPool(σ(Conv1(F α ))) Among them, A α represents the attention weight; σ represents the Sigmoid activation function; The weight fusion layer performs convolution and average pooling operations on the second feature map output by the segmentation layer in sequence to obtain the second local feature map, which is expressed as: F β+ =AvgPool(Conv1(F β )) Among them, F β Represents the second feature map of the segmentation layer output; F β+ represents the second local feature map; The weight fusion layer uses the attention mechanism to transform the second local feature map F β+ and attention weight A α Fusion is performed to obtain an enhanced feature map, the expression is: Among them, F αβ+ represents the enhanced feature map; Represents a pixel-by-pixel multiplication operation; The addition and fusion layer adds the first local feature map and the enhanced feature map to obtain a comprehensive feature map, which is expressed as: in, represents a comprehensive feature map; Re-parameterized convolution module for comprehensive feature maps After performing 1*1 convolution, fusion and 3*3 convolution in sequence, the fusion enhancement feature is obtained.

7. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 4, characterized in that: The output module includes an upsampling layer and a loss calculation layer; the upsampling layer includes a 1*1 convolution layer, a transposed convolution layer, and an activation function layer connected in sequence, and the fusion enhancement feature After the convolution operation of the 1*1 convolution layer, the input is sent to the transposed convolution layer, which enhances the fusion features after convolution. Upsample to the same resolution as the remote sensing ship image, and the activation function layer converts the upsampled feature map into a probability distribution to obtain the ship target segmentation result; The loss calculation layer calculates the loss value using a mixed loss function to optimize model training based on the ship target segmentation results output by the activation function layer and the ship target segmentation results marked in the remote sensing ship image.

8. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 7, characterized in that: The mixed loss function used in the loss calculation layer includes the cross entropy loss function and the Dice loss function; the mixed loss function is expressed as: L=ηL c +γL δ Among them, η and γ represent the hyperparameters used to balance the weights of different loss functions; L c represents the cross entropy loss; L δ represents Dice loss; Cross entropy loss L c It is expressed as: Among them, i represents the category index, Y i represents the segmentation result of the real ship target annotated in the image, P i It represents the ship target segmentation result predicted and output by the remote sensing image ship segmentation model; Dice loss L δ It is expressed as: Among them, λ represents a smoothing term to prevent the denominator from being zero; P represents the distribution of ship target segmentation results predicted and output by the remote sensing image ship segmentation model; Y represents the distribution of the actual ship target segmentation results marked in the image.

9. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 8, characterized in that: The Adam optimizer is used in the training process of the remote sensing image ship segmentation model to minimize the mixed loss function and optimize the model.

10. The remote sensing image target segmentation method based on multi-level feature aggregation and segmentation according to claim 1, characterized in that: Step 1 also includes data enhancement processing on the collected remote sensing ship images, expanding the images through random rotation, flipping and scaling of the images, and constructing a training data set.

Citation Information

Patent Citations

  • Rotating ship detection method based on remote sensing image feature extraction

    CN116630808A

  • Bulk cargo ship image point cloud fusion scanning segmentation method

    CN117252895A

  • Remote sensing image semantic segmentation method based on multi-scale feature fusion and attention mechanism

    CN117765409A

  • Cross-modal ship remote sensing data generation and key part segmentation method in near-field security and protection

    CN117953323A

  • Multi-scale feature optimized remote sensing image segmentation model and method

    CN118154868A

Cited By

  • Remote sensing image segmentation method and system fusing stage perception and multi-dimensional orientation mechanism

    CN121053541A

  • Remote sensing image segmentation method and system of fusion stage perception and multi-dimensional orientation mechanism

    CN121053541B