A method for ship segmentation in remote sensing images based on the U-KAN network model
By adopting the U-KAN network model in the remote sensing image ship segmentation technology and integrating the hybrid architecture of U-Net and KAN networks, the problem of insufficient robustness and generalization capabilities of the existing technology in complex backgrounds and diversified ship morphology processing is solved, and the ship segmentation effect with high precision and robustness is achieved.
Patent Information
- Application Number
- CN202411640212.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-18
AI Technical Summary
When existing remote sensing image ship segmentation technology deals with complex backgrounds and diversified ship forms, its robustness and generalization capabilities are weak, making it difficult to achieve efficient and accurate ship information extraction.
The remote sensing image ship segmentation method based on the U-KAN network model is adopted, and the hybrid network architecture of the U-Net network and KAN network is integrated to achieve the combination of multi-scale feature fusion and adaptive activation functions, enhancing the model's adaptability to complex backgrounds and diverse ship forms.
It significantly improves the accuracy and robustness of ship segmentation of remote sensing images, can effectively capture ship targets of different scales and complexities, and achieves effective fusion between shallow detail features and deep semantic features.
Smart Images

Figure CN119600604B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image processing, and more specifically to a remote sensing image ship segmentation method based on a U-KAN network model. Background Art
[0002] Ship segmentation technology in remote sensing images has important application value in many fields such as ocean monitoring, shipping safety, and national defense applications, and has always attracted the attention of researchers. With the development of remote sensing technology, it has become easier to obtain high-resolution remote sensing images, but how to efficiently and accurately extract ship information from these images remains a major challenge.
[0003] At present, the marine environment is complex and changeable, the background noise is large, and the ship shapes are diverse. Traditional methods based on visual saliency or template matching are easily disturbed when dealing with such complex backgrounds, resulting in low detection accuracy. In addition, these methods rely on preset specific features and have poor adaptability to changes in ship shapes, making it difficult to deal with different types of ship targets.
[0004] Furthermore, ships may appear in a variety of scales in remote sensing images, ranging from large aircraft carriers to small speedboats. Although traditional deep learning models can capture rich target information through multi-level feature extraction, they can only extract fixed hierarchical features when dealing with multi-scale targets, and it is difficult to fully express target features of different scales and complexities. This results in limited performance of the model when detecting small-scale ships or segmenting fine structures, and it is easy to lose details or blur the edges of the target.
[0005] In addition, existing remote sensing image segmentation methods are insufficient in fusing shallow detail features with deep semantic features. Although some methods try to retain low-level detail features through skip connections, in practical applications, these methods often fail to effectively fuse feature information at different levels, resulting in segmentation results that are not detailed enough, especially when dealing with complex backgrounds or target edges.
[0006] Existing models have weak robustness and generalization capabilities when processing images with complex backgrounds or large changes in target morphology, and are prone to false detection or missed detection. In addition, the fixed network structure and activation function limit the adaptability of the model.
[0007] Therefore, how to design a remote sensing image ship segmentation method based on the U-KAN network model to achieve effective fusion between shallow detail features and deep semantic features and overcome the limitations of existing technologies in dealing with complex backgrounds and diverse ship forms is an urgent problem that technicians in this field need to solve. Summary of the invention
[0008] In view of this, the present invention provides a remote sensing image ship segmentation method based on the U-KAN network model. By integrating the hybrid network architecture of the U-Net network and the KAN network, the combination of multi-scale feature fusion and adaptive activation function is realized, which significantly improves the accuracy and robustness of remote sensing image ship segmentation.
[0009] In order to achieve the above object, the present invention adopts the following technical solution:
[0010] A remote sensing image ship segmentation method based on a U-KAN network model comprises the following steps:
[0011] S1. Construct a U-KAN network model; the U-KAN network model is a hybrid network architecture integrating a U-Net network and a KAN network; wherein the U-Net network includes an encoder and a decoder with jump connections; and the KAN network is nested with multiple KAN layers;
[0012] S2, training the U-KAN network model with remote sensing image data sets to obtain an optimized U-KAN network model;
[0013] S3. Input the remote sensing image to be processed into the optimized U-KAN network model to obtain the corresponding ship segmentation result.
[0014] Furthermore, in S1, the encoder combines multiple first convolution blocks to reduce the spatial dimension of the feature map by downsampling to perform multi-scale feature extraction;
[0015] The first convolution block includes: convolution layer, batch normalization layer, ReLU activation function and pooling layer; the encoder output E(x) is expressed as:
[0016] E(x)=[C 1 (x),C 2 (P 1 (C 1 (x))),…,C n (P n-1 (C n-1 (…P 1 (C 1 (x))…))]
[0017] Among them, C i represents the i-th convolutional layer, P i represents the remote sensing image input by the i-th pooling layer, x.
[0018] Furthermore, in S1, the decoder combines multiple second convolution blocks to restore the spatial dimension of the feature map by upsampling, and after each upsampling, fuses the feature map from the jump connection and the feature map output by the encoder;
[0019] The second convolution block includes: a convolution layer, a batch normalization layer, and a ReLU activation function; the decoder output D(E(x), S) is expressed as:
[0020] D(E(x),S)=U{S,[C n+1 (U 1 (C n (…P 1 (C 1 (x))…))),…,C 2n (U n (P 1 (C 1 (x))))]}
[0021] Among them, S represents the feature map of the skip connection, U i represents the i-th upsampling and convolution operation.
[0022] Furthermore, in S1, the KAN network is represented as:
[0023] KAN(Z)=(Φ k-1 °Φ k-2 °…°Φ 1 °Φ 0 )Z
[0024] Among them, ° represents the composite operation of functions, taking one function as the input of another function; Φ i represents the i-th KAN layer of the KAN network. The input dimension and output dimension of each KAN layer are n respectively. in and n out , including n in ×n out Learnable activation functions
[0025] The calculation result Z from the kth KAN layer to the k+1th KAN layer in the KAN network k+1 It is expressed as:
[0026] Z k+1 =Φ k Z k
[0027]
[0028] in, represents the learnable activation function of the kth KAN layer.
[0029] Furthermore, in S1, constructing the U-KAN network model includes: based on the KAN network, embedding a plurality of KAN modules in the skip connection between the encoder and the decoder;
[0030] Each KAN module includes: a tokenization layer, a KAN layer, a depth-wise separable convolutional layer, and a layer normalization layer.
[0031] Furthermore, the tokenization layer is used to vectorize the encoder output and map the vector to a D-dimensional embedding space through a trainable linear projection; wherein the vectorization processing includes:
[0032] Design a multi-scale tokenization layer; each scale in the tokenization layer corresponds to a different patch size;
[0033] Perform tokenization operation on each scale to generate the corresponding vector representation;
[0034] The vectors at all scales are combined to generate a comprehensive feature description.
[0035] Furthermore, the KAN layer is used to dynamically adjust the response capability of the network model to features of different scales through an adaptive activation function; the adaptive activation function is expressed as:
[0036] F KAN (S, E) = σ(ω·concat(S, E)+b)
[0037] Among them, F KAN (S, E) represents the output of the KAN layer, concat(S, E) represents the concatenation of the feature map S of the jump connection and the feature map E output by the encoder, ω and b represent the network weight and bias respectively.
[0038] Furthermore, in S1, the U-KAN network model is also combined with an attention mechanism to increase the network model's attention to important feature areas and perform initial fusion of multi-scale features; the attention mechanism is expressed as:
[0039]
[0040] Where A(y) represents the attention weight; e(·) represents the attention score function, which maps the input feature y to a real value.
[0041] Furthermore, in S2, the mixed loss function L of the U-KAN network model training is seg It is expressed as:
[0042] L seg =λ ce ·L ce +λ d ·L d
[0043] Among them, L ce represents the cross entropy loss function, L d represents the Dice coefficient loss function, λce , d is the weighting coefficient.
[0044] Furthermore, the S3 includes:
[0045] S31, using an encoder to perform multi-scale feature extraction on the remote sensing image to be processed;
[0046] S32, combine the attention mechanism to initially fuse the multi-scale features to obtain the initial fused features; and pass the multi-scale features to the decoder through the jump connection embedded in the KAN module;
[0047] S33, the decoder fuses the feature map from the jump connection embedded in the KAN module and the initial fused feature to obtain the fused feature and generate a binary mask with the same size as the input image;
[0048] S34. Based on the binary mask, a clear ship boundary contour is generated through threshold processing, and a visualized ship segmentation result is generated in combination with the color information of the remote sensing image to be processed.
[0049] It can be seen from the above technical solution that compared with the prior art, the technical solution of the present invention has the following advantages:
[0050] Beneficial effects:
[0051] 1. Compared with the existing remote sensing image ship segmentation technology, this method integrates U-Net and KAN networks in network architecture design to form a hybrid network architecture. It not only inherits the powerful semantic segmentation capability of U-Net, but also enhances the ability to capture multi-scale features through the KAN network, and can dynamically adjust the network's response to features of different scales, significantly improving the model's adaptability to complex backgrounds and diverse ship forms.
[0052] 2. Based on the KAN network, the jump connection between the encoder and the decoder is embedded in the KAN module, which eliminates the dependence on the traditional linear weight matrix, so that the model can achieve superior performance at a smaller scale. The design of the tokenization layer further improves the efficiency of feature vectorization processing, laying a solid foundation for subsequent feature fusion and segmentation tasks. In particular, the adaptive activation function in the KAN layer can dynamically adjust the activation mode according to the input data, so that the model can respond to different types of inputs more flexibly, especially when dealing with small targets in complex backgrounds.
[0053] 3. The U-KAN model also introduces an attention mechanism, which enables the network to pay more attention to important feature areas and further improve the accuracy of segmentation. At the same time, the mixed loss function used in model training can effectively measure the difference between the predicted segmentation map and the actual segmentation map, guide the model to continuously optimize parameters, and thus achieve the U-KAN model's accurate segmentation of remote sensing image ships. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0055] Figure 1 A flow chart of a method for segmenting a remote sensing image ship based on a U-KAN network model provided in an embodiment of the present invention;
[0056] Figure 2 A schematic diagram of the U-KAN network model structure provided by an embodiment of the present invention;
[0057] Figure 3 A schematic diagram of the processing process of a remote sensing image to be processed by the U-KAN network model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] like Figure 1 As shown, this embodiment provides a remote sensing image ship segmentation method based on the U-KAN network model, comprising the following steps:
[0060] S1. Construct a U-KAN network model; the U-KAN network model is a hybrid network architecture integrating a U-Net network and a KAN network; wherein the U-Net network includes an encoder and a decoder with jump connections; and the KAN network is nested with multiple KAN layers;
[0061] S2, training the U-KAN network model with remote sensing image data sets to obtain an optimized U-KAN network model;
[0062] S3. Input the remote sensing image to be processed into the optimized U-KAN network model to obtain the corresponding ship segmentation result.
[0063] This method combines the hybrid network architecture of U-Net and KAN (Kolmogorov-Arnold Networks) to achieve the combination of multi-scale feature fusion and adaptive activation function, significantly improving the accuracy and robustness of ship segmentation in remote sensing images. It can not only effectively capture ship targets of different scales and complexities in remote sensing images, but also achieve effective fusion between shallow detail features and deep semantic features, thus overcoming the limitations of existing technologies in dealing with complex backgrounds and diverse ship morphologies.
[0064] The above steps and related features are further described in detail below:
[0065] In this embodiment S1, a U-KAN network model is constructed; the U-KAN network model is a hybrid network architecture integrating a U-Net network and a KAN network; wherein the U-Net network includes an encoder and a decoder with jump connections; and the KAN network is nested with multiple KAN layers;
[0066] Specifically, the specific implementation of the encoder-decoder architecture in the U-KAN network model is described in detail:
[0067] The encoder combines multiple first convolutional blocks to reduce the spatial dimension of the feature map by downsampling and performs multi-scale feature extraction;
[0068] The first convolution block includes: convolution layer, batch normalization layer, ReLU activation function and pooling layer; the encoder output E(x) is expressed as:
[0069] E(x)=[C 1 (x),C 2 (P 1 (C 1 (x))),…,C n (P n-1 (C n-1 (…P 1 (C 1 (x))…))]
[0070] Among them, C i represents the i-th convolutional layer, P i represents the remote sensing image input by the i-th pooling layer, x.
[0071] Furthermore, the decoder combines multiple second convolution blocks to restore the spatial dimension of the feature map by upsampling, and after each upsampling, the feature map from the skip connection and the feature map output by the encoder are fused;
[0072] The second convolution block includes: a convolution layer, a batch normalization layer, and a ReLU activation function; the decoder output D(E(x), S) is expressed as:
[0073] D(E(x),S)=U{S,[C n+1 (U 1 (C n (…P 1 (C 1 (x))…))),…,C 2n (U n (P 1 (C 1 (x))))]}
[0074] Among them, S represents the feature map of the skip connection, U i represents the i-th upsampling and convolution operation.
[0075] Here, the encoder combines multiple first convolution blocks and reduces the spatial dimension of the feature map by downsampling, thereby achieving the goal of multi-scale feature extraction. The decoder combines multiple second convolution blocks and restores the spatial dimension of the feature map by upsampling. After each upsampling, the decoder fuses the feature map from the jump connection with the feature map output by the encoder. This enables the model to effectively capture features of different scales and maintains spatial resolution through jump connections, thereby improving overall performance.
[0076] like Figure 2 As shown in the figure, in order to improve the selection and fusion of features, the model can process complex data sets more efficiently and obtain better segmentation results. Based on the KAN network, multiple KAN modules are embedded in the jump connection between the encoder and the decoder;
[0077] The KAN network is represented as:
[0078] KAN(Z)=(Φ k-1 °Φ k-2 °…°Φ 1 °Φ 0 )Z
[0079] Among them, ° represents the composite operation of functions, taking one function as the input of another function; Φ i represents the i-th KAN layer of the KAN network. The input dimension and output dimension of each KAN layer are n respectively. in and n out , including n in ×n out Learnable activation functions
[0080] The calculation result Z from the kth KAN layer to the k+1th KAN layer in the KAN networkk+1 It is expressed as:
[0081] Z k+1 =Φ k Z k
[0082]
[0083] in, represents the learnable activation function of the kth KAN layer.
[0084] Each KAN module is responsible for processing the feature map output by the encoder, including: tokenization layer, KAN layer, depth-wise separable convolution layer, and layer normalization layer.
[0085] The specific process of each KAN module processing the feature map output by the encoder includes:
[0086] First, the encoder output is vectorized through a tokenization layer, and the vector is mapped to a D-dimensional embedding space through a trainable linear projection. The vectorization process includes: designing a multi-scale tokenization layer; each scale in the tokenization layer corresponds to a different patch size; performing a tokenization operation on each scale to generate a corresponding vector representation; and merging the vectors at all scales to generate a comprehensive feature description.
[0087] In the tokenization layer, tokens of different patch sizes are generated to capture feature variations at different scales. The tokens at each scale are converted into vector representations of fixed length. These vectors are then merged to form a comprehensive feature description for subsequent KAN layer processing.
[0088] Next, the KAN layer is used to dynamically adjust the network model's responsiveness to features of different scales through an adaptive activation function; the adaptive activation function is expressed as:
[0089] F KAN (S, E) = σ(ω·concat(S, E)+b)
[0090] Among them, F KAN (S, E) represents the output of the KAN layer, concat(S, E) represents the concatenation of the feature map S of the jump connection and the feature map E output by the encoder, ω and b represent the network weight and bias respectively.
[0091] The adaptive activation function can automatically adjust its parameters according to the input features to optimize the model's responsiveness to features of different scales.
[0092] Subsequently, the output of the KAN layer enters the depthwise separable convolution layer, which is mainly used to further extract features. Compared with standard convolution, depthwise separable convolution has lower computational cost and can effectively capture the correlation between features.
[0093] Finally, the features are standardized through the layer normalization layer to ensure consistent feature distribution on each channel, which is beneficial to subsequent calculations and gradient propagation.
[0094] Furthermore, the U-KAN network model also combines the attention mechanism to increase the network model's attention to important feature areas and perform initial fusion of multi-scale features; the attention mechanism is expressed as:
[0095]
[0096] Where A(y) represents the attention weight; e(·) represents the attention score function, which maps the input feature y to a real value.
[0097] The above U-KAN network model combines the U-Net network and the KAN (Kolmogorov-Arnold Networks) network to form a hybrid network structure. The U-Net part realizes multi-scale feature extraction and spatial dimension recovery of the input image through the encoder-decoder framework and its jump connection mechanism; on this basis, based on the KAN network, multiple KAN modules are embedded in the jump connection, especially the adaptive activation function in the CAN layer, which further enhances the model's feature processing and generalization capabilities. And combined with the attention mechanism, it strengthens the capture of key area features, thereby showing superior performance when processing complex backgrounds and diverse ship forms.
[0098] In this embodiment S2, a U-KAN network model is trained in combination with a remote sensing image data set to obtain an optimized U-KAN network model;
[0099] In the process of training the U-KAN network model, a high-quality remote sensing image dataset is prepared, which should contain a large number of annotated samples covering different types of ships, background environments, and lighting conditions. The diversity and richness of the dataset are crucial to the generalization ability of the model. Subsequently, the dataset is divided into training, validation, and test sets to ensure that the model can fully learn during the training process and evaluate its performance in the validation and testing stages.
[0100] During the training process, the model is trained end-to-end using optimization algorithms such as stochastic gradient descent (SGD) or Adam, combined with a hybrid loss function. The hybrid loss function can effectively measure the difference between the model prediction and the true label, prompting the model to converge quickly and reach the optimal state during the training process.
[0101] Specifically, the mixed loss function L for U-KAN network model training is seg It is expressed as:
[0102] L seg =λ ce ·L ce +λ d ·L d
[0103] Among them, L ce represents the cross entropy loss function, L d represents the Dice coefficient loss function, λ ce , d is the weighting coefficient.
[0104] In addition, in order to prevent overfitting, data enhancement techniques (such as rotation, flipping, cropping, etc.) are used to increase the diversity of training data, while regularization methods (such as L2 regularization) and early stopping strategies are used to control the complexity of the model.
[0105] Through the above steps, an optimized U-KAN network model is finally obtained.
[0106] In this embodiment S3, the remote sensing image to be processed is input into the optimized U-KAN network model to obtain the corresponding ship segmentation result.
[0107] In this step, there is a remote sensing image containing one or more ships, with the background being the sea surface, which may contain complex factors such as waves and clouds. Figure 3 As shown, the optimized U-KAN network model is used to segment the ship in the image, which specifically includes the following steps:
[0108] S31, using an encoder to perform multi-scale feature extraction on the remote sensing image to be processed;
[0109] The remote sensing image to be processed is input into the encoder part of the U-KAN network model. The encoder downsamples through multiple convolutional blocks, gradually reducing the spatial dimension of the feature map while increasing the number of channels to capture more abstract features.
[0110] S32, combine the attention mechanism to initially fuse the multi-scale features to obtain the initial fused features; and pass the multi-scale features to the decoder through the jump connection embedded in the KAN module;
[0111] The attention mechanism is applied to the multi-scale feature map output by the encoder to perform weighted fusion of features of different scales to generate initial fused features. The attention mechanism calculates the importance weight of each feature, allowing the model to pay more attention to important feature areas, including the edges and details of the ship.
[0112] The initial fused features are passed to the decoder part through jump connections. Each jump connection embeds a KAN module, which includes a tokenization layer, a KAN layer, a depth-separable convolution layer, and a layer normalization layer. These modules further process the features, especially when dealing with small objects or subtle structures, to enhance the model's ability to respond to features of different scales.
[0113] S33, the decoder fuses the feature map from the jump connection embedded in the KAN module and the initial fused feature to obtain the fused feature and generate a binary mask with the same size as the input image;
[0114] The decoder performs upsampling through multiple convolution blocks to gradually restore the spatial dimensions of the feature map. After each upsampling, the decoder fuses the feature map from the jump connection with the feature map output by the encoder to generate a fused feature.
[0115] Then a binary mask of the same size as the input image is generated. Each pixel value in this binary mask indicates whether the pixel belongs to the ship area. For example, a pixel with a value of 1 represents a ship, and a pixel with a value of 0 represents the background.
[0116] S34. Based on the binary mask, a clear ship boundary contour is generated through threshold processing, and a visualized ship segmentation result is generated in combination with the color information of the remote sensing image to be processed.
[0117] The generated binary mask is thresholded to generate a clear ship boundary outline. Specifically, a suitable threshold is selected, and the pixel values in the mask greater than the threshold are set to 1 (indicating a ship), and the pixel values less than or equal to the threshold are set to 0 (indicating the background). This ensures that the boundary of the segmentation result is clearer.
[0118] Furthermore, the generated binary mask is overlaid on the original remote sensing image, and the final visual ship segmentation result is generated by combining the color information of the image. The ship area is highlighted in the image by color overlay or transparency adjustment. Specifically, the ship area is displayed in red and the background is displayed in blue, so as to intuitively show the position and shape of the ship.
[0119] This embodiment proposes a remote sensing image ship segmentation method based on the U-KAN network model. The method combines the advantages of the U-Net network and the KAN (Kolmogorov-Arnold Networks) network. By constructing an encoder-decoder architecture with jump connections and embedding the KAN module in the jump connections, efficient extraction and fusion of multi-scale features are achieved. During the training process, the model is optimized using a high-quality remote sensing image dataset to ensure that the model has good generalization ability and robustness. In practical applications, the remote sensing image to be processed is input into the optimized U-KAN network model, and multi-scale feature extraction is performed through the encoder. The decoder restores the spatial dimension of the feature map and generates a binary mask. Finally, the attention mechanism and threshold processing are combined to generate a clear ship segmentation result, which significantly improves the recognition accuracy of ship targets in remote sensing images.
[0120] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0121] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image ship segmentation method based on the U-KAN network model, characterized in that: The following steps are involved: S1. Construct a U-KAN network model; the U-KAN network model is a hybrid network architecture integrating a U-Net network and a KAN network; wherein the U-Net network includes an encoder and a decoder with jump connections; and the KAN network is nested with multiple KAN layers; Building a U-KAN network model includes: based on the KAN network, embedding multiple KAN modules into the jump connection between the encoder and the decoder; wherein each KAN module includes: a tokenization layer, a KAN layer, a depth-separable convolution layer, and a layer normalization layer; The tokenization layer is used to vectorize the encoder output and map the vector to a D-dimensional embedding space through a trainable linear projection; wherein the vectorization process includes: Design a multi-scale tokenization layer; each scale in the tokenization layer corresponds to a different patch size; Perform tokenization operation on each scale to generate the corresponding vector representation; Merge vectors at all scales to generate a comprehensive feature description; The KAN layer is used to dynamically adjust the response capability of the network model to features of different scales through an adaptive activation function; the adaptive activation function is expressed as: F KAN (S,E)=σ(ω concat(S,E)+b) Among them, F KAN (S, E) represents the output of the KAN layer, concat(S, E) represents the concatenation of the feature map S of the jump connection and the feature map E output by the encoder, ω and b represent the network weight and bias respectively; S2, training the U-KAN network model with remote sensing image data sets to obtain an optimized U-KAN network model; S3. Input the remote sensing image to be processed into the optimized U-KAN network model to obtain the corresponding ship segmentation result.
2. According to the method for ship segmentation in remote sensing images based on the U-KAN network model in claim 1, it is characterized in that: In S1, the encoder combines multiple first convolution blocks to reduce the spatial dimension of the feature map by downsampling to perform multi-scale feature extraction; The first convolution block includes: convolution layer, batch normalization layer, ReLU activation function and pooling layer; the encoder output E(x) is expressed as: E(x)=[C1(x),C2(P1(C1(x))),…,C n (P n-1 (C n-1 (…P1(C1(x))…))] Among them, C i represents the i-th convolutional layer, P i represents the remote sensing image input by the i-th pooling layer, x.
3. According to the method for ship segmentation in remote sensing images based on the U-KAN network model in claim 1, it is characterized in that: In S1, the decoder combines multiple second convolution blocks to restore the spatial dimension of the feature map by upsampling, and after each upsampling, fuses the feature map from the jump connection and the feature map output by the encoder; The second convolution block includes: a convolution layer, a batch normalization layer, and a ReLU activation function; the decoder output D(E(x), S) is expressed as: D(E(x),S)=U{S,[C n+1 (U1(C n (…P1(C1(x))…))),…,C 2n (U n (P1(C1(x))))]} Among them, S represents the feature map of the skip connection, U i represents the i-th upsampling and convolution operation.
4. According to the method for ship segmentation in remote sensing images based on the U-KAN network model in claim 1, it is characterized in that: In S1, the KAN network is represented as: KAN(Z)=(Φ k-1 °F k-2 °…°Φ1°Φ0)Z Among them, Z represents input data, ° represents the composite operation of functions, taking one function as the input of another function; Φ i represents the i-th KAN layer of the KAN network. The input dimension and output dimension of each KAN layer are n respectively. in and n out , including n in ×n out Learnable activation functions The calculation result Z from the kth KAN layer to the k+1th KAN layer in the KAN network k+1 It is expressed as: WITH k+1 =Φ k WITH k in, represents the learnable activation function of the kth KAN layer.
5. According to claim 1, a remote sensing image ship segmentation method based on U-KAN network model is characterized in that: In S1, the U-KAN network model is also combined with an attention mechanism to increase the network model's attention to important feature areas and perform initial fusion of multi-scale features; the attention mechanism is expressed as: Where A(y) represents the attention weight; e(·) represents the attention score function, which maps the input feature y to a real value.
6. The method for ship segmentation in remote sensing images based on the U-KAN network model according to claim 1, characterized in that: In S2, the mixed loss function L for U-KAN network model training is seg It is expressed as: L seg =λ ce ·L ce +λ d ·L d Among them, L ce represents the cross entropy loss function, L d represents the Dice coefficient loss function, λ ce , d is the weighting coefficient.
7. The method for ship segmentation in remote sensing images based on the U-KAN network model according to claim 1, characterized in that: The S3 includes: S31, using an encoder to perform multi-scale feature extraction on the remote sensing image to be processed; S32, combine the attention mechanism to initially fuse the multi-scale features to obtain the initial fused features; and pass the multi-scale features to the decoder through the jump connection embedded in the KAN module; S33, the decoder fuses the feature map from the jump connection embedded in the KAN module and the initial fused feature to obtain the fused feature and generate a binary mask with the same size as the input image; S34. Based on the binary mask, a clear ship boundary contour is generated through threshold processing, and a visualized ship segmentation result is generated in combination with the color information of the remote sensing image to be processed.
Citation Information
Patent Citations
Remote-sensing image river extraction device and method
CN111383230A
KIA Net network model and image thereof, and high-precision wafer defect detection and segmentation method
CN118762013A