Remote sensing image ship semantic segmentation method and system based on feature hyper-fusion module

By combining the feature hyper-fusion module with multi-scale convolution, dilated convolution and dual-channel attention mechanism, the problems of scale variation and background interference in the semantic segmentation of ships in remote sensing images are solved, the segmentation accuracy and robustness are improved, and the computational efficiency is optimized.

CN120125996BActive Publication Date: 2025-09-19耕宇牧星(北京)空间科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510183503.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-09-19
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

Existing semantic segmentation methods for ships in remote sensing images have difficulty achieving high-precision segmentation when faced with problems of different scales, complex backgrounds, and insufficient feature fusion. In particular, when the contrast between the ship target and the background is low or there is noise, the model has difficulty effectively distinguishing the target from the background.

Method used

A method based on feature super-fusion module is adopted. Through multi-scale convolution, dilated convolution, global average pooling and dual-channel attention mechanism, combined with cross entropy loss and focal loss, feature fusion and enhancement are optimized to improve the model's detection ability and semantic segmentation accuracy for ship targets.

Benefits of technology

The accuracy and robustness of ship semantic segmentation in remote sensing images are significantly improved, background interference is reduced, the recognition performance of the model in complex scenes is enhanced, and computational efficiency is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125996B_ABST
    Figure CN120125996B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module, belonging to the field of remote sensing image processing technology. The method comprises the following steps: obtaining a remote sensing ship image dataset and performing preliminary feature extraction on the remote sensing ship images in the remote sensing ship image dataset to obtain initial features of the remote sensing ship images; using the feature hyper-fusion module to perform feature fusion and enhancement on the initial features of the remote sensing ship images to obtain a fused and enhanced feature map of the remote sensing ship images; performing feature extraction on the fused and enhanced feature map of the remote sensing ship images to obtain a feature representation for semantic segmentation, and obtaining a semantic segmentation result based on the feature representation of semantic segmentation. The present invention effectively improves the model's ability to detect and semantically segment ship targets through means such as feature fusion, enhancement, and adaptive calibration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing image processing, and in particular relates to a method and system for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module. Background Art

[0002] The task of semantic segmentation of ships in remote sensing images is an important application in remote sensing image processing, especially in the fields of ocean monitoring, shipping management, and environmental protection. With the continuous development of remote sensing technology, the resolution of remote sensing images has continued to increase, and the areas and types of targets covered have also increased. This has made the task of semantic segmentation of ships increasingly complex. Ship targets in remote sensing images are often affected by various factors such as different scales, shapes, materials, backgrounds, and weather conditions. Therefore, traditional semantic segmentation techniques based on manual features and shallow learning methods have been unable to meet the requirements of high-precision semantic segmentation. With the rise of deep learning technology, automatic feature extraction methods based on convolutional neural networks (CNNs) have shown great potential and advantages in the semantic segmentation of remote sensing images.

[0003] Currently, many methods for semantic segmentation of ships in remote sensing images rely on deep learning networks to automatically extract features and perform semantic segmentation. Common network architectures include convolutional neural networks (CNNs), fully convolutional neural networks (FCNs), residual networks (ResNets), and other variants. However, due to the diverse scales and shapes of ship targets in remote sensing images, as well as complex backgrounds, these deep networks face the following challenges:

[0004] 1) Scale Inconsistency: Ships in remote sensing images vary significantly in scale. Traditional convolutional layers, due to their fixed receptive fields, struggle to capture both large-scale and small-scale ship features. To address this issue, some studies have introduced multi-scale convolutional layers or dilated convolutional layers. However, these methods still suffer from insufficient information fusion and are unable to effectively integrate information from different scales.

[0005] 2) Background interference: Ship targets in remote sensing images are often interfered by complex backgrounds. Especially when the contrast between the ship target and the background is low or there is background noise such as water surface and clouds, it is difficult for the model to effectively distinguish the target from the background.

[0006] 3) Feature fusion and enhancement: Many current methods use simple splicing or weighted fusion strategies for feature fusion, which may lead to information loss or redundancy in the fusion process of multiple layers of features, affecting the final performance of the model.

[0007] Therefore, how to improve the semantic segmentation task of remote sensing images of ships and better solve problems such as scale changes of ship targets, background interference or insufficient feature fusion is still an urgent problem to be solved by technicians in this field. Summary of the Invention

[0008] In view of this, the present invention provides a method and system for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module, which is used to at least solve some of the technical problems in the background technology.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] The present invention first discloses a method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module, comprising the following steps:

[0011] Acquire a remote sensing ship image dataset, perform preliminary feature extraction on the remote sensing ship images in the remote sensing ship image dataset, and obtain initial features of the remote sensing ship images;

[0012] The feature hyper-fusion module is used to fuse and enhance the initial features of the remote sensing ship image to obtain the fusion enhanced feature map of the remote sensing ship image;

[0013] Feature extraction is performed on the fusion enhanced feature map of the remote sensing ship image to obtain the feature representation for semantic segmentation, and the semantic segmentation result is obtained based on the feature representation of semantic segmentation.

[0014] Furthermore, preliminary feature extraction is performed on the remote sensing ship images in the remote sensing ship image dataset to obtain initial features of the remote sensing ship images, specifically including the following steps:

[0015] S11. preprocessing the remote sensing ship images in the remote sensing ship image dataset;

[0016] S12. Utilize an initial feature extraction module composed of multiple convolution units to extract preliminary features from the preprocessed remote sensing ship image.

[0017] Furthermore, in step S12, the initial feature extraction module composed of multiple convolution units includes an initial convolution unit, a deep convolution unit, and a jump connection unit; the initial convolution unit and the deep convolution unit each include a convolution layer, a nonlinear activation function layer, a batch normalization layer, and a pooling layer connected in sequence;

[0018] The initial convolution unit is used to extract shallow features from the pre-processed remote sensing ship image;

[0019] The deep convolution unit is used to extract the shallow features again to obtain the deep features of the remote sensing ship image;

[0020] The jump connection unit is used to combine the shallow features output by the initial convolution unit with the deep features output by the deep convolution unit to obtain the final initial features of the remote sensing ship image.

[0021] Furthermore, the feature hyper-fusion module is used to fuse and enhance the initial features of the remote sensing ship image to obtain a fusion-enhanced feature map of the remote sensing image, which specifically includes the following steps:

[0022] The initial features are sequentially input into the 5*5 convolution layer, the first batch normalization layer and the first activation function layer in the feature super fusion module to obtain the first feature map

[0023] The initial features are sequentially input into the dilated convolution layer, the second batch normalization layer, and the second activation function layer in the feature super-fusion module to obtain the second feature map.

[0024] The activation functions in the first activation function layer and the second activation function layer are both LeakyReLu activation functions; c is the number of channels, h and w are the height and width of the feature map respectively;

[0025] Based on element addition, the first feature map X α and the second feature map X β Fusion is performed to obtain the first fusion feature map

[0026] in Represents element addition operation;

[0027] The first fusion feature map X G Input the maximum pooling layer and the fully connected layer in sequence, and perform Sigmoid function activation operation on the output of the fully connected layer to obtain the channel attention map A α And supplement channel attention map A β , the formula is:

[0028] A α =Sigmoid(FC(MaxPool(X G )));

[0029] Aβ=1-A α ;

[0030] Among them, A α ∈[0, 1] c×1×1 Indicates the importance of each channel;

[0031] Using channel attention map A α For the first feature map X α Perform adaptive correction to obtain the first correction feature map X′α :

[0032] Using the supplementary channel attention map A β For the second feature map X β Perform adaptive correction to obtain the second correction feature map X′ β :

[0033] in, Represents element-wise multiplication operation;

[0034] Based on the convolution operation, the first corrected feature map X′ α and the second correction characteristic map X′ β The fusion is performed to obtain the second fusion feature map X′:

[0035]

[0036] The initial features are sequentially input into a convolutional layer, the third batch normalization layer, and the activation function is used for nonlinear transformation to obtain the intermediate feature map, and then the intermediate feature map is input into a convolutional layer again to obtain the enhanced feature map X′ O ;

[0037] The second fusion feature map and the enhanced feature map X′ O Perform feature fusion and use the ReLU activation function to activate and obtain the final fusion enhanced feature map

[0038]

[0039] Among them, ReLU represents the ReLU activation function.

[0040] Furthermore, feature extraction is performed on the fusion enhanced feature map of the remote sensing ship image to obtain feature representation for semantic segmentation, and a semantic segmentation result is obtained based on the feature representation of semantic segmentation, which specifically includes the following steps:

[0041] Inputting the fused enhanced feature map into a flattening layer composed of multiple fully connected layers, using the flattening layer to map the fused enhanced feature map to a low-dimensional space, and extracting feature representation for semantic segmentation;

[0042] The feature representation obtained for semantic segmentation is nonlinearly transformed using an activation function to obtain a nonlinear feature map Y;

[0043] The nonlinear feature map Y is input into the semantic segmentation layer to obtain the semantic segmentation result.

[0044] Furthermore, the nonlinear feature map Y is input into the semantic segmentation layer to obtain a semantic segmentation result, which is expressed by the following formula:

[0045]

[0046] Among them, w cls and b cls Represent the weight matrix and bias of the semantic segmentation layer, Represents the probability distribution of each class, and softmax represents the Softmax activation function.

[0047] Furthermore, during the model training process to obtain semantic segmentation results, the following loss functions are specifically included:

[0048]

[0049] Among them, λ is a hyperparameter, Lcls represents the cross entropy loss, represents Dice loss, and Lfl represents focal loss.

[0050] Furthermore, the focus loss specifically includes the following formula:

[0051]

[0052] Among them, α i is the balancing factor for each category, and γ is the adjustment factor.

[0053] The present invention also discloses a remote sensing image ship semantic segmentation system based on a feature hyper-fusion module, which includes an initial feature extraction module, a feature hyper-fusion module and a semantic segmentation result acquisition module;

[0054] The initial feature extraction module is used to obtain a remote sensing ship image dataset and perform preliminary feature extraction on the remote sensing ship images in the remote sensing ship image dataset to obtain initial features of the remote sensing ship images;

[0055] The feature super fusion module is used to fuse and enhance the initial features to obtain fusion enhanced features;

[0056] The semantic segmentation result acquisition module is used to extract features from the fused enhanced feature map, obtain feature representation for semantic segmentation, and obtain a semantic segmentation result based on the feature representation of semantic segmentation.

[0057] Preferably, in the above system disclosed in the present invention, the feature hyper-fusion module performs feature fusion and enhancement on the initial features to obtain fusion-enhanced features, which specifically includes the following steps:

[0058] The initial features are sequentially input into the 5*5 convolution layer, the first batch normalization layer and the first activation function layer in the feature super fusion module to obtain the first feature map

[0059] The initial features are sequentially input into the dilated convolution layer, the second batch normalization layer, and the second activation function layer in the feature super-fusion module to obtain the second feature map.

[0060] The activation functions in the first activation function layer and the second activation function layer are both LeakyReLu activation functions; c is the number of channels, h and w are the height and width of the feature map respectively;

[0061] Based on element addition, the first feature map X α and the second feature map X β Fusion is performed to obtain the first fusion feature map

[0062] in Represents element addition operation;

[0063] The first fusion feature map X G Input the maximum pooling layer and the fully connected layer in sequence, and perform Sigmoid function activation operation on the output of the fully connected layer to obtain the channel attention map A α And supplement channel attention map A β , the formula is:

[0064] A α =Sigmoid(FC(MaxPool(X G )));

[0065] A β =1-A α ;

[0066] Among them, A α ∈[0, 1] c×1×1 Indicates the importance of each channel;

[0067] Using channel attention map A α For the first feature map X α Perform adaptive correction to obtain the first correction feature map X′ α :

[0068] Using the supplementary channel attention map A β For the second feature map X β Perform adaptive correction to obtain the second correction feature map X′ β :

[0069] in, Represents element-wise multiplication operation;

[0070] Based on the convolution operation, the first corrected feature map X′α and the second correction characteristic map X′ β The fusion is performed to obtain the second fusion feature map X′:

[0071]

[0072] The initial features are sequentially input into a convolutional layer, the third batch normalization layer, and the activation function is used for nonlinear transformation to obtain the intermediate feature map, and then the intermediate feature map is input into a convolutional layer again to obtain the enhanced feature map X′ O ;

[0073] The second fusion feature map and the enhanced feature map X′ O Perform feature fusion and use the ReLU activation function to activate and obtain the final fusion enhanced feature map

[0074]

[0075] Among them, ReLU represents the ReLU activation function.

[0076] It can be seen from the above technical solution that, compared with the prior art, the present invention discloses a method and system for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module, which has the following beneficial effects:

[0077] Through the design of the feature hyper-fusion module, the present invention can effectively combine multi-scale information and contextual features to enhance the detection capability of ship targets, especially in complex remote sensing image backgrounds, significantly improving the semantic segmentation accuracy and robustness.

[0078] The present invention adopts a dual-channel attention mechanism to dynamically adjust the feature map and adaptively calibrate the weights, thereby effectively reducing the interference of irrelevant information, improving the model's ability to focus on the key features of the ship, enhancing the recognition performance in different scenarios, and optimizing the computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0080] Figure 1 A schematic diagram of the overall system architecture provided by the present invention that can implement the semantic segmentation method of the present invention.

[0081] Figure 2A schematic diagram of the structure of the initial feature extraction module provided in an embodiment of the present invention.

[0082] Figure 3 A schematic diagram of the structure of the feature hyper-convergence module provided in an embodiment of the present invention.

[0083] Figure 4 Schematic diagram of the structure of the semantic segmentation result acquisition module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0084] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0085] This paper proposes a method for semantic segmentation of ships in remote sensing images based on a feature hyperfusion module. This method combines multi-scale convolution, dilated convolution, global average pooling (GAP), and a dual-channel attention mechanism. Through feature fusion, enhancement, and adaptive calibration, it effectively improves the model's ability to detect and segment ship targets. First, multi-scale and dilated convolutions are used. To effectively capture ship features at different scales, this paper proposes a combination of 5×5 and dilated convolution layers for preliminary feature extraction. The 5×5 convolution layers help capture local details, while the dilated convolution layers capture a wider range of contextual information while maintaining the receptive field. The combination of multi-scale features helps the model better cope with the variations in ship targets at different scales, enhancing semantic segmentation accuracy. Second, a dual-attention mechanism is also used. To address the contrast issue between ship targets and background in remote sensing images, this paper designs a dual-channel attention mechanism. By calculating and supplementing channel attention maps, the model can dynamically adjust its focus on different ship features based on the image content. The channel attention mechanism effectively emphasizes key ship features and mitigates the impact of background noise, thereby improving semantic segmentation performance. Finally, using a composite loss function, this paper constructs a composite loss function to effectively handle the class imbalance problem in semantic segmentation tasks. By combining cross-entropy loss and focal loss, this ensures that the model not only correctly performs semantic segmentation but also pays more attention to samples that are difficult to segment, improving semantic segmentation performance under class imbalance.

[0086] In a specific embodiment, the present invention discloses a method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module, comprising the following steps:

[0087] Step 1: Obtain a remote sensing ship image dataset, and perform preliminary feature extraction on the remote sensing ship images in the remote sensing ship image dataset to obtain the initial features of the remote sensing ship images.

[0088] Step 2: Construct a feature hyper-fusion module and use it to fuse and enhance the initial features to obtain fusion-enhanced features.

[0089] Step 3: Perform feature extraction on the fused enhanced feature map to obtain feature representation for semantic segmentation, and obtain the semantic segmentation result based on the feature representation of semantic segmentation.

[0090] In order to implement the above method, the embodiment of the present invention also discloses a remote sensing image ship semantic segmentation system based on a feature hyper-fusion module, such as Figure 1 As shown, it includes an initial feature extraction module, a feature super-fusion module and a semantic segmentation result acquisition module; wherein, the initial feature extraction module is used to obtain a remote sensing ship image dataset, and perform preliminary feature extraction on the remote sensing ship images of the remote sensing ship image dataset to obtain the initial features of the remote sensing ship images; the feature super-fusion module is used to perform feature fusion and enhancement on the initial features to obtain fusion enhanced features; the semantic segmentation result acquisition module is used to perform feature extraction on the fusion enhanced feature map, obtain feature representation for semantic segmentation, and obtain semantic segmentation results based on the feature representation of semantic segmentation.

[0091] The remote sensing image ship semantic segmentation method based on the feature hyper-fusion module of the present invention not only improves the accuracy of segmentation, but also reduces the resource requirements of model training and deployment through efficient parameter fine-tuning and the use of adapters, making it suitable for remote sensing image processing tasks in various computing environments.

[0092] In the embodiment of the present invention, the goal of step 1 of the semantic segmentation method is to extract an initial feature map from the input remote sensing ship image, and generate initial features (or initial feature maps) through image preprocessing (such as normalization, resizing and data enhancement) and preliminary convolution, activation, pooling and skip connection operations, laying the foundation for subsequent feature fusion and enhancement.

[0093] Specifically, step one includes the following steps:

[0094] Step 1.1 Preprocess the remote sensing ship image dataset.

[0095] First, the remote sensing ship images in the remote sensing ship image dataset are normalized to ensure that the pixel values ​​of each input image are within an appropriate range (usually [0, 1] or [-1, 1]). Normalization helps accelerate training and improve model stability.

[0096] Secondly, resizing. The sizes of each remote sensing ship image in the remote sensing ship image dataset may be inconsistent. Therefore, the normalized image needs to be resized to ensure that all input remote sensing ship images are of the same size. Usually, each remote sensing ship image is resized to a fixed size of h×w based on network requirements.

[0097] Finally, to increase the diversity of the training data and prevent overfitting, we perform data augmentation on the image data in the remote sensing ship image dataset, including rotation, flipping, scaling, and cropping. These operations help improve the model's generalization ability, especially in remote sensing ship images, where the scale, angle, and background of the ship can vary greatly.

[0098] Step 1.2: Extract the initial features of each remote sensing ship image in the preprocessed remote sensing ship image dataset.

[0099] In a specific embodiment, the initial features of the image are extracted by an initial feature extraction module composed of several convolution units. Specifically, Figure 2 As shown, the initial feature extraction module includes an initial convolution unit and a deep convolution unit connected to each other. The initial convolution unit and the deep convolution unit have the same structure, both of which include a convolution layer, a nonlinear activation function layer, a batch normalization layer and a pooling layer connected in sequence. The nonlinear activation function layer is used to perform a nonlinear transformation on the output of the convolution layer using a nonlinear activation function (such as ReLU), effectively solving the gradient vanishing problem and allowing the network to learn more complex feature representations. The addition of the batch normalization layer can improve training efficiency, reduce internal covariate shift, help accelerate convergence, and improve the stability of the feature extraction network. The addition of the pooling layer can reduce the size of the feature map through the pooling operation and maintain the most significant feature information, which will help reduce the amount of calculation and improve the robustness of the feature. Specifically, the pooling operation includes maximum pooling or average pooling operation.

[0100] In this embodiment, Figure 2 As shown in FIG, the initial feature extraction module is composed of an initial convolution unit and a deep convolution unit. The initial convolution unit can obtain shallow features of the pre-processed remote sensing ship image (usually including edges, colors, simple textures, etc.). The shallow features are input to the deep convolution unit again to obtain the deep features of the input remote sensing ship image (usually including more complex textures, shapes, context information, etc.). In order to avoid information loss, a skip connection is used in this embodiment to combine the shallow features output by the initial convolution unit with the deep features output by the deep convolution unit to obtain the final initial features X of the input remote sensing ship image. O .

[0101] The initial feature extraction module uses skip connections to combine shallow features with deep features, which helps to avoid losing low-level feature information when extracting high-level features in the deep layer of the network. This approach helps to maintain the details and semantic information of the image, especially in remote sensing image analysis, where details and semantic information are crucial. The initial feature X O It will be used as the input of the feature hyper-fusion module step of the present invention.

[0102] In a specific operation process, the convolution layer in the initial feature extraction module of the present invention can perform local operations on the image through a sliding window to extract shallow features in the remote sensing ship image.

[0103] Step 2: Use the feature hyper-fusion module to fusion the initial feature X O Perform feature fusion and enhancement to obtain fusion enhancement features. The feature super fusion module constructed in the embodiment of the present invention has an overall structure as follows Figure 3 shown.

[0104] The feature hyper-fusion module constructed in this paper processes initial features separately through multi-scale convolutional layers and dilated convolutional layers, and fuses them through global average pooling to obtain comprehensive information. Next, a dual-channel attention mechanism dynamically adjusts the weights of the feature maps, reducing interference from irrelevant information and thus improving ship detection. Finally, through feature fusion, enhancement, and further processing, features with higher discriminative power are generated, preparing for subsequent semantic segmentation tasks.

[0105] Specifically, refer to Figure 3 As shown, the feature hyper-fusion module constructed by the present invention specifically performs the following steps:

[0106] Step 2.1: Initial feature X obtained in step 1 O After passing through the 5×5 convolution layer and the dilated convolution layer, and respectively undergoing batch normalization and activation function, the feature maps are obtained. and Where c is the number of channels, h and w are the height and width of the feature map respectively, and the formula is expressed as:

[0107] X α =LeakyReLu(BN(Conυ 5×5 (X O ))),

[0108] X β =LeakyReLu(BN(DConυ 5×5 (X O )))

[0109] Here, Conv represents a convolutional layer, DConv represents a dilated convolutional layer, BN represents batch normalization, and LeakyReLu represents the LeakyReLu activation function. Ships in remote sensing images often have different scales and shapes. The feature maps extracted by the 5×5 convolutional layer can capture local details, while the dilated convolutional layer (with a different receptive field) can capture broader contextual information. This combination of multi-scale features is crucial for detecting ship targets of different scales.

[0110] Step 2.2: Global average pooling (GAP) and fusion operation. Add X α and X β Fusion into a new feature map It can be recorded as the first fusion feature map, and the formula is expressed as:

[0111]

[0112] in, In the ship detection task, fusing information from different scales can help the model better understand the location, shape, and background differences of ships in remote sensing images.

[0113] Step 2.3: Next, use the new feature map Get the channel attention map and the supplementary channel attention map.

[0114] In this process, the channel attention map can be implemented in three different ways.

[0115] The first way is to get the feature map Input into the global average pooling layer, fully connected layer, batch normalization layer, and ReLU activation function layer in sequence to obtain the channel attention map A α .

[0116] In this way, the feature map First, a global average pooling operation is performed, and then the fully connected layer (FC) is input. Then, the feature distribution is adjusted through batch normalization, and finally the nonlinear characteristics are added through the ReLU activation function to obtain the channel attention map A. α .

[0117] The global average pooling operation compresses the feature map to a 1×1 size, aggregating global information from the entire image and reducing the impact of local noise. Batch normalization improves the stability of image set training and accelerates convergence. The ReLU activation function effectively introduces nonlinear features and suppresses negative values, preventing the vanishing gradient problem. These methods can improve network performance without significantly increasing the amount of computation.

[0118] The second way is to get the feature map Input into the global average pooling layer, the fully connected layer, and the ReLU activation function layer in sequence to obtain the channel attention map A α The network structure diagram in this way can be referred to Figure 3 .

[0119] Compared to the first approach, this method removes the batch normalization layer and directly feeds the output of the fully connected layer into the ReLU activation function. This approach works well in some cases, particularly when the network is small or the input data features are already well-distributed, reducing computational overhead. It is suitable for remote sensing ship image datasets with a well-defined initial weight distribution or when the amount of training data is small.

[0120] The third way is to get the feature map Input into a maximum pooling layer, a fully connected layer and a Sigmoid activation function layer in sequence to obtain the channel attention map A α , this method can be expressed by the following formula:

[0121] A α =Sigmoid(FC(MaxPool(X G )));

[0122] Among them, MaxPool represents the maximum pooling layer, FC represents the fully connected layer, and Sigmoid represents the Sigmoid activation function layer.

[0123] Then the feature map X G Perform a global average pooling operation to obtain the intermediate feature map. The global average pooling operation compresses the feature map to a size of 1×1, which can aggregate the global information of the entire image and thus reduce the impact of local noise.

[0124] The third approach first downsamples the feature map using a max pooling layer to reduce computational effort and feature map size. A fully connected (FC) operation is then used to generate a channel-wise attention map, which is then output using a sigmoid activation function. In this approach, the max pooling operation effectively extracts the most important features, reduces the amount of information, suppresses noise, and helps the network focus on global features. The sigmoid activation function outputs values ​​ranging from 0 to 1, making it ideal for representing attention weights because it expresses the "importance" of each channel. In summary, this approach reduces the dimensionality of the feature map through pooling, reducing computational burden and making it suitable for processing larger input remote sensing ship images. The probability values ​​output by the sigmoid output allow the channel-wise attention map to be directly used to weight the input feature map, resulting in high interpretability.

[0125] The channel attention map A obtained by the above method α ∈[0, 1] c×1×1 Indicates the importance of each channel. In addition, it is necessary to calculate the supplementary channel attention map A based on the channel attention map β , whose value is 1-A α .

[0126] Different ships have different material and color characteristics. The attention mechanism helps the network dynamically adjust its focus on different ship features based on the image content. By calculating the supplementary channel attention map, the model can adaptively adjust the weights in different feature maps. This reverse attention mechanism helps reduce the influence of irrelevant information and improves the model's performance in complex scenarios.

[0127] Step 2.4: Use dual attention to calibrate features. Use channel attention map A α And supplement channel attention map A β X α and X β Perform adaptive calibration to obtain a new feature map X′ α and X′ β :

[0128]

[0129] in, Represents element-wise multiplication. By calibrating the feature map through the dual attention mechanism, the model can focus more on the key areas of the ship, which effectively reduces the influence of background noise.

[0130] Step 2.5: Feature map fusion and output. Finally, X′ α and X′ β The fusion is performed and input into the convolutional layer to obtain a new fusion feature map, which is recorded as the second fusion feature map X′:

[0131]

[0132] in, By fusing feature maps after different attention calibrations, the model can make full use of multi-dimensional information and enhance the recognition ability of different types of ships.

[0133] Step 2.6: Further processing and extraction of initial features. Initial features X O Input to a convolution layer, then pass through the batch normalization layer, and use the activation function (such as reul activation function) for nonlinear transformation to obtain the intermediate feature map, and then pass through another convolution layer to obtain X' OThis operation further extracts key patterns in the feature map and enhances feature expression. The convolutional layer can capture local structural information in the image, and batch normalization helps reduce internal covariate shift and improve training stability.

[0134] Step 2.7: Generate the fused and enhanced features. First, generate X′ generated in step 2.6 O Add it to the feature map X′ generated in step 2.5. Then, pass it through an activation function to get the final feature

[0135] Among them, ReLU represents the ReLU activation function.

[0136] Step 3: Use the semantic segmentation result acquisition module to extract features again on the fusion enhanced feature map obtained in the previous step, obtain the feature representation for semantic segmentation, and obtain the semantic segmentation result based on the feature representation of semantic segmentation.

[0137] The overall structure of the semantic segmentation result acquisition module in the present invention is as follows: Figure 4 As shown in the figure, the fully connected layer first flattens the enhanced feature maps and applies a nonlinear activation function to extract the feature representations used for semantic segmentation. These features are then fed into the semantic segmentation layer, which uses a softmax activation function to output the probability distribution for each class. Finally, a composite loss function is constructed by combining cross-entropy loss and focal loss. This ensures that the model not only accurately segments the multi-class semantic segmentation task but also effectively handles class imbalance.

[0138] Specifically, the semantic segmentation result acquisition module of the present invention performs the following steps:

[0139] Step 3.1: First, the enhanced feature map generated in step 2.7 To flatten the image, it is fed into the flattening layer (composed of multiple fully connected layers). These fully connected layers map the high-dimensional feature maps to a low-dimensional space, extracting the most informative features required for ship semantic segmentation. Next, an activation function (such as ReLU) is applied to the outputs of the multiple fully connected layers to introduce nonlinear transformations, resulting in Y.

[0140] Step 3.2: The feature map Y is input to a semantic segmentation layer (usually a fully connected layer) whose output dimension is equal to the number of categories of the semantic segmentation task. For ship detection, a softmax activation function can be used for multi-semantic segmentation:

[0141]

[0142] Among them, w cls and b clsare the weight matrix and bias of the semantic segmentation layer, Is the final output of the model. For multi-class semantic segmentation tasks, softmax is usually the Softmax activation function to obtain the probability distribution of each class.

[0143] The composite loss function of the semantic segmentation result acquisition module is constructed through the following steps.

[0144] First, for the problem of multi-category semantic segmentation of remote sensing images, the common loss function is cross entropy loss, which is formulated as:

[0145]

[0146] Where C is the number of categories, y i is the true label (usually one-hot encoding). Next, the focal loss is applied to address the class imbalance problem. It introduces a regulation factor based on the standard cross entropy loss, making the model pay more attention to samples that are difficult to semantically segment. The formula of focal loss is:

[0147]

[0148] Among them, α i Is the balance factor for each category (helps to deal with the problem of category imbalance), usually set to a constant, γ is the adjustment factor used to adjust the weight of easy semantic segmentation and difficult semantic segmentation samples, usually choose γ∈[0,5], a larger γ will increase the attention to difficult semantic segmentation samples. Add Dice loss for semantic segmentation Finally, construct the composite loss function:

[0149]

[0150] Where λ is a hyperparameter that controls the weight of the focal loss. The cross-entropy loss, Dice loss, and focal loss work together to ensure that the model not only performs correct semantic segmentation but also improves performance in the presence of class imbalance.

[0151] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0152] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module, characterized in that: The following steps are involved: Acquire a remote sensing ship image dataset, perform preliminary feature extraction on the remote sensing ship images in the remote sensing ship image dataset, and obtain initial features of the remote sensing ship images; The feature hyper-fusion module is used to fuse and enhance the initial features of the remote sensing ship image to obtain the fusion enhanced feature map of the remote sensing ship image; specifically, it includes: The initial features are sequentially input into the 5*5 convolution layer, the first batch normalization layer and the first activation function layer in the feature super fusion module to obtain the first feature map The initial features are sequentially input into the dilated convolution layer, the second batch normalization layer, and the second activation function layer in the feature super-fusion module to obtain the second feature map. The activation functions in the first activation function layer and the second activation function layer are both LeakyReLu activation functions; c is the number of channels, h and w are the height and width of the feature map respectively; Based on element addition, the first feature map X α and the second feature map X β Fusion is performed to obtain the first fusion feature map in Represents element addition operation; The first fusion feature map X G Input the maximum pooling layer and the fully connected layer in sequence, and perform Sigmoid function activation operation on the output of the fully connected layer to obtain the channel attention map A α And supplement channel attention map A β , the formula is: A α =Sigmoid(FC(MaxPool(X G ))); A β =1-A α ; Among them, A α ∈[0, 1] c×1×1 Indicates the importance of each channel; Using channel attention map A α For the first feature map X α Perform adaptive correction to obtain the first correction feature map X′ α : Using the supplementary channel attention map A β For the second feature map X β Perform adaptive correction to obtain the second correction feature map X′ β : in, Represents element-wise multiplication operation; Based on the convolution operation, the first corrected feature map X′ α and the second correction characteristic map X′ β The fusion is performed to obtain the second fusion feature map X′: The initial features are sequentially input into a convolutional layer, the third batch normalization layer, and the activation function is used for nonlinear transformation to obtain the intermediate feature map, and then the intermediate feature map is input into a convolutional layer again to obtain the enhanced feature map X′ O ; The second fusion feature map and the enhanced feature map X′ O Perform feature fusion and use the ReLU activation function to activate and obtain the final fusion enhanced feature map Among them, ReLU represents the ReLU activation function; Feature extraction is performed on the fusion enhanced feature map of the remote sensing ship image to obtain the feature representation for semantic segmentation, and the semantic segmentation result is obtained based on the feature representation of semantic segmentation.

2. The method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module according to claim 1, characterized in that: Perform preliminary feature extraction on the remote sensing ship images in the remote sensing ship image dataset to obtain the initial features of the remote sensing ship images, specifically including the following steps: S11. preprocessing the remote sensing ship images in the remote sensing ship image dataset; S12. Utilize an initial feature extraction module composed of multiple convolution units to extract preliminary features from the preprocessed remote sensing ship image.

3. The method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module according to claim 2, characterized in that: In step S12, the initial feature extraction module composed of multiple convolution units includes an initial convolution unit, a deep convolution unit, and a skip connection unit; the initial convolution unit and the deep convolution unit each include a convolution layer, a nonlinear activation function layer, a batch normalization layer, and a pooling layer connected in sequence; The initial convolution unit is used to extract shallow features from the pre-processed remote sensing ship image; The deep convolution unit is used to extract the shallow features again to obtain the deep features of the remote sensing ship image; The jump connection unit is used to combine the shallow features output by the initial convolution unit with the deep features output by the deep convolution unit to obtain the final initial features of the remote sensing ship image.

4. The method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module according to claim 1, characterized in that: Feature extraction is performed on the fusion enhanced feature map of the remote sensing ship image to obtain feature representation for semantic segmentation, and semantic segmentation results are obtained based on the feature representation of semantic segmentation. Specifically, the following steps are included: Inputting the fused enhanced feature map into a flattening layer composed of multiple fully connected layers, using the flattening layer to map the fused enhanced feature map to a low-dimensional space, and extracting feature representation for semantic segmentation; The feature representation obtained for semantic segmentation is nonlinearly transformed using an activation function to obtain a nonlinear feature map Y; The nonlinear feature map Y is input into the semantic segmentation layer to obtain the semantic segmentation result.

5. The method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module according to claim 4, characterized in that: The nonlinear feature map Y is input into the semantic segmentation layer to obtain the semantic segmentation result, which is expressed by the following formula: Among them, w cls and b cls Represent the weight matrix and bias of the semantic segmentation layer, Represents the probability distribution of each class, and softmax represents the Softmax activation function.

6. The method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module according to claim 1, characterized in that: The model training process for obtaining semantic segmentation results includes the following loss functions: Among them, λ is a hyperparameter, Lcls represents the cross entropy loss, represents Dice loss, and Lfl represents focal loss.

7. The method for semantic segmentation of ships in remote sensing images based on a feature hyper-fusion module according to claim 6, characterized in that: The focus loss specifically includes the following formula: Among them, α i is the balancing factor for each category, and γ is the adjustment factor.

8. A remote sensing image ship semantic segmentation system based on feature hyper-fusion module, characterized by: It includes the initial feature extraction module, feature super-fusion module and semantic segmentation result acquisition module; The initial feature extraction module is used to obtain a remote sensing ship image dataset and perform preliminary feature extraction on the remote sensing ship images in the remote sensing ship image dataset to obtain initial features of the remote sensing ship images; The feature super-fusion module is used to fuse and enhance the initial features to obtain fusion-enhanced features; specifically, it includes the following steps: The initial features are sequentially input into the 5*5 convolution layer, the first batch normalization layer and the first activation function layer in the feature super fusion module to obtain the first feature map The initial features are sequentially input into the dilated convolution layer, the second batch normalization layer, and the second activation function layer in the feature super-fusion module to obtain the second feature map. The activation functions in the first activation function layer and the second activation function layer are both LeakyReLu activation functions; c is the number of channels, h and w are the height and width of the feature map respectively; Based on element addition, the first feature map X α and the second feature map X β Fusion is performed to obtain the first fusion feature map in Represents element addition operation; The first fusion feature map X G Input the maximum pooling layer and the fully connected layer in sequence, and perform Sigmoid function activation operation on the output of the fully connected layer to obtain the channel attention map A α And supplement channel attention map A β , the formula is: A α =Sigmoid(FC(MaxPool(X G ))); A β =1-A α ; Among them, A α ∈[0, 1] c×1×1 Indicates the importance of each channel; Using channel attention map A α For the first feature map X α Perform adaptive correction to obtain the first correction feature map X′ α : Using the supplementary channel attention map A β For the second feature map X β Perform adaptive correction to obtain the second correction feature map X′ β : in, Represents element-wise multiplication operation; Based on the convolution operation, the first corrected feature map X′ α and the second correction characteristic map X′ β The fusion is performed to obtain the second fusion feature map X′: The initial features are sequentially input into a convolutional layer, the third batch normalization layer, and the activation function is used for nonlinear transformation to obtain the intermediate feature map, and then the intermediate feature map is input into a convolutional layer again to obtain the enhanced feature map X′ O ; The second fusion feature map and the enhanced feature map X′ O Perform feature fusion and use the ReLU activation function to activate and obtain the final fusion enhanced feature map Among them, ReLU represents the ReLU activation function; The semantic segmentation result acquisition module is used to extract features from the fused enhanced feature map, obtain feature representation for semantic segmentation, and obtain a semantic segmentation result based on the feature representation of semantic segmentation.

Citation Information

Patent Citations

  • Medical image segmentation method based on multi-scale cross-layer attention fusion network

    CN117152433A

  • Multi-scale deep supervision based reverse attention model

    US20220292394A1