SegFormer-based seaborne image semantic segmentation method

Through the SegFormer network architecture and ECA channel attention mechanism, combined with the mixed loss function, the problem of insufficient recognition of small and medium-sized objects in marine images is solved, and a high-precision semantic segmentation effect is achieved.

CN120339628APending Publication Date: 2025-07-18DALIAN MARITIME UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510739106.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art lacks the ability to identify and segment small objects in offshore images, making it difficult to accurately capture their boundaries and detailed information, affecting the segmentation accuracy.

Method used

The SegFormer network architecture is adopted, combining the ECA channel attention mechanism and mixed loss function to establish a semantic segmentation model of marine images, and feature extraction and classification are performed through the embedding layer, Transformer module and multi-layer perceptron layer, self-attention calculation and pixel comparison loss function are introduced to enhance the recognition ability of small-size target features.

Benefits of technology

It improves the stability and accuracy of semantic segmentation of marine images, and can effectively identify and segment small targets in different environments to meet the needs of navigation applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339628A_ABST
    Figure CN120339628A_ABST
Patent Text Reader

Abstract

The invention discloses a seaborne image semantic segmentation method based on SegFormer. The seaborne image semantic segmentation method comprises the following steps: S1, acquiring a seaborne semantic segmentation data set; s2, establishing a seaborne image semantic segmentation model based on a SegFormer network architecture, training the seaborne image semantic segmentation model based on the seaborne semantic segmentation data set, setting a mixed loss function as a loss function in a training process, and when the loss function converges, obtaining the trained seaborne image semantic segmentation model; and S3, based on the trained seaborne image semantic segmentation model, performing seaborne image semantic segmentation processing on a to-be-identified image. According to the sea image semantic segmentation model, more attention is paid to important channels for extracting object edge features and small-size target features, the problem of serious imbalance of categories is avoided, the stability and accuracy of sea image semantic segmentation are finally improved, and different environment and application requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method for semantic segmentation of maritime images based on SegFormer. Background Art

[0002] In recent years, major shipping countries have actively promoted the research on ship intelligence, aiming to improve the safety, environmental protection, economy and reliability of ship operation. In 2024, the "Intelligent Ship Specification" published by the China Classification Society further clarified the four basic capabilities that intelligent ships need to possess, including perception, memory and thinking, learning and adaptation, and behavior decision-making. Among them, endowing ships with thinking ability is the key to the full autonomy of intelligent ships, which also marks the deepening of perception ability. In this process, image description technology, as an important part, plays an indispensable role. By accurately understanding and describing the environment and surrounding objects, image description technology not only improves the perception ability of ships, but also provides a basis for intelligent decision-making. Therefore, image description has a crucial position and significance in the development of intelligent ships. Deep learning has a multi-layer network structure and strong fitting ability. Through training and learning with a large amount of data, it can extract the abstract features and semantic information of objects to the greatest extent. Compared with traditional algorithms, its accuracy has been effectively improved. With the continuous development of computer hardware and the rise of deep learning, in the field of computer vision, the characteristics of strong information capture and learning ability of deep learning are used to achieve semantic segmentation of images.

[0003] There are a large number of small targets in the maritime scene, and these small targets are crucial for navigation safety. In maritime images, small targets often only account for a very small proportion of the whole image. In the existing technology, during the feature extraction process, it often tends to focus more on the significant features or large areas in the image, and has poor ability in the recognition and segmentation of small targets. It often fails to accurately identify and segment these small targets. This bias makes it difficult for existing models to accurately capture their boundaries and detail information when dealing with small targets, thereby affecting the segmentation accuracy. Summary of the Invention

[0004] The present invention provides a method for semantic segmentation of maritime images based on SegFormer to overcome the technical problem of low recognition accuracy of small-size target features in the existing semantic segmentation methods.

[0005] To achieve the above object, the technical solution of the present invention is:

[0006] A method for semantic segmentation of maritime images based on SegFormer, the specific steps include:

[0007] S1: Obtain a maritime semantic segmentation dataset;

[0008] S2: Establish a maritime image semantic segmentation model based on the SegFormer network architecture, train the maritime image semantic segmentation model based on the maritime semantic segmentation dataset, and set a hybrid loss function as the loss function during the training process. When the loss function converges, obtain the trained maritime image semantic segmentation model;

[0009] The maritime image semantic segmentation model includes an encoder and a decoder, where:

[0010] The encoder includes an embedding layer and a Transformer module established based on the ECA channel attention mechanism;

[0011] The embedding layer is used to perform embedding processing on the maritime semantic segmentation dataset to obtain an embedded feature vector;

[0012] The Transformer module established based on the ECA channel attention mechanism is used to perform self-attention calculation on the embedded feature vector to obtain several feature maps with different resolutions;

[0013] The decoder includes a multi-layer perceptron layer and a multi-layer perceptron module;

[0014] The multi-layer perceptron layer is used to perform linear transformation and non-linear activation processing on several feature maps with different resolutions to unify the dimensions of the feature maps with different resolutions;

[0015] The multi-layer perceptron module is used to perform classification processing on the feature maps with unified dimensions to obtain the semantic segmentation result;

[0016] S3: Based on the trained maritime image semantic segmentation model, perform maritime image semantic segmentation processing on the image to be recognized.

[0017] Further, the Transformer module established based on the ECA channel attention mechanism includes a first Transformer unit, a second Transformer unit, a third Transformer unit, and a fourth Transformer unit;

[0018] The first Transformer unit is used to perform self-attention calculation on the embedded feature vector to obtain a first feature map with a quarter resolution, and transmit it to the second Transformer unit and the decoder;

[0019] The second Transformer unit is used to perform self-attention calculation processing on the feature map with a quarter resolution to obtain a feature map with an eighth resolution, and transmit it to the third Transformer unit and the decoder;

[0020] The third Transformer unit is used to perform self-attention calculation processing on the feature map with an eighth resolution to obtain a feature map with a sixteenth resolution, and transmit it to the fourth Transformer unit and the decoder;

[0021] The fourth Transformer unit is used to perform self-attention calculation processing on the feature map with a sixteenth resolution to obtain a feature map with a thirty-second resolution, and transmit it to the decoder.

[0022] Further, each Transformer unit includes a Transformer Block and an ECA module;

[0023] The Transformer Block is used to perform multi-level local feature extraction processing on the input feature map and transmit it to the ECA module;

[0024] The ECA module is used to perform cross-channel interaction processing on the multi-level local features output by the Transformer Block.

[0025] Further, the hybrid loss function is expressed as:

[0026]

[0027] In the formula, is the cross-entropy loss function; is the pixel contrast loss function.

[0028] Further, the cross-entropy loss function is expressed as:

[0029]

[0030] The pixel contrast loss function is expressed as:

[0031]

[0032] where i represents the normalized feature vector of each pixel, P i and N i respectively represent the positive and negative sample sets of i, τ is a hyperparameter, i+ is the positive sample feature vector, and i - is the negative sample feature vector; y c is the score output by the marine image semantic segmentation model for the correct class c; y c′ is the score output by the marine image semantic segmentation model for all possible classes c'.

[0033] Further, the ECA module includes a global average pooling layer, a one-dimensional convolutional layer, a Sigmoid activation function, and an output layer;

[0034] The global average pooling layer is used to perform global average pooling operation on the input feature map;

[0035] The one-dimensional convolutional layer is used to perform one-dimensional convolution operation on the output data of the global average pooling layer;

[0036] The Sigmoid activation function is used to calculate the weights of each channel in the output data of the one-dimensional convolutional layer;

[0037] The output layer is used to multiply the weights output by the Sigmoid activation function by the feature map input to the global average pooling layer to obtain the output feature map.

[0038] Beneficial effects: For the actual navigation scenario, the present invention establishes a maritime image semantic segmentation model based on the SegFormer network architecture and introduces the ECA channel attention mechanism to enhance the recognition ability and perception ability of the maritime image semantic segmentation model for small-size target features, alleviate the feature semantic loss, so as to realize the semantic segmentation of pixel-level navigation images in the way of multi-feature cascade; meanwhile, a hybrid loss function is set as the loss function in the training process, which can make the maritime image semantic segmentation model pay more attention to the important channels for extracting object edge features and small-size target features, avoid the problem of serious class imbalance, and finally improve the stability and accuracy of the maritime image semantic segmentation to meet different environmental and application requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0040] Figure 1 It is a flowchart of a maritime image semantic segmentation method based on SegFormer in the present invention;

[0041] Figure 2 It is a schematic structural diagram of the maritime image semantic segmentation model in the embodiment of the present invention;

[0042] Figure 3 It is a schematic structural diagram of the ECA mechanism in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] This embodiment provides a method for semantic segmentation of maritime images based on SegFormer. As Figure 1 shown, the specific steps include:

[0045] S1: Obtain a maritime semantic segmentation dataset;

[0046] Specifically, in this embodiment, a real ship navigation video is obtained, image data is collected frame by frame based on the real ship navigation video, and pictures with qualified clarity are selected. The data is enriched by adjusting cropping, sharpening, contrast, and brightness. Then, the obtained pictures are labeled for semantic segmentation (classifying each pixel point into one of the following categories: ocean, sky, ship, dock, buoy, city, island, and background, a total of eight categories) to construct a maritime semantic segmentation dataset;

[0047] S2: Establish a maritime image semantic segmentation model based on the SegFormer network architecture, train the maritime image semantic segmentation model based on the maritime semantic segmentation dataset, and set a hybrid loss function as the loss function during the training process. When the loss function converges, a trained maritime image semantic segmentation model is obtained;

[0048] As Figure 2 shown, the maritime image semantic segmentation model includes an encoder and a decoder, where:

[0049] The encoder includes an embedding layer and a Transformer module established based on the ECA channel attention mechanism;

[0050] The embedding layer (overlap patch embeddings) is used to perform embedding processing on the maritime semantic segmentation dataset to obtain an embedded feature vector;

[0051] The Transformer module established based on the ECA channel attention mechanism is used to perform self-attention calculation on the embedded feature vector to obtain several feature maps with different resolutions;

[0052] The decoder includes a multi-layer perceptron layer and a multi-layer perceptron module;

[0053] The multi-layer perceptron layer (MLP Layer) is used to perform linear transformation and non-linear activation processing on feature maps of several different resolutions to unify the dimensions of feature maps with different resolutions;

[0054] The multi-layer perceptron module (MLP) is used to classify feature maps with unified dimensions (size is N cls the total number of categories to be distinguished) to obtain the semantic segmentation result;

[0055] S3: Based on the trained marine image semantic segmentation model, perform marine image semantic segmentation processing on the image to be recognized.

[0056] In a specific embodiment, the Transformer module based on the ECA channel attention mechanism includes a first Transformer unit, a second Transformer unit, a third Transformer unit, and a fourth Transformer unit;

[0057] The first Transformer unit is used to perform self-attention calculation on the embedded feature vector to obtain a first feature map with a quarter resolution (size is C1 is the number of channels), and transmit it to the second Transformer unit and the decoder;

[0058] The second Transformer unit is used to perform self-attention calculation processing on the feature map with a quarter resolution to obtain a feature map with an eighth resolution (size is C2 is the number of channels), and transmit it to the third Transformer unit and the decoder;

[0059] The third Transformer unit is used to perform self-attention calculation processing on the feature map with an eighth resolution to obtain a feature map with a sixteenth resolution (size is C3 is the number of channels), and transmit it to the fourth Transformer unit and the decoder;

[0060] The fourth Transformer unit is used to perform self-attention calculation processing on the feature map with a sixteenth resolution to obtain a feature map with a thirty-second resolution (size is C4 is the number of channels), and transmit it to the decoder.

[0061] In a specific embodiment, each Transformer unit includes a Transformer Block and an ECA module;

[0062] The Transformer Block is used to perform multi-level local feature extraction processing on the input feature map and transmit it to the ECA module;

[0063] The ECA module is used to perform cross-channel interaction processing on the multi-level local features output by the Transformer Block.

[0064] Specifically, for the problems in nautical images such as monotonous background, but being affected by factors such as jitter, water droplets, weather, etc., resulting in missing local information, etc., this embodiment introduces an Figure 3 EfficientChannelAttention (ECA) module as shown. ECA is a lightweight channel attention mechanism that can avoid dimensionality reduction and efficiently realizes local cross-channel interaction with one-dimensional convolution, extracting the dependency relationship between channels. By adding the ECA module after the Transformer Block, the maritime image semantic segmentation model can more accurately obtain the relationship between feature channels and suppress unimportant channel features, so as to obtain more representative high-level semantic features and improve the expression ability of features.

[0065] In a specific embodiment, as Figure 3 shown, the ECA module includes a global average pooling layer GAP, a one-dimensional convolutional layer, a Sigmoid activation function, and an output layer;

[0066] The global average pooling layer is used to perform global average pooling operation on the input feature map;

[0067] The one-dimensional convolutional layer is used to perform one-dimensional convolution operation on the output data of the global average pooling layer;

[0068] The Sigmoid activation function is used to calculate the weights of each channel in the output data of the one-dimensional convolutional layer;

[0069] The output layer is used to multiply the weights output by the Sigmoid activation function by the feature map input to the global average pooling layer to obtain the output feature map.

[0070] In a specific embodiment, the hybrid loss function is expressed as:

[0071]

[0072] wherein, is the cross-entropy loss function; is the pixel contrast loss function.

[0073] In a specific embodiment, the cross-entropy loss function is expressed as:

[0074]

[0075] A pixel contrast loss function branch is added to the decoder part. This branch is only used during the training process and contains two 1×1 convolutional layers and one ReLU layer. The pixel contrast loss function is expressed as:

[0076]

[0077] where i represents the normalized feature vector of each pixel, P i and N i respectively represent the positive and negative sample sets of i, τ is a hyperparameter, i+ is the positive sample feature vector, and i - is the negative sample feature vector; y c is the score output by the marine image semantic segmentation model for the correct class c; y c′ is the score output by the marine image semantic segmentation model for all possible classes c'.

[0078] Specifically, in nautical images, some relatively key elements account for a very small proportion of the entire image, resulting in severe class imbalance. In response to this situation, in this embodiment, a pixel contrast loss function is introduced during the training process as an auxiliary function to optimize the problem of class imbalance.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A semantic segmentation method for maritime images based on SegFormer, characterized in that, The specific steps include: S1: Obtain a maritime semantic segmentation dataset; S2: Establish a maritime image semantic segmentation model based on the SegFormer network architecture, train the maritime image semantic segmentation model based on the maritime semantic segmentation dataset, and set a hybrid loss function as the loss function during the training process. When the loss function converges, obtain the trained maritime image semantic segmentation model; The maritime image semantic segmentation model includes an encoder and a decoder, where: The encoder includes an embedding layer and a Transformer module established based on the ECA channel attention mechanism; The embedding layer is used to perform embedding processing on the maritime semantic segmentation dataset to obtain an embedded feature vector; The Transformer module established based on the ECA channel attention mechanism is used to perform self-attention calculation on the embedded feature vector to obtain several feature maps with different resolutions; The decoder includes a multi-layer perceptron layer and a multi-layer perceptron module; The multi-layer perceptron layer is used to perform linear transformation and non-linear activation processing on several feature maps with different resolutions to unify the dimensions of the feature maps with different resolutions; The multi-layer perceptron module is used to perform classification processing on the feature maps with unified dimensions to obtain a semantic segmentation result; S3: Based on the trained maritime image semantic segmentation model, perform maritime image semantic segmentation processing on the image to be recognized.

2. The method for semantic segmentation of maritime images based on SegFormer according to claim 1, characterized in that, The Transformer module established based on the ECA channel attention mechanism includes a first Transformer unit, a second Transformer unit, a third Transformer unit, and a fourth Transformer unit; The first Transformer unit is used to perform self-attention calculation on the embedded feature vector to obtain a first feature map with a quarter resolution, and transmit it to the second Transformer unit and the decoder; The second Transformer unit is used to perform self-attention calculation processing on the feature map with a quarter resolution to obtain a feature map with an eighth resolution, and transmit it to the third Transformer unit and the decoder; The third Transformer unit is used to perform self-attention calculation processing on the feature map with an eighth resolution to obtain a feature map with a sixteenth resolution, and transmit it to the fourth Transformer unit and the decoder; The fourth Transformer unit is used to perform self-attention calculation processing on the feature map with a sixteenth resolution to obtain a feature map with a thirty-second resolution, and transmit it to the decoder.

3. The method for semantic segmentation of maritime images based on SegFormer according to claim 2, characterized in that, Each Transformer unit includes a Transformer Block and an ECA module; The Transformer Block is used to perform multi-level local feature extraction processing on the input feature map and transmit it to the ECA module; The ECA module is used to perform inter-channel interaction processing on the multi-level local features output by the Transformer Block.

4. The method for semantic segmentation of marine images based on SegFormer according to claim 3, characterized in that, The hybrid loss function is expressed as: In the formula, is the cross-entropy loss function; is the pixel contrast loss function.

5. The method for semantic segmentation of maritime images based on SegFormer according to claim 4, wherein The cross-entropy loss function is expressed as: The pixel contrast loss function is expressed as: Among them, i represents the normalized feature vector of each pixel, P i and N i respectively represent the positive and negative sample sets of i, τ is a hyperparameter, i+ is the positive sample feature vector, and i- is the negative sample feature vector; y c is the score output by the maritime image semantic segmentation model for the correct class c; y c′ is the output score of the maritime image semantic segmentation model for all possible classes c'.

6. The method for semantic segmentation of marine images based on SegFormer according to claim 5, wherein The ECA module includes a global average pooling layer, a one-dimensional convolutional layer, a Sigmoid activation function, and an output layer; The global average pooling layer is used to perform global average pooling operations on the input feature map; The one-dimensional convolutional layer is used to perform one-dimensional convolutional operations on the output data of the global average pooling layer; The Sigmoid activation function is used to calculate the weights of each channel in the output data of the one-dimensional convolutional layer; The output layer is used to multiply the weights output by the Sigmoid activation function by the feature map input to the global average pooling layer to obtain the output feature map.