Marine animal segmentation method based on deep multi-level feature aggregation
By constructing a deep multi-level feature aggregation method, combining the frozen SAM encoder, adapter, super feature map extraction module and progressive prediction decoder, the problem of inaccurate segmentation of marine animals in complex underwater environments is solved, and high-performance segmentation of marine animals is achieved.
Patent Information
- Application Number
- CN202510405898.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively capture the details and structural information of marine animals in complex underwater environments, resulting in inaccurate segmentation of marine animals, especially when image quality declines under light scattering and absorption effects.
A marine animal segmentation method based on deep multi-level feature aggregation is constructed. By freezing the pre-trained SAM encoder and introducing an adapter, combining the super feature map extraction module and an incremental prediction decoder, the multi-level feature map is fused with the fusion attention module to realize information extraction of global context and local details.
High-performance marine animal segmentation is achieved in complex underwater scenarios, improving the accuracy and stability of segmentation, and being able to accurately segment marine animals in various variable marine scenarios.
Smart Images

Figure CN120340066A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision, artificial intelligence, and underwater robots, and relates to visual self-attention, fully supervised training, and visual foundation models. Specifically, it is a marine animal segmentation method based on deep multi-level feature aggregation. Background Art
[0002] Marine Animal Segmentation (MAS) is a key and fundamental task in the fields of visual intelligence and underwater robots. Its goal is to identify and segment marine animals from underwater images or videos. Functionally, accurate marine animal segmentation is crucial for multiple research fields, including marine biology, ecology, and environmental protection. However, MAS faces unique challenges compared to typical land image segmentation. In fact, the underwater environment faces complex light scattering and absorption effects, resulting in degraded image quality, reduced contrast, and blurred objects. In addition, marine animals often have camouflage characteristics, making the segmentation task of marine organisms more complex. To address these challenges, advanced sensing technologies are required.
[0003] With the advancement of deep learning technology, Convolution Neural Network (CNN) has become a mainstream choice for solving the MAS task. For example, some researchers are committed to designing efficient deep convolutional neural networks for extracting the structural texture and context clues of underwater images. However, these CNN-based methods are greatly limited in their performance due to the lack of global understanding ability. At the same time, Vision Transformer has shown strong capabilities in capturing long-range dependencies and global context information and has achieved better performance in general image segmentation tasks. Therefore, some researchers have combined the advantages of CNN and Transformer and proposed a network model coupling CNN and Transformer for underwater salient object segmentation. However, these Transformer-based methods require a large amount of training data to achieve satisfactory results. Different from previous work, the method of the present invention uses a visual foundation model to construct a deep multi-level feature aggregation method and has achieved better MAS results in complex underwater scenes.
[0004] Recently, the Segment Anything Model (SAM) has been proposed and demonstrated great potential in general image segmentation tasks. However, the training data of SAM are mainly natural images, which limits its performance in other complex scenarios and it is difficult to achieve excellent performance in specific tasks. At the same time, due to the simplification of the model decoder, SAM only utilizes the end features of the vision Transformer, which may lead to the loss of details of complex objects, thus affecting accurate object segmentation. To address the above drawbacks, some research works inject domain-specific cues into SAM by using simple adapters. In addition, some researchers have verified that complex adapters can highlight task-specific information. However, the exploration of the application of SAM in MAS is still relatively limited. Some researchers directly fine-tune SAM for underwater image segmentation. Others implement MAS by adjusting the structure of SAM and using adaptive prompts. But these research works cannot capture the details and structural information of complex marine animals. For this reason, the present invention improves SAM and uses deep multi-level feature aggregation to achieve high-performance MAS. Summary of the Invention
[0005] The present invention constructs a method for segmenting marine animals based on deep multi-level feature aggregation, named MAS-SAM. By freezing the pre-trained parameters of the SAM encoder and introducing an effective adapter, the present invention constructs an Adapter-informed SAM Encoder (ASE) for injecting marine domain knowledge to extract unique features of marine animal images. In addition, the present invention also constructs a Hypermap Extraction Module (HEM) for extracting multi-level feature maps from ASE. This module provides comprehensive guidance for the subsequent mask prediction process. To improve the decoder of SAM, the present invention introduces a Progressive Prediction Decoder (PPD) to aggregate the features in the original prompt, ASE, and HEM. By combining with the Fusion Attention Module (FAM), the method in the present invention can make full use of the multi-level feature maps and extract richer marine animal-related information from global context to local details, so as to accurately segment marine animals in various complex and changeable marine scenarios.
[0006] Technical solution of the present invention:
[0007] A method for segmenting marine animals based on deep multi-level feature aggregation, the steps are as follows:
[0008] Step 1: Obtain animal pictures in the ocean scene as training and test data, and accurate segmentation mask annotation is required for the data;
[0009] Step 2: Retain the feature extraction backbone of SAM, and inject Low Rank Adaptation (LoRA) and Adapter into the Multi-Head Self Attention (MHSA) and Feed Forward Network (FFN) of each Transformer module in SAM respectively; the feature extractor after the change is called ASE;
[0010] Let X i ∈R N×D be the input of the i-th Transformer module, where N represents the number of visual tokens and D represents the embedding dimension; the multi-head self-attention mechanism modified by LoRA is expressed as the following formula:
[0011]
[0012] K i =W k (X i )
[0013]
[0014] where W q ,W k ,W v are the weights of the three linear projection layers for generating the original Query, Key, and Value matrices respectively; Q i ,K i ,V i are the inputs of the multi-head self-attention respectively; is the weight of the compressed linear projection layer in LoRA, is the weight of the upsampling linear projection layer in LoRA, used to reduce and restore the feature dimension respectively; M is the dimension of the compressed projection; is the intermediate state generated by the multi-head self-attention mechanism; in this way, the present invention can freeze the pre-trained weights (W q ,W k ,W v ) and greatly reduce the number of trainable parameters by using low-rank matrix decomposition.
[0015] In addition, an adapter is inserted into the feed forward network as follows:
[0016]
[0017] Among them, MLP is a multi-layer perceptron; σ is an activation function; is the intermediate state obtained by the feedforward network; are the weights of two linear projection layers, respectively used to reduce and restore the feature dimension; P is the dimension projected in the adapter; The present invention can achieve parameter-efficient fine-tuning and adapt the pre-trained SAM encoder to the marine scenario.
[0018] Due to the complex underwater environment, making full use of local detail information and global context information is crucial for achieving robust and accurate marine animal segmentation. Different Transformer layers can capture semantic information at different levels. Generally, the shallow layers retain more local details, while the deep layers express more context information. Therefore, in order to enable the model constructed by the present invention to utilize richer marine information, the present invention constructs a super feature map extraction module HEM for fusing multi-level feature maps from ASE; The multi-level feature maps extracted by ASE provide comprehensive guidance for the subsequent mask prediction process; First, an image I ∈ R H×W×3 is input into ASE, and the outputs of different Transformer layers are obtained, where H and W are the height and width of the image; Select the 3rd, 6th, 9th, and 12th Transformer layers, and obtain multi-level features, namely X i (i = 3, 6, 9, 12); Then, transform them into spatial feature maps To consider these multiple feature maps simultaneously, perform the following feature aggregation:
[0019] R i = φ 1×1 (F i )
[0020]
[0021] Among them, φ 1×1 and φ 3×3 represent convolutions with a convolution kernel size of 1x1 and 3x3 respectively, is the intermediate state of the first super feature map;
[0022] After that, a channel attention mechanism is introduced to generate the super feature map:
[0023]
[0024] Among them, GAP represents global average pooling, δ is the Sigmoid activation function, H j , (j = 1, 2, 3, 4) is the super feature map obtained after fusion;
[0025] To fully integrate the super feature maps obtained by HEM and the multi-level feature maps obtained by ASE, the present invention constructs a fusion attention module FAM; first, the multi-level feature maps from ASE are upsampled, and the input features are adjusted to the same size; then, they are fused in the following way:
[0026]
[0027] D j+1 = FAM(U i , D j , H j )
[0028] Among them, is the operation of linear interpolation upsampling, and U i is the result obtained after upsampling the features of ASE; D j is the output of the j-th layer in PPD; for FAM, the channel attention mechanism is used to evaluate the importance of features; at the same time, the residual structure is used to enhance the feature representation ability; this process is expressed as:
[0029] Y j = φ 1×1 ([U i , D j , H j )
[0030] W j = δ(φ 1×1 (GAP(Y j )) + φ 1×1 (GMP(Y j )))
[0031]
[0032] Among them, Y j is the intermediate state of fusing U i , D j , H j ; W j is the weight of the feature channel; GMP represents global maximum pooling; finally, the final prediction mask result is obtained through the gradually expanding feature prediction map;
[0033] Step 3: During the training process, three levels of deep supervision are adopted, that is, pixel-level supervision is binary cross-entropy loss, region-level supervision is SSIM loss, and overall-level supervision is IoU loss; therefore, L is defined as the total loss consisting of three parts:
[0034] L = L BCE + L SSIM + L IoU .
[0035] Advantages of the present invention:
[0036] (1) The present invention constructs a marine animal segmentation method based on deep multi-level feature aggregation, named MAS-SAM. It can achieve high-performance marine animal segmentation by customizing SAM and aggregating multi-level features.
[0037] (2) The present invention constructs a super feature map extraction module HEM, which generates multi-level features based on the encoder of SAM and provides comprehensive guidance for subsequent mask prediction.
[0038] (3) The present invention constructs a progressive prediction decoder PPD, which effectively improves the representation ability of the SAM decoder and can capture marine biological information from global context clues to fine-grained local details. Brief Description of the Drawings
[0039] Figure 1 It is the overall model flow chart of the present invention.
[0040] Figure 2 It is the ASE structure framework diagram of the present invention.
[0041] Figure 3 It is the schematic diagram of the attention fusion module FAM of the present invention. Detailed Embodiment
[0042] The following further illustrates the detailed embodiments of the present invention in combination with the drawings and technical solutions. The steps are as follows:
[0043] (1) The model of the present invention is implemented using the PyTorch toolbox and an RTX 3090 GPU. Among them, the encoder of SAM and the prompt encoder are initialized from the pre-trained SAM-B, while other modules are randomly initialized. During training, the encoder of SAM is frozen, and only the remaining modules are fine-tuned. The present invention uses the AdamW optimizer to update the model parameters. The initial learning rate and weight decay are set to 0.001 and 0.1 respectively. The present invention reduces the learning rate to 1 / 10 of the original after every 20 training epochs. The total number of training epochs is set to 50. The batch size is set to 8. The input images are uniformly adjusted to 512×512×3.
[0044] (2) Construct and implement the proposed SAM encoder ASE with marine domain knowledge injection. As Figure 1As shown, the present invention injects a learnable matrix of rank 4 into the MHSA part of the original structure of SAM. The dimension of the original frozen matrix is 1024. Subsequently, the present invention adds an adapter in the FFN part. This adapter compresses the feature map in SAM to 1 / 16 of the original size, and then restores the feature map to the original feature map size after learning the knowledge of the marine field. With very few learnable parameters, the present invention can both retain the strong image understanding ability of SAM and learn the knowledge of the marine field.
[0045] (3) The present invention constructs a super feature map extraction module HEM for utilizing the multi-level feature maps from ASE. The present invention first extracts the feature maps of the 3rd, 6th, 9th, and 12th levels, and then performs feature weighting through feature channel attention. Subsequently, the present invention uses the super feature map as the guiding feature map in the decoder. Specifically, the present invention enlarges the feature map at different scales to cope with the gradually enlarged feature maps in the decoder.
[0046] (4) Finally, the present invention constructs a progressive decoding predictor PPD. The present invention fully fuses and gradually enlarges the super feature map in HEM and the feature map in the encoder ASE. The specific way of fusion is to recalculate the feature map channel weights through the channel attention fusion module FAM. Finally, it is gradually enlarged to the original image size to obtain the final prediction map.
[0047] After true value supervised training, the trained model will be used for the segmentation prediction of marine animals. During the inference process, the present invention understands the marine images of various scenes through ASE to find the attention areas. Subsequently, the semantic information of marine animals is further understood through the HEM module. Finally, the present invention gradually predicts a mask consistent with the original image size through PPD to obtain an accurate segmentation mask of marine animals.
Claims
1. A method for segmenting marine animals based on deep multi-level feature aggregation, characterized in that The steps are as follows: Step 1: Obtain animal pictures in the ocean scene as training and test data, and precise segmentation mask annotation is required for the data; Step 2: Retain the feature extraction backbone part of SAM, and inject low-rank adaptation and adapters into the multi-head self-attention and feed-forward network of each Transformer module in SAM respectively; the feature extractor after the change is named ASE; Let X i ∈R N×D be the input of the i-th Transformer module, where N represents the number of visual tokens and D represents the embedding dimension; the multi-head self-attention mechanism modified by LoRA is expressed as the following formula: K i = W k (X i ) Among them, W q , W k , W v are the weights of three linear projection layers for generating the original query, key, and value matrices respectively; Q i , K i , V i are the inputs of the multi-head self-attention respectively; is the weight of the compressed linear projection layer in LoRA, is the weight of the upsampling linear projection layer in LoRA, which is used to reduce and restore the feature dimensions respectively; M is the dimension of the compressed projection; is the intermediate state generated by the multi-head self-attention mechanism; in this way, the present invention can freeze the pre-trained weights (W q , W k , W v ), and greatly reduce the number of trainable parameters by using low-rank matrix factorization; In addition, an adapter is inserted into the feed-forward network as follows: Among them, MLP is a multi-layer perceptron; σ is an activation function; is the intermediate state obtained by the feed-forward network; are the weights of two linear projection layers, respectively used to reduce and restore the feature dimension; P is the dimension projected in the adapter; The present invention can achieve parameter-efficient fine-tuning and adapt the pre-trained SAM encoder to the marine scenario; Due to the complex underwater environment, making full use of local detail information and global context information is crucial for achieving robust and accurate marine animal segmentation. Different Transformer layers can capture semantic information at different levels. Generally, shallow layers retain more local details, while deep layers express more context information. Therefore, in order to enable the model constructed by the present invention to utilize richer marine information, the present invention constructs a super feature map extraction module HEM for fusing multi-level feature maps from ASE; the multi-level feature maps extracted by ASE provide comprehensive guidance for the subsequent mask prediction process; first, an image I ∈ R H×W×3 is input into ASE, and the outputs of different Transformer layers are obtained, where H and W are the height and width of the image; the 3rd, 6th, 9th, and 12th Transformer layers are selected, and multi-level features, namely X i (i = 3, 6, 9, 12) are obtained; then, they are transformed into spatial feature maps To consider these multiple feature maps simultaneously, the following feature aggregation is performed: R i = φ 1×1 (F i ) Among them, φ 1×1 and φ 3×3 represent convolutions with kernel sizes of 1x1 and 3x3 respectively, is the intermediate state of the first super feature map; After that, a channel attention mechanism is introduced to generate super feature maps: Among them, GAP represents global average pooling, δ is the Sigmoid activation function, and H j , (j = 1, 2, 3, 4) is the super feature map obtained after fusion; To fully fuse the super feature maps obtained by HEM and the multi-level feature maps obtained by ASE, the present invention constructs a fusion attention module FAM; first, the multi-level feature maps from ASE are upsampled and the input features are adjusted to the same size; then, they are fused in the following way: D j+1 = FAM(U i , D j , H j ) Among them, is an operation adopted in linear interpolation, and U i is the result obtained after the features are upsampled by ASE; D j is the output of the j-th layer in PPD; for FAM, a channel attention mechanism is used to evaluate the importance of features; meanwhile, a residual structure is used to enhance the feature representation ability; this process is expressed as: Y j = φ 1×1 ([U i , D j , H j ) W j = δ(φ 1×1 (GAP(Y j )) + φ 1×1 (GMP(Y j ))) Among them, Y j is the intermediate state of fusing U i , D j , H j ; W j is the weight of the feature channel; GMP represents global maximum pooling; finally, the final prediction mask result is obtained through the gradually expanding feature prediction map; Step 3: During the training process, three levels of deep supervision are adopted, that is, the pixel-level supervision is the binary cross-entropy loss, the region-level supervision is the SSIM loss, and the overall-level supervision is the IoU loss; therefore, L is defined as the total loss consisting of three parts: L = L BCE + L SSIM + L IoU 。