Vision-language model-based polar sea ice semantic segmentation method
Through the visual-language model's semantic segmentation method, the challenge of polar sea ice segmentation under visible light images is solved, flexible and accurate sea ice segmentation is achieved, segmentation accuracy and efficiency are enhanced, and it is suitable for polar environmental monitoring and climate research.
Patent Information
- Application Number
- CN202510507701.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art polar sea ice segmentation faces challenges such as extreme light changes, low contrast and similar colors of sea ice and sea water under visible light images, making it difficult to achieve high-precision segmentation. The application of language-image pre-training framework in polar sea ice segmentation tasks has not been fully explored.
The polar sea ice semantic segmentation method based on the vision-language model is adopted. By obtaining the visible polar sea ice data set, the visual-language model is used for encoding, the text embedding vector and image embedding vector are fused, and the text semantic information is decoded, the polar sea ice segmentation mask is obtained, and the cross-modal fusion module and loss function are introduced for training.
It realizes flexible, accurate and efficient segmentation of polar sea ice, solves the shortcomings of the existing methods that require manual interaction, can handle complex sea ice segmentation tasks, enhances fine-grained alignment capabilities, and improves segmentation accuracy.
Smart Images

Figure CN120259670A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image target segmentation, and particularly relates to a polar sea ice semantic segmentation method based on a vision-language model. Background Art
[0002] The polar sea ice segmentation technology based on computer vision and image processing is of great significance in polar environmental monitoring and climate research. At present, a large number of studies mainly perform sea ice segmentation based on remote sensing images (such as SAR, multispectral images). These methods can effectively overcome problems such as illumination and texture complexity in specific bands. However, the research on polar sea ice segmentation in visible light images is relatively insufficient. Although visible light images have higher resolution and more intuitive visual information, they face challenges such as extreme illumination changes, low contrast, and similar colors between sea ice and sea water in the polar environment, resulting in difficulty for existing methods to achieve high-precision segmentation. Although the language-image pre-training framework has achieved remarkable success in various visual tasks such as object detection and semantic segmentation, the application of semantic prior knowledge in the VLM in the polar sea ice segmentation task has not been explored. Therefore, there is an urgent need to develop a polar sea ice semantic segmentation method based on a vision-language model, make full use of the advantages of visible light images, make up for the deficiencies of current research, and provide more comprehensive technical support for polar monitoring and research. Summary of the Invention
[0003] To solve the above technical problems, the present invention proposes a polar sea ice semantic segmentation method based on a vision-language model, which can more flexibly select and segment target regions and more accurately and efficiently segment different types of ice.
[0004] The present invention proposes the following solutions:
[0005] A polar sea ice semantic segmentation method based on a vision-language model, comprising:
[0006] Obtaining a visible light polar sea ice data set;
[0007] Encoding the visible light polar sea ice data set by using a vision-language model to obtain a text embedding vector and an image embedding vector;
[0008] Fusing the text embedding vector and the image embedding vector to obtain image semantic information and text semantic information;
[0009] Decoding the image semantic information based on the text semantic information to obtain a polar sea ice segmentation mask.
[0010] Optionally, obtaining a visible light polar sea ice data set includes:
[0011] Obtaining video data collected by a polar icebreaker;
[0012] Annotate the video data, perform data augmentation on the annotated video data, and obtain the visible light polar sea ice dataset.
[0013] Optionally, encoding the visible light polar sea ice dataset to obtain the text embedding vector includes:
[0014] Preprocess the input text in the visible light polar sea ice dataset, convert and obtain word vectors;
[0015] Input the word vectors into the CLIP text editor to obtain the text embedding vector.
[0016] Optionally, encoding the visible light polar sea ice dataset to obtain the image embedding vector includes:
[0017] Slice the input image in the visible light polar sea ice dataset into several blocks, and add positional encoding to the blocks to obtain embedding vectors;
[0018] Slice the embedding vectors to obtain several sliced blocks;
[0019] Use the self-attention mechanism to calculate the sliced blocks to obtain the image embedding vector.
[0020] Optionally, fusing the text embedding vector and the image embedding vector to obtain image semantic information and text semantic information includes:
[0021] Input the text embedding vector and the image embedding vector into the cross-modal fusion module to obtain image semantic information and text semantic information, where the cross-modal fusion module includes: a text-image cross-attention unit and an image-text cross-attention unit;
[0022] The text-image cross-attention unit is used to calculate the feature correlation of each text token to the image region with the text as the query and the image as the key-value;
[0023] The image-text cross-attention unit is used to calculate the semantic association degree of the image region to the text token with the image as the query and the text as the key-value.
[0024] Optionally, calculating the feature correlation of each text token to the image region includes:
[0025] Separate the image embedding vector in the channel dimension to obtain a first image embedding vector and a second image embedding vector;
[0026] Perform local spatial feature extraction and convolution operations on the first image embedding vector to obtain a third image embedding vector;
[0027] Perform a linear mapping on the text embedding vector to obtain a first text embedding vector;
[0028] Perform an attention calculation on the third image embedding vector and the first text embedding vector to obtain regions with high correlation of the text in the image features.
[0029] Optionally, calculating the semantic association degree of the image region to the text token includes:
[0030] Adopt step-by-step cross-attention to inversely calculate the third image embedding vector and the first text embedding vector;
[0031] Concatenate the third image embedding vector and the second image embedding vector to obtain reverse query of the image features for associated descriptive features in the text embedding vector.
[0032] Optionally, based on the text semantic information, decoding the image semantic information to obtain a polar sea ice segmentation mask includes:
[0033] Input the image semantic information into a Mask decoder to obtain a polar sea ice segmentation mask, where the Mask decoder includes: an upsampling module, a segmentation head module, and a text head module;
[0034] The upsampling module is used to restore the image size;
[0035] The segmentation head module is used to generate a mask map;
[0036] The text head module is used to calculate the cosine similarity matrix between the mask map and the corresponding text semantic information.
[0037] Optionally, based on the text semantic information, decoding the image semantic information to obtain a polar sea ice segmentation mask further includes: training the text semantic information and the image semantic information using a loss function, and the loss function includes: a text-image contrast loss function, a Dice loss function, and a Focal loss function.
[0038] Compared with the prior art, the present invention has the following advantages and technical effects:
[0039] (1) The present invention improves and innovates on the original SAM segmentation algorithm. By means of text-driven guidance, on the one hand, the polar sea ice segmentation becomes more flexible, solving the disadvantages of SAM that require manual point annotation and box annotation, and can quickly obtain segmentation results related to the description without manual interaction. On the other hand, based on text description as guidance, SAM can handle more complex sea ice segmentation tasks, and can perform segmentation tasks for specific categories of sea ice through different text inputs, achieving more accurate segmentation.
[0040] (2) The present invention proposes a cross-modal fusion module TIFM, which can effectively fuse the text information encoded by the CLIP text encoder and the image information encoded by the SAM image encoder. By calculating the bidirectional attention between the text and the image, it can achieve the dynamic association between the text and the image, enhance the fine-grained alignment ability, and provide a basis for the downstream sea ice segmentation task. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0042] Figure 1 is a training flow chart of a polar sea ice semantic segmentation method based on a vision-language model according to an embodiment of the present invention;
[0043] Figure 2 is a flow chart of a polar sea ice semantic segmentation method based on a vision-language model according to an embodiment of the present invention;
[0044] Figure 3 is a schematic structural diagram of the cross-modal fusion module TIFM according to an embodiment of the present invention;
[0045] Figure 4 is a schematic structural diagram of the text-image cross-attention unit and the image-text cross-attention unit according to an embodiment of the present invention;
[0046] Figure 5 is to guide the segmentation of the corresponding region map by specifying the text of the model segmentation region according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0048] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] This embodiment proposes a polar sea ice semantic segmentation method based on a vision-language model, as Figure 2 shown, and specifically includes the following steps:
[0050] Obtain a visible light polar sea ice data set;
[0051] Encode the visible light polar sea ice dataset using a vision - language model to obtain text embedding vectors and image embedding vectors;
[0052] Fuse the text embedding vectors and image embedding vectors to obtain image semantic information and text semantic information;
[0053] Based on the text semantic information, decode the image semantic information to obtain the polar sea ice segmentation mask.
[0054] Specifically, the following technical solutions are proposed in this embodiment: construct a visible light polar sea ice dataset and perform data augmentation;
[0055] Construct a network model, use the text encoder of the CLIP model to encode the text to obtain text embedding vectors, use the image encoder of the SAM model to encode the image to obtain image embedding vectors; design a text - image cross - modal fusion module to obtain text - guided image semantic information and text semantic information aligned with the image; design a Mask decoder to upsample the image semantic information to obtain the polar sea ice segmentation mask;
[0056] Design a loss function and train it, including text - image contrast loss, Dice loss, and Focal loss.
[0057] This embodiment can more flexibly select and segment the target area, and more accurately and efficiently segment different types of ice. In addition, a cross - modal fusion module is introduced. Through the calculation of the mutual attention mechanism between text and image, the text can be fused with the image, enabling the text to accurately select the corresponding image area. At the same time, combined with the powerful zero - shot segmentation ability of SAM itself, it can achieve a more accurate and high - precision visible light polar sea ice segmentation task.
[0058] Furthermore, obtaining the visible light polar sea ice dataset includes:
[0059] Obtain the video data collected by polar icebreakers;
[0060] Annotate the video data and perform data augmentation on the annotated video data to obtain the visible light polar sea ice dataset.
[0061] Specifically, the construction of the visible light polar sea ice dataset includes: using X - AnyLabeling to annotate the videos collected by polar icebreakers; the polar sea ice dataset is mainly divided into four categories: sky, water, thin ice, and thick ice; perform data augmentation on the annotated polar sea ice dataset, including: flipping, translation, rotation, contrast enhancement, saturation enhancement, Mosaic enhancement, etc.; divide the dataset into a training set, a validation set, and a test set according to 7:2:1.
[0062] Furthermore, encoding the visible light polar sea ice dataset to obtain text embedding vectors includes:
[0063] Preprocess the input text in the visible light polar sea ice dataset, convert it, and obtain word vectors;
[0064] Input the word vectors into the CLIP text editor to obtain text embedding vectors.
[0065] Specifically, the process of encoding text using the text encoder of the CLIP model includes: First, perform tokenize preprocessing on the input text to convert each word or sub-word into the corresponding vocabulary index, obtaining an array, usually an index array; then completely freeze the text encoder of the CLIP pre-trained model and input the obtained index array into the text encoder to obtain text embedding vectors.
[0066] Furthermore, encoding the visible light polar sea ice dataset to obtain image embedding vectors includes:
[0067] Slice the input image in the visible light polar sea ice dataset into several blocks, and add positional encoding to the blocks to obtain embedding vectors;
[0068] Slice the embedding vectors to obtain several sliced blocks;
[0069] Use the self-attention mechanism to calculate the sliced blocks to obtain image embedding vectors.
[0070] Specifically, the process of encoding an image using the image encoder of the SAM model includes: Slice the input image into multiple patches of size PxP, add positional encoding to obtain embedding vectors, and then input them into several Transformer blocks. After calculation by the self-attention mechanism, the final image embedding vectors are obtained.
[0071] Furthermore, fusing the text embedding vectors and the image embedding vectors to obtain image semantic information and text semantic information includes:
[0072] Input the text embedding vectors and the image embedding vectors into the cross-modal fusion module to obtain image semantic information and text semantic information, where the cross-modal fusion module includes: a text-image cross-attention unit and an image-text cross-attention unit;
[0073] The text-image cross-attention unit is used to calculate the feature correlation of each text token to the image region with the text as the query and the image as the key-value;
[0074] The image-text cross-attention unit is used to calculate the semantic association degree of the image region to the text token with the image as the query and the text as the key-value.
[0075] Specifically, for the design text-image cross-modal fusion module TIFM, the process of obtaining text-guided image semantic information and text semantic information aligned with the image includes: The text embedding vector and the image embedding vector are input into the cross-modal fusion module, which mainly includes two core modules. One is the Text-Image Cross Attention Model, and the other is the Image-Text Cross Attention Model. The Text-Image Cross Attention Model is used to calculate the regions with high correlation of the text in the image features, and the Image-Text Cross Attention is used to query the associated descriptive features of the image in the text embedding vector in reverse. Through this two-way design, the fine-grained alignment ability between the text and the image is enhanced, and finally the fused image embedding vector and the text embedding vector are obtained.
[0076] Furthermore, calculating the feature correlation of each text token to the image region includes:
[0077] The image embedding vector is separated in the channel dimension to obtain the first image embedding vector and the second image embedding vector;
[0078] The first image embedding vector is subjected to local spatial feature extraction and convolution operation to obtain the third image embedding vector;
[0079] The text embedding vector is linearly mapped to obtain the first text embedding vector;
[0080] The third image embedding vector and the first text embedding vector are subjected to attention calculation to obtain the regions with high correlation of the text in the image features.
[0081] Furthermore, calculating the semantic association degree of the image region to the text token includes:
[0082] The third image embedding vector and the first text embedding vector are calculated in reverse by using step-by-step cross attention;
[0083] The third image embedding vector and the second image embedding vector are concatenated to obtain the associated descriptive features of the image feature query in the text embedding vector in reverse.
[0084] Furthermore, based on the text semantic information, decoding the image semantic information to obtain the polar sea ice segmentation mask includes:
[0085] The image semantic information is input into the Mask decoder to obtain the polar sea ice segmentation mask, where the Mask decoder includes: an upsampling module, a segmentation head module, and a text head module;
[0086] Upsampling module, used to restore the image size;
[0087] Segmentation head module, used to generate mask map;
[0088] The text header module is used to calculate the cosine similarity matrix between the mask image and the corresponding text semantic information.
[0089] Specifically, the process of designing a Mask decoder and inputting the fused semantic information into the decoder to generate the final polar sea ice segmentation image includes: the Mask decoder is mainly divided into a segmentation head and a text head. The segmentation head is mainly responsible for generating masks and IOU scores: first, the fused image embedding vector is reshaped to the original image size and upsampled to obtain a feature map. Secondly, the fused text embedding vector is passed through two different MLP layers to obtain a Mask token and an IOU score respectively. The Mask token is used to perform matrix multiplication with the feature map to obtain a Mask with the same number of output categories. The IOUScore is used to measure the confidence of the mask map. The text head is mainly responsible for calculating the cosine similarity matrix between the mask map and the corresponding text: first, the fused image embedding vector and the text embedding vector are passed through different MLP layers for feature mapping, and the feature dimension is unified. Then, L2 is used for normalization. Finally, the cosine similarity matrix is calculated. The cosine similarity matrix is used for subsequent text image contrast loss calculation. The cosine similarity matrix is combined with the cross entropy loss to bring the matching text-image pairs closer and push the unmatched pairs farther.
[0090] Furthermore, based on the text semantic information, the image semantic information is decoded to obtain the polar sea ice segmentation mask, which also includes: using a loss function to train the text semantic information and the image semantic information, and the loss function includes: a text image contrast loss function, a Dice loss function and a Focal loss function.
[0091] Specifically, the process of designing and training loss functions, including text-image contrast loss, Dice loss, and Focal loss, includes: text-image contrast loss, which enables the model to learn to align the features of text and image during training, combines the cosine similarity matrix and uses cross-entropy loss to bring matching text-image pairs closer and push unmatched pairs further away; Dice loss, which is mainly used to measure the degree of overlap between the predicted segmentation area and the true area. Dice loss is usually defined as 1-Dice coefficient; Focal Loss reduces the contribution of easily classified background pixels, allowing the model to pay more attention to difficult-to-classify areas, thereby alleviating category imbalance and improving segmentation accuracy.
[0092] The following is combined with Figures 1-5 This embodiment is described in detail:
[0093] In this embodiment, a polar sea ice semantic segmentation method based on a vision-language model is provided, which is of great significance for polar exploration, ensuring shipping safety, etc. As Figure 1 shown, the main steps are as follows:
[0094] Step S1: Construct a polar sea ice dataset:
[0095] The polar environment is photographed by a camera mounted on a polar icebreaker.
[0096] Step S1 further includes the following steps:
[0097] Step S11: Collect a video sequence with a size of 1920*1080 for a total of 8 hours. The video sequence is preprocessed, and 50 representative video sequences are selected. One picture is selected every 10 frames, and a total of 5120 polar sea ice pictures are collected.
[0098] Step S12: Use the X-Anylabling software for semantic segmentation image annotation, which is divided into four categories: sky, water, thin ice, and thick ice.
[0099] Step S13: Perform data augmentation operations, including flipping, translation, contrast enhancement, saturation enhancement, Mosaic enhancement, etc., and divide the dataset into a training set, a validation set, and a test set according to 7:2:1. The dataset is used as the input of the model in the form of text-image pairs.
[0100] Step S2: Construct a network model:
[0101] Construct a CLIP-guided SAM visible light polar segmentation network, which mainly includes a CLIP text encoder, a SAM image encoder, a cross-modal fusion module, and a Mask decoder.
[0102] Step S2 further includes the following steps:
[0103] Step S21: Use the CLIP text encoder to extract text features to obtain text embedding vectors.
[0104] The CLIP text encoder is stacked by several Transformer blocks. This encoder is responsible for mapping the input text to a high-dimensional embedding space to generate the corresponding text feature representation T text ∈R B×K×D, where K is the length of the input text sequence, which is the number of categories to be segmented here, and D is the text feature dimension. Since CLIP has been fully pre-trained on a large-scale multi-modal dataset, its text encoder already has powerful feature extraction capabilities. Therefore, during the training process, the CLIP text encoder remains completely frozen and does not participate in parameter updates to ensure the stable performance of the learned text feature representation ability and improve the training efficiency of the model.
[0105] Step S22: Use the image encoder of SAM to extract image features to obtain an image embedding vector:
[0106] The image encoder of SAM is based on the powerful Vision Transformer (ViT) architecture, which is optimized for multi-task visual understanding and segmentation tasks. This encoder can efficiently extract multi-scale features from the input image and map them to a high-dimensional embedding space to generate rich image representations. During the training process, the SAM image encoder globally models the relationships between image regions through layer-by-layer Transformer calculations and finally outputs an image embedding vector F that is downsampled 16 times, stable and semantically consistent. img ∈R B×C×H×W . Since SAM has been pre-trained on a large amount of data, its encoder remains completely frozen in the current task to ensure its generalization ability and improve the inference efficiency and model stability.
[0107] Step S23: Design a text-image cross-modal fusion module TIFM to obtain text-guided image semantic information and text semantic information aligned with the image;
[0108] As Figures 3-4 shown: The input of the cross-modal fusion module comes from the text embedding vector T of the CLIP text encoder text ∈R B×K×D and the image embedding vector F of the SAM image encoder img ∈R B×C×H×W , which mainly includes two stages: The first stage performs cross-modal fusion of Text-Image. Under the image branch, first, F img ∈R B×C×H×W is separated in the channel dimension and split into where undergoes a shallow local spatial feature extraction through Conv and then passes through the BottleNeck module. The BottleNeck module is a convolutional layer with channel compression-expansion containing residual connections, and finally outputs with the shape unchanged At the same time, the text T text ∈R B×K×D undergoes a linear mapping through a layer of Linear and remains unchanged in shape to obtain Ttext ∈R B×K×D , and is input to the Text-Image Cross Attention Model for attention calculation. At this time, the text is the Query, and the Key and Value are the images (where the image embedding vector needs to be feature-mapped and reshaped into and linearly mapped through Linear to which is consistent with the dimension of the text embedding vector). The text, as the Query, can provide semantic guidance, and through attention calculation, it can perceive which parts of the image are most relevant to the text, and then add them to the Query to fuse the information of the text and the image. By combining the weighted image features with the original text query vector, the model can better understand the semantic relationship between the text and the image, thereby improving the performance of downstream tasks. In the second stage, cross-modal fusion of Image-Text is performed. The image is used as the Query, and the Query after fusion in the first stage is used as the Key and Value in the second stage for the second attention fusion. Using this step-by-step interactive attention calculation can further optimize the feature alignment between the image and the text, and finally add it to the image information to enhance the image information and make it more capable of semantic understanding to assist in completing downstream tasks. Finally, the fused text embedding vector T text ∈R B×L×D and the fused and reshaped image embedding vector where and are concatenated and fused through a 1x1 Conv for channel fusion and then output to obtain F img ∈R B ×C×H×W .
[0109] Step S24: Design a Mask decoder to upsample the image semantic information to obtain the polar sea ice segmentation mask:
[0110] The Mask decoder mainly includes an upsampling module, a segmentation head, and a text head. The upsampling module mainly upsamples the fused image embedding vector F img ∈R B×C×H×W to obtain F img ∈R B×C×16H×16W ; the segmentation head mainly processes the fused text embedding vector T text ∈R B×K×D by connecting to K MLP layers to generate K mask tokens, namely M token ∈R B ×K×D , where K is the number of categories, and then F img ∈R B×C×16H×16WAdjust the channel to degree F img ∈R B×D×(16*16HW) At the same time, perform matrix multiplication with the mask token and Reshape to finally obtain K masks Mask, that is, M ∈ R B×K×16H×16W ; In addition, for T text ∈R B ×K×D Connect to an MLP layer to generate the IOU Score, and obtain T iou ∈R B×K×1 ; The text head is mainly used to calculate the cosine similarity matrix between the text and the image, and the input is the text embedding vector T text ∈R B×K×D and the generated image mask M ∈ R B×K×16H×16W , for M ∈ R B×K×16H×16W Calculate the average feature of each category, that is, K M ∈ R B×16H×16W×1 and the reshaped F img ∈R B ×16H×16W×D Perform element-wise multiplication and global average summation to obtain I k ∈R B×K×D , the formula is as follows:
[0111]
[0112] Among them, i represents the row, j represents the column, k represents the kth Mask, and ε is used to prevent division by zero.
[0113] Then, perform L2 normalization on T text ∈R B×K×D and I k ∈R B×K×D respectively, and calculate the cosine similarity matrix to measure the similarity between the text and the image. The formula is as follows:
[0114] S = α · l2_norm(I k ) · l2_norm(T text ) T
[0115] Among them, S ∈ R B×K×K is the cosine similarity matrix, and α is a learnable weight used to adjust the distribution.
[0116] Step S3: Design the loss function and train, including the text-image contrast loss, Dice loss, and Focal loss.
[0117] Text-image contrast loss: The goal is to make the text of each category similar to the corresponding image features and dissimilar to the rest. Specifically, its loss function is as follows:
[0118] L sim= cross_entropy(S, GT)
[0119] where S ∈ R B×K×K is the cosine similarity matrix, and GT is the label, indicating that the image features of each category should correspond to the text embedding with the same subscript.
[0120] Dice Loss: The Dice coefficient (or IoU) measures the overlap between the predicted region and the ground truth region. Specifically, its loss function is as follows:
[0121]
[0122] where A is the predicted region and B is the ground truth region.
[0123] Focal Loss: It is used to handle the class imbalance problem in the visible light polar sea ice dataset. Specifically, the loss function is as follows:
[0124] L Focal = -α · (1 - p t ) γ · log(p t )
[0125] where α is the adjustment factor and p t is the probability of being predicted as the target class.
[0126] The final total loss is shown as follows:
[0127] L total = λ1L sim + λ2L Dice + λ3L Focal
[0128] where λ1, λ2, λ3 are the loss weight coefficients.
[0129] Step S4: Train the model:
[0130] Use the Adam algorithm to perform iterative optimization of the model. The number of iterations epoch is 1000, the batch size is 32, the learning rate is 0.001, the momentum is 0.937, the decay rate is 0.0005, and the environment configuration is Python
[0131] 3.10, Pytorch 2.5.1 + cu121, and the graphics card is NVIDIA GeForce RTX 4090D, 24G. The performance metrics on the validation set after the final training are shown in Table 1. It can be seen that the visible light polar sea ice SAM segmentation algorithm based on CLIP text guidance proposed by the present invention can have high accuracy and segmentation precision in the sea ice segmentation scenario.
[0132] Table 1
[0133]
[0134] Step S5: Specify the text of the segmentation area to guide the model to generate the corresponding segmentation map:
[0135] As Figure 5 shown, the text describes the text of the area to be segmented, and the picture is the corresponding segmentation map generated through text guidance. It can be seen that by inputting the described text, only the area corresponding to the described text can be segmented, and any area can be specified for simultaneous segmentation, making the sea ice segmentation task more flexible and efficient while maintaining the segmentation accuracy.
[0136] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A polar sea ice semantic segmentation method based on a vision-language model, characterized in that, Comprising: Obtain a visible light polar sea ice dataset; Encode the visible light polar sea ice dataset using a vision-language model to obtain a text embedding vector and an image embedding vector; Fuse the text embedding vector and the image embedding vector to obtain image semantic information and text semantic information; Decode the image semantic information based on the text semantic information to obtain a polar sea ice segmentation mask.
2. The polar sea ice semantic segmentation method based on a vision-language model according to claim 1, characterized in that, Obtaining a visible light polar sea ice dataset includes: Obtain video data collected by a polar icebreaker; Annotate the video data, and perform data augmentation on the annotated video data to obtain the visible light polar sea ice dataset.
3. A polar sea ice semantic segmentation method based on a vision-language model according to claim 1, characterized in that Encoding the visible light polar sea ice dataset to obtain the text embedding vector includes: Preprocess the input text in the visible light polar sea ice dataset, and convert it to obtain word vectors; Input the word vectors into a CLIP text encoder to obtain the text embedding vector.
4. A polar sea ice semantic segmentation method based on a vision-language model according to claim 1, characterized in that Encoding the visible light polar sea ice dataset to obtain the image embedding vector includes: Slice the input image in the visible light polar sea ice dataset into several blocks, and add position encoding to the blocks to obtain embedding vectors; Slice the embedding vectors to obtain several sliced blocks; Use the self-attention mechanism to calculate the sliced blocks to obtain the image embedding vector.
5. A polar sea ice semantic segmentation method based on a vision-language model according to claim 1, characterized in that Fusing the text embedding vector and the image embedding vector to obtain image semantic information and text semantic information includes: Input the text embedding vector and the image embedding vector into a cross-modal fusion module to obtain image semantic information and text semantic information, where the cross-modal fusion module includes: a text-image cross-attention unit and an image-text cross-attention unit; The text-image cross-attention unit is used to calculate the feature correlation of each text token to the image region with the text as the query and the image as the key-value; The image-text cross-attention unit is used to calculate the semantic association degree of the image region to the text token with the image as the query and the text as the key-value.
6. The polar sea ice semantic segmentation method based on a vision-language model according to claim 5, wherein, Calculating the feature correlation of each text token to the image region includes: Separate the image embedding vector in the channel dimension to obtain a first image embedding vector and a second image embedding vector; Perform local spatial feature extraction and convolution operations on the first image embedding vector to obtain a third image embedding vector; Linearly map the text embedding vector to obtain a first text embedding vector; Perform attention calculation on the third image embedding vector and the first text embedding vector to obtain the regions in the image features with high correlation to the text.
7. A method for polar sea ice semantic segmentation based on a vision-language model according to claim 6, characterized in that, Calculating the semantic association degree of the image region to the text token includes: Use step-by-step cross-attention to reversely calculate the third image embedding vector and the first text embedding vector; Concatenate the third image embedding vector and the second image embedding vector to obtain the reverse of the image features to query the associated descriptive features in the text embedding vector.
8. A polar sea ice semantic segmentation method based on a vision-language model according to claim 1, characterized in that Decoding the image semantic information based on the text semantic information to obtain a polar sea ice segmentation mask includes: Input the image semantic information into the Mask decoder to obtain the polar sea ice segmentation mask, where the Mask decoder includes: an upsampling module, a segmentation head module, and a text head module; The upsampling module is used to restore the image size; The segmentation head module is used to generate a mask map; The text head module is used to calculate the cosine similarity matrix between the mask map and the corresponding text semantic information.
9. A polar sea ice semantic segmentation method based on a vision-language model according to claim 1, characterized in that Based on the text semantic information, decoding the image semantic information to obtain the polar sea ice segmentation mask further includes: training the text semantic information and the image semantic information using a loss function, where the loss function includes: a text-image contrast loss function, a Dice loss function, and a Focal loss function.
Citation Information
Cited By
Remote sensing scene graph guided semantic information reasoning method and device, equipment and medium
CN120599616A
Anaphora image segmentation method based on cross-modal mirror image alignment and double contrast learning
CN122115877A