Remote sensing image semantic segmentation method based on image-text multi-modal feature fusion
By constructing text prompts of the data set object category in the semantic segmentation of remote sensing images and combining image-text feature fusion, using TIFF module and loss function optimization, the problem of insufficient segmentation of remote sensing images in complex scenarios is solved, and a higher accuracy and robust segmentation effect is achieved.
Patent Information
- Application Number
- CN202510734933.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-08-19
AI Technical Summary
The existing remote sensing image semantic segmentation method ignores the potential of multimodal data sources of images and text, resulting in insufficient performance when processing complex scenarios. In particular, the difference in image and text feature representation and structure leads to difficulty in learning the model, low utilization of text cues, and failure to effectively integrate interactive information in the training target.
By constructing text prompts for the name of the object category of the data set, only text prompts for label files are added during the training stage, combining image encoder and text encoder, the TIFF module is used to achieve text-image features interaction and fusion, and the self-attention and cross-attention mechanisms are used to enhance feature consistency and suppress redundancy, and the model is optimized using cross-entropy and dice loss functions.
It significantly improves the accuracy and robustness of semantic segmentation of remote sensing images, improves the utilization rate of text features, and the model can more accurately capture the deep correlation between the image and text, improving the quality of segmentation results.
Smart Images

Figure CN120510385A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a remote sensing image semantic segmentation method based on image-text multimodal feature fusion, and belongs to the technical field of image processing. Background Art
[0002] Semantic segmentation of remote sensing images involves assigning each pixel in a remote sensing image to a predefined category, such as water, forest, or city. This technology has been widely used in fields such as environmental monitoring, urban planning, and disaster assessment. However, existing research has largely focused on learning representations in the visual feature space, relying primarily on pixel-level image features for classification and ignoring the potential of multimodal data sources such as images and text. This limitation results in insufficient performance when dealing with complex scenes, especially those with rich semantic information.
[0003] In remote sensing image applications, images provide visual information of objects, while text can provide rich semantic information through descriptive language. Fusion of these two modalities enables the model to more comprehensively understand and parse complex geographic information. Currently, most multimodal methods use all category names in the dataset as textual cues during the training phase and apply them to all training images. However, each image typically contains only a small number of relevant categories and is unrelated to other categories, resulting in low utilization of textual cues. In addition, text and images differ significantly in representation and structure. Directly using textual features to guide semantic features in images makes model learning difficult. Currently, the training goal of most methods is to maximize the similarity between images and text, but they do not show the interaction information between the fused images and text, which limits the learning effect of the model. Summary of the Invention
[0004] The purpose of the present invention is to provide a remote sensing image semantic segmentation method based on image-text multimodal feature fusion to address the above-mentioned shortcomings. This method significantly improves the accuracy and robustness of remote sensing image semantic segmentation through the effective fusion of image-text multimodal features.
[0005] The technical solution adopted by the present invention is: A remote sensing image semantic segmentation method based on image-text multimodal feature fusion includes the following steps: S1. Obtain original remote sensing images, preprocess them, and divide the datasets. S2. Construct text hints for each category based on the feature category names in the dataset. Input the text hints into a text encoder to obtain text features. Construct label text hints based on the feature category names in the label file. Input the label text hints into a text encoder to obtain label text features. The text features and label text features are concatenated together only during the training phase to form the final text features. S3. Divide the remote sensing image into multiple non-overlapping image blocks, and obtain the image block embedding vector after flattening and mapping all the image blocks. X , add the position code of each image block to the embedding vector of each image block to obtain X’ , construct a vector with the same dimension as the embedding vector [ CLS ], the embedded vector after adding position encoding X’ and[ CLS ] vectors are input into the image encoder respectively to obtain image features and vectors representing global image semantic information; S4. The obtained text features, image features and vectors representing global image semantic information are input into the TIFF module to realize the interaction and fusion between text and image features. In the TIFF module, the text features are first enhanced with the vector representing global image semantic information. The enhanced text features are further enhanced using the self-attention mechanism to ensure internal feature consistency and suppress redundant information. The enhanced text features and image features are then interacted using the cross-attention mechanism to realize text-image feature fusion. The intermediate fusion features in the cross-attention are processed by the activation function as the final fusion features. F ; S5. Fusion features F Input the segmentation head to obtain the final segmentation map; S6. Optimize and train the segmentation model by calculating the loss function and verify it on the test set to obtain the final segmentation model.
[0006] In the above method, the text prompt in step S2 is mapped to d dimensional embedding vector space, which is then input into the text encoder to obtain text features. The text encoder is composed of multiple layers of Transformer encoders stacked together.
[0007] In step S3, vector [ CLS ] is initialized to a random value and is gradually updated during training as the network learns. The image encoder is composed of multiple layers of Vision Transformer encoders. Within these layers, the [CLS] vector processes and aggregates information from different image patches layer by layer through a self-attention mechanism, ultimately containing the global semantic information of the image. The embedding vector is input into the multi-layer encoder, and the output vector of the last encoder layer is selected as the image feature.
[0008] The TIFF module described in step S4 first uses the vector representing the global image semantic information g Enhanced text features T : , in, is the enhanced text feature, is the Hadamard product operation, concat (•) is the splicing operation, conv (•) is a 1×1 convolution. The enhanced text features use the self-attention mechanism to further enhance internal feature consistency and suppress redundant information. The calculation formula is: , , in, φ (•) is a linear transformation, which is implemented as a fully connected layer. d is the vector dimension, softmax (•) is the softmax function, norm (•) is the normalization layer. The fusion feature acquisition process is as follows: first, linear transformation is applied to map the text features into a query vector (Q, query), and then linear transformation is applied to map the image features into f image Mapped into key vector (K, key) and value vector (V, value): , Calculate the attention score between Query and Key, and then pass sigmoid () After the activation function, the final fusion feature is directly obtained: , , K is the number of feature categories in the dataset, N Equal to the number of image blocks.
[0009] The segmentation head described in step S5 includes an upsampling layer and an argmax layer. First, the feature The size of , P is the size of the image block, and then upsampled to the same size as the original image by bilinear interpolation, that is, Finally, the argmax layer selects the channel with the highest probability in each category as the predicted category.
[0010] The loss function described in step S6 is a weighted sum of the cross entropy loss function and the dice loss function.
[0011] Another object of the present invention is to provide an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the remote sensing image semantic segmentation method based on image-text multimodal feature fusion as described above is implemented.
[0012] The beneficial effects of the present invention are: This method incorporates textual hints from label files during the training phase, improving the utilization of textual features and enabling more accurate guidance for the model when processing complex object categories. The TIFF module enables efficient interaction and fusion between text and image features, enabling the model to more accurately capture the deep connections between image and text, thereby improving the quality of segmentation results. This method significantly enhances the accuracy and robustness of semantic segmentation of remote sensing images and has broad potential for application. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a model network structure diagram of the present invention; Figure 3 This is a structural diagram of the TIFF module of the present invention; Figure 4 Schematic diagram of the process of obtaining image features and vectors representing global image semantic information in the present invention. DETAILED DESCRIPTION
[0014] The present invention is further described below with reference to specific embodiments.
[0015] Example 1 A remote sensing image semantic segmentation method based on image-text multimodal feature fusion includes the following steps: S1. Obtain the original remote sensing image, preprocess it, and divide the dataset: The original remote sensing images are divided into training set and test set in a ratio of 7:3. Depend on N train image-label pairs, the test set Depend on N test image-label pairs, where I i For an RGB image, M i is the corresponding image label.
[0016] S2. Construct text hints for each category based on the feature category names in the dataset. Input the text hints into the text encoder to obtain text features. Construct label text hints based on the feature category names in the label file. Input the label text hints into the text encoder to obtain label text features. Only during the training phase are the text features and label text features concatenated together to form the final text features: Assume that there are a total of K categories, and construct the text prompt "a satellite photo with a " for each category based on the name of the ground feature category in the dataset. , which are replaced by category names, such as airplane, bare-soil, buildings, etc., as guidance prompts for remote sensing images.
[0017] The text features are extracted through the text encoder. Specifically, the text encoder is composed of 12 layers of Transformer encoders. The input text prompt is pre-processed by word segmentation and the word embedding layer is mapped to a dimensional embedding vector space, and then input into the text encoder to obtain text features: , in, f text is the output text feature, Embedding () is the word embedding layer, prompt Is a text prompt. After the text encoder, K Text features , d is the feature dimension, It is i Text features of categories.
[0018] In order to enable text features to provide more accurate guidance prompts for the model and improve the utilization of text prompts, text prompts of label files are added during the training phase. Specifically, assuming that the label file contains k categories ( k ≤ K ), according to this k The name of the feature category is constructed to construct the label text prompt "a satellite photo with a", , and then pass the text encoder to obtain the label text features , Will T dataset and T label Stitch together to obtain the final text features T : , In order to ensure the effectiveness of the model, the text prompts of the label file are only added in the training phase, and the final text features are T = T dataset , which ensures that the model effectively learns the rich semantic information of textual prompts during training, thereby enhancing the generalization ability of the final model in practical applications.
[0019] S3. Divide the remote sensing image into multiple non-overlapping image blocks, and obtain the image block embedding vector after flattening and mapping all the image blocks. X , add the position code of each image block to the embedding vector of each image block to obtain X’ , construct a vector with the same dimension as the embedding vector [ CLS ], the embedded vector after adding position encoding X’ and[ CLS ] vectors are input into the multi-layer Vision Transformer encoder to obtain image features and vectors representing global image semantic information: The image encoder is composed of 4 layers of Vision Transformer encoders stacked together. Assume that the input image size is H × W ×3, divide the input image into non-overlapping patches, the size of each patch is P × P ×3, a total of N =( H × W ) / P 2 Then, each image block is flattened into a one-dimensional vector and mapped to an embedding vector of fixed dimension through a fully connected layer, whose vector representation is ,in d Is the dimension of the embedding vector. For all image blocks, the image block embedding vector is obtained after flattening and mapping. Preserve the position information of the image block and add the position code to the embedding vector of each image block: , in, p’ i is the image patch embedding vector with position encoding added, e i is the code of the corresponding position.
[0020] In order to obtain the global semantic information of the entire image, a vector is constructed The image block embedding vector after position encoding The vector is fed into the multi-layer Vision Transformer encoder, and its output is: , in, is the image feature, is the feature of the image block, yes CLS The output corresponding to the vector, in the multi-layer encoder, the [CLS] vector will process and aggregate information from different image blocks layer by layer through the self-attention mechanism, so as to ultimately contain the global semantic information of the image.
[0021] S4. The obtained text features, image features, and vector representing global image semantic information are input into the TIFF module to realize the interaction and fusion between text and image features. In the TIFF module, the vector representing global image semantic information is first used to enhance the text features. The enhanced text features are further enhanced using the self-attention mechanism to ensure internal feature consistency and suppress redundant information. The enhanced text features and image features are then interacted using the cross-attention mechanism to achieve text-image feature fusion. The intermediate fused features in the cross-attention process are processed by the activation function and used as the final fused features: Due to the inherent complexity of remote sensing images, it is difficult for the text information extracted by the text encoder to accurately correspond to a specific remote sensing image. To this end, this paper designs a TIFF (Text-Image Feature Fusion) module to effectively combine the high-level semantic features of the image with the text features. Specifically, we first use the enhanced text features that represent the global image semantic information: , in, is the enhanced text feature, is the Hadamard product operation, concat (•) is the splicing operation, conv (•) is the convolution.
[0022] Typically, each remote sensing image only contains a subset of ground object categories. Even with the addition of textual hints from the label file, some textual features are still irrelevant to the current training image. Therefore, a self-attention mechanism is used to promote the interaction between different features, thereby enhancing the consistency of internal textual features and suppressing redundant information. The calculation formula is: , , in, φ (•) is a linear transformation, which is implemented as a fully connected layer.d is the vector dimension, softmax (•) is the softmax function, norm (•) is the normalization layer.
[0023] Then, the cross attention mechanism is used to interactively realize the fusion of text and image features. Specifically, a linear transformation is applied to map the text features into a query vector (Q, query), and a linear transformation is applied to map the image features into a query vector (Q, query). f image Mapped into key vector (K, key) and value vector (V, value): , The final fusion feature is the intermediate product of the cross attention mechanism calculation process, that is, the attention score between the query and the key is calculated, and then the final fusion feature is directly obtained after the activation function. F : , , is the number of object categories in the dataset, which is equal to the number of image blocks.
[0024] S5. Fusion features F Input the segmentation head and get the final segmentation map: The segmentation head consists of an upsampling layer and an argmax layer. First, the features The size of reshape is , and then upsampled to the same size as the original image by bilinear interpolation, that is, Finally, the argmax layer selects the channel with the highest probability in each category as the predicted category to obtain the segmented image .
[0025] S6. Optimize and train the above segmentation model by calculating the loss function and verify it on the test set to obtain the final segmentation model: The cross entropy loss function is used to measure the difference between the model output and the true label. Assuming the image size is H × W , and there are C categories, for each pixel ( h , w ), whose true category label is y i , the category probability output by the model is p i , the cross entropy loss function can be expressed as: , Furthermore, in order to alleviate the significant class imbalance problem in remote sensing images, that is, the background area often occupies the majority, and the foreground target area (such as roads, rivers, houses, etc.) is relatively small, dice loss is used as an auxiliary loss function: , Where C is the total number of categories, y i is the true label value, which is 0 or 1. is the model’s predicted probability that the pixel is the foreground target area, N is the total number of pixels, ε is the smoothing factor, generally set to 1×10 -6 , to prevent the denominator from being 0.
[0026] The total loss function can be expressed as: , in α , β is the coefficient combining cross entropy loss and dice loss.
[0027] Example 2 An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the remote sensing image semantic segmentation method based on image-text multimodal feature fusion as described in Example 1 above is implemented.
[0028] The above is a further description of the present invention in conjunction with the embodiments, and the protection scope of the present invention is not limited thereto.
Claims
1. A remote sensing image semantic segmentation method based on image-text multimodal feature fusion, characterized by: The steps are as follows: S1. Obtain original remote sensing images, preprocess them, and divide the datasets. S2. Construct text hints for each category based on the feature category names in the dataset. Input the text hints into a text encoder to obtain text features. Construct label text hints based on the feature category names in the label file. Input the label text hints into a text encoder to obtain label text features. The text features and label text features are concatenated together only during the training phase to form the final text features. S3. Divide the remote sensing image into multiple non-overlapping image blocks, and obtain the image block embedding vector after flattening and mapping all the image blocks. X , add the position code of each image block to the embedding vector of each image block to obtain X’ , construct a vector with the same dimension as the embedding vector [ CLS ], the embedded vector after adding position encoding X’ and[ CLS ] vectors are input into the image encoder respectively to obtain image features and vectors representing global image semantic information; S4. The obtained text features, image features, and vectors representing global image semantic information are input into the TIFF module to realize the interaction and fusion between text and image features. In the TIFF module, the text features are first enhanced with the vector representing global image semantic information. The enhanced text features are further enhanced using the self-attention mechanism to ensure internal feature consistency and suppress redundant information. The enhanced text features and image features are then interacted using the cross-attention mechanism to realize text-image feature fusion. The intermediate fusion features in the cross-attention are processed by the activation function as the final fusion features. F ; S5. Fusion features F Input the segmentation head to obtain the final segmentation map; S6. Optimize and train the segmentation model by calculating the loss function and verify it on the test set to obtain the final segmentation model.
2. The remote sensing image semantic segmentation method based on image-text multimodal feature fusion according to claim 1, characterized in that: In step S2, the text prompt is mapped to d dimensional embedding vector space, which is then input into the text encoder to obtain text features. The text encoder is composed of multiple layers of Transformer encoders stacked together.
3. The remote sensing image semantic segmentation method based on image-text multimodal feature fusion according to claim 1, characterized in that: The image encoder in step S3 is composed of multiple layers of Vision Transformer encoders stacked together. In the multi-layer encoder, [ CLS ]The vector will process and aggregate information from different image blocks layer by layer through the self-attention mechanism, so that it will eventually contain the global semantic information of the image. The embedded vector after position encoding is input into the multi-layer encoder, and the output vector of the last layer of encoder is selected as the image feature.
4. The remote sensing image semantic segmentation method based on image-text multimodal feature fusion according to claim 1, characterized in that: The TIFF module described in step S4 first uses the vector representing the global image semantic information g Enhanced text features T : , in, is the enhanced text feature, is the Hadamard product operation, concat (•) is the splicing operation, conv (•) is a 1×1 convolution.
5. The remote sensing image semantic segmentation method based on image-text multimodal feature fusion according to claim 1, characterized in that: The enhanced text features described in step S4 use the self-attention mechanism to further enhance internal feature consistency and suppress redundant information. The calculation formula is: , , in, φ (•) is a linear transformation, which is implemented as a fully connected layer. d is the vector dimension, softmax (•) is the softmax function, norm (•) is the normalization layer.
6. The remote sensing image semantic segmentation method based on image-text multimodal feature fusion according to claim 1, characterized in that: The fused features described in step S4 F The acquisition process is: first apply linear transformation to transform the text features Mapped into query vector Q, apply linear transformation to transform image features f image Mapped into a key vector K Sum vector V : , calculate Q and K The attention score between sigmoid () After the activation function, the final fusion feature is directly obtained F : , , K is the number of feature categories in the dataset, N Equal to the number of image blocks.
7. The remote sensing image semantic segmentation method based on image-text multimodal feature fusion according to claim 1, characterized in that: The segmentation head described in step S5 includes an upsampling layer and an argmax layer. First, the feature The size of , P is the size of the image block, and then upsampled to the same size as the original image by bilinear interpolation, that is, Finally, the argmax layer selects the channel with the highest probability in each category as the predicted category.
8. The remote sensing image semantic segmentation method based on image-text multimodal feature fusion according to claim 1, characterized in that: The loss function described in step S6 is a weighted sum of the cross entropy loss function and the dice loss function.
Citation Information
Cited By
Weak supervision image defect segmentation method and system based on text guidance
CN120807555A