Ultrahigh-resolution remote sensing image semantic segmentation method based on balanced anchor point sampling and prototype representation

By adopting a method of equalized anchor sampling and prototype characterization in ultra-high resolution remote sensing image object segmentation model, combining multi-size anchor area clipping and pre-training visual language model, multiple challenges of existing models in the training and inference process are solved, achieving higher performance and robustness.

CN119964000AActive Publication Date: 2025-05-09ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510042137.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing ultra-high resolution remote sensing image geographic segmentation models have multiple challenges in the training and inference process, including the inability to capture complete geographic information, the lack of high-quality text information for feature enhancement, as well as the poor migration effect of the model on multi-source data and the high training cost.

Method used

Using a method based on balanced anchor sampling and prototype characterization, multimodal features are extracted through the combination of multi-size anchor area tailoring and pre-trained visual language model to improve the performance and robustness of the semantic segmentation model.

Benefits of technology

The performance and robustness of the semantic segmentation model of higher ultra-high resolution remote sensing images is achieved, which can better capture multi-scale context features and improve the model's recognition ability of a few categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964000A_ABST
    Figure CN119964000A_ABST
Patent Text Reader

Abstract

The invention discloses an ultrahigh-resolution remote sensing image semantic segmentation method based on balanced anchor point sampling and prototype representation, and the method comprises the steps: combining text features provided by a pre-trained visual language model with a conventional segmentation model; comprising the steps of data preprocessing, multi-size anchor point region clipping, target detection mark frame and sorting, visual feature extraction, text feature extraction, prototype representation injection, model training and the like. According to the method, the improvement of the overall performance of the model is promoted by pre-extracting the complete ground features, and the text information is injected into the traditional semantic segmentation model to carry out detailed semantic segmentation on the ultrahigh-resolution remote sensing image, so that the performance and robustness of the ultrahigh-resolution remote sensing image semantic segmentation model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image semantic segmentation, and in particular relates to an ultra-high resolution remote sensing image semantic segmentation method based on balanced anchor point sampling and prototype representation. Background Art

[0002] Ultra-High Resolution (UHR) images have developed rapidly in recent years, mainly due to the advancement of remote sensing technology and the improvement of satellite image acquisition capabilities. These images have become increasingly important in many fields because of their fine details and rich ground information.

[0003] In the past few years, vision-language pre-training models have achieved remarkable results in the field of remote sensing, such as GeoRSCLIP[Zhang Z, Zhao T, Guo Y, et al. RS5M and GeoRSCLIP: A large scale vision-language dataset and a large vision-language model for remote sensing[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024], RemoteCLIP[Liu F, Chen D, Guan Z, et al. RemoteClip: A vision language foundation model for remote sensing[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024] and SAM[Kirillov A, Mintun E, Ravi N, et al. Segment anything[C] / / Proceedings of the IEEE / CVFInternational Conference on Computer Vision.2023:4015-4026], etc. These models have achieved very significant results on specific tasks and image or text datasets; however, when the image data is expanded to the object segmentation of ultra-high resolution images, due to the limitations of the pre-training tasks and the changes in image resolution, the performance of the model drops sharply, and the results are unsatisfactory. The main challenges come from two aspects: on the one hand, the existing remote sensing image object segmentation models can only process data through random cropping during the training phase, or divide the entire ultra-high resolution image into fixed-size blocks for training, which usually cannot capture complete object information, nor can they apply high-quality text information for feature enhancement; on the other hand, the existing pre-trained visual language models in the remote sensing field are usually pre-trained on multi-source data, which are not effective in migrating downstream tasks, and the training cost is usually high.

[0004] In a typical ultra-high resolution remote sensing image object segmentation task, the main challenge is the huge size of the image. Due to the limitations of existing hardware and algorithm conditions, it is impossible to perform end-to-end training and reasoning for an entire image. The general solution is to split the ultra-high resolution image into small images of fixed size, convert the data set and send it to the semantic segmentation model for training. During the reasoning process, the sliding window algorithm gradually predicts an entire image to obtain the final segmentation result. This training will lead to two serious problems: one is that the distribution of the original data category is always maintained during the training process. If there is a long tail problem in the data set, the head category will be overfitted during training, the generalization ability will be reduced, and the tail category will always be in an underfitting state; the other is that the model always finds it difficult to see the complete object, and always distinguishes the object by texture features. It lacks feature information such as shape, so it is difficult to distinguish subdivided categories such as lakes and rivers, residential areas and industrial areas.

[0005] There are special methods for ultra-high resolution remote sensing image segmentation. One idea is to use a multi-branch network to extract features of different scales. One network extracts local features and the other network extracts contextual features. Then, by designing feature fusion modules and feature refinement modules, contextual features are aggregated into local features to obtain an enhanced feature, which is then used for subsequent processing. For example, GLNet[Zhang S, Song L, Gao C, et al. Glnet: Global local network for weakly supervised action localization[J]. IEEE Transactions on Multimedia, 2019, 22(10): 2610-2622], FctlNet[Li Q, Yang W, Liu W, et al. From contexts to locality: Ultra-high resolution image segmentation via locality-aware contextual correlation[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 7252-7261], WiCoNet[Ding L, Lin D, Lin S, et al. Looking outside the window:Wide-context transformer for the semanticsegmentation of high-resolution remote sensing images.arXiv 2021[J].arXivpreprint arXiv:2106.15754]、ISDNet[Guo S,Liu L,Gan Z,et al.Isdnet:Integratingshallow and deep networks for efficient ultra-high resolution segmentation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition.2022:4361-4370] and other methods all adopt this idea.Another idea is to reduce the complexity of network operations and try to process an entire remote sensing satellite image in one forward process. This method usually uses traditional feature extraction methods such as wavelet transform and combines them with lightweight convolutional neural networks for processing, such as WSDNet[Ji D, Zhao F, Lu H, et al. Ultra-high resolution segmentation with ultra-rich context: A novel benchmark[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023: 23621-23630] and GPWFormer[Ji D, Zhao F, Lu H. Guided patch-grouping wavelet transformer with spatial congruence for ultra-highresolution segmentation[J]. arXiv preprint arXiv: 2307.00711, 2023]. Based on these two methods, methods such as patch selection and scale adaptive selection have been derived, and some progress has been made; however, a patch is often a small pixel block of 14×14 or 32×32. The method of the present invention is based on a larger anchor point area and can provide richer contextual information.

[0006] During the investigation, we found that in the field of ultra-high resolution image segmentation, the existing segmentation models (such as Deeplabv3+[Chen LC, Zhu Y, Papandreou G, et al. Encoder-decoder with atrousseparable convolution for semantic image segmentation[C] / / Proceedings of theEuropean conference on computer vision(ECCV).2018:801-818], Segformer[Xie E, Wang W, Yu Z, et al. SegFormer: Simple and efficient design for semanticsegmentation with transformers[J]. Advances in neural information processing systems, 2021, 34:12077-12090], etc.) have not been combined with the multimodal features provided by pre-trained text image models (such as GeoRSCLIP, RemoteCLIP, etc.). The former usually provides high-quality segmentation results within a specific domain, but lacks rich knowledge in the entire field; the latter usually contains rich knowledge and is good at capturing the category features of objects, but it is difficult to generate high-quality pixel-level classification results. Since the rich information of ultra-high-resolution remote sensing images is difficult to feed into traditional segmentation model training, we need to inject the information of the pre-trained visual language model as prototype representation into the traditional segmentation model to promote the performance improvement of the semantic segmentation model. Summary of the invention

[0007] In view of the above, the present invention provides a semantic segmentation method for ultra-high resolution remote sensing images based on balanced anchor point sampling and prototype representation, which proposes a new multi-size anchor point region cropping method, which can better capture multi-scale contextual features, and sample according to the principle of minority category priority and high category richness priority, and at the same time can combine the pre-trained visual language model with the traditional segmentation model to perform detailed semantic segmentation of ultra-high resolution remote sensing images, thereby improving the performance and robustness of the semantic segmentation model of ultra-high resolution remote sensing images.

[0008] A semantic segmentation method for ultra-high resolution remote sensing images based on balanced anchor point sampling and prototype representation includes the following steps:

[0009] (1) Preprocess and enhance the original ultra-high resolution image to obtain the enhanced image I aug ;

[0010] (2) For image I aug Multi-scale cropping is performed based on the anchor point area to obtain a multi-scale anchor point area cropping group;

[0011] (3) Use the pre-trained target detection model to detect image I aug Conduct detection and generate multiple remote sensing object frames;

[0012] (4) Sort the remote sensing object frames according to the rarity of the main category and the category richness, and mix the multi-scale anchor region cropping group and the remote sensing object frame images as a batch of training samples based on the sorting results;

[0013] (5) Obtain a large number of remote sensing-specific vocabulary descriptions, and obtain a remote sensing text vocabulary library T through classification and deduplication. Fill the vocabulary t in the library into the pre-prepared template to obtain the corresponding text description T t , and then the text description T t Input to the pre-trained visual language model to extract the corresponding prototype representation F t , t∈T;

[0014] (6) Constructing a remote sensing image semantic segmentation model, which includes:

[0015] Visual encoder, used to extract visual features of images;

[0016] Semantic segmentation decoder, which performs semantic segmentation based on the fusion of visual and text features to obtain the classification result of the image;

[0017] (7) The above model is trained using training samples combined with prototype representation. Finally, the ultra-high resolution image to be classified is processed and input into the trained model to predict the output and obtain the corresponding classification result.

[0018] Furthermore, the specific implementation method of step (1) is as follows: firstly, the original ultra-high resolution image is preprocessed, and only the RGB three-channel bands are retained to obtain the processed image I rgb ; Then for image I rgb Data enhancement processing includes random flipping, rotation, color space jitter, image sharpening, and random noise transformation to obtain the enhanced image I aug .

[0019] Furthermore, the specific implementation of step (2) is as follows:

[0020] 2.1 In Image I augA point is randomly selected from the image, and a fixed-size area is cut out with this point as the upper left vertex, namely the anchor area. The anchor area must have at least 2 categories, and the pixel category with the highest proportion cannot be the background and its proportion cannot exceed h (0<h<1) times the total area of ​​the anchor area.

[0021] 2.2 Setting a series of trimming scales 2 ,l 3 ,…,l K ], for any scale l k , the size of the cropped image at this scale is k times the size of the anchor point area, k is a natural number and 2≤k≤K, K is a natural number greater than 2;

[0022] 2.3 According to the scale l k The cropped image size is in image I aug The search is performed with a horizontal step size of (k-1)crop_w and a vertical step size of (k-1)crop_h. The searched image must completely contain the anchor area. crop_h and crop_w are the height and width of the anchor area, respectively. Then, a random image is selected from the searched images as the scale l. k The cropped image I lk ; The search process is to traverse the image according to the above step size, record the image blocks containing the anchor point area, and obtain the searched image. The specific number of searched images varies with the size of the original image;

[0023] 2.4 According to steps 2.1 to 2.3, a multi-size anchor region cropping group is obtained, which includes the anchor region image I l1 and cropped images at multiple scales[I l2 ,I l3 ,…,I lk ].

[0024] Furthermore, in step (3), an instance segmentation model (such as SAM-HQ [Ke L, Ye M, Danelljan M, et al. Segment anything in high quality [J]. Advances in Neural Information Processing Systems, 2024, 36]) is used to detect image I aug The instances in the example are selected, and the threshold is set to filter out the instances with higher confidence. A series of remote sensing object frames are generated according to the instance mask. Each remote sensing object frame is represented by a four-tuple (w, h, c x ,c y ), where w and h are the width and height of the remote sensing feature frame, respectively, and c x ,c yis the center point coordinate of the remote sensing object frame.

[0025] Furthermore, the specific implementation of step (4) is as follows:

[0026] 4.1 Count the pixel proportion of each category in the original ultra-high resolution image according to the labels. Each remote sensing feature frame has a main category and several subcategories. The main category is the category with the highest pixel proportion in the remote sensing feature frame.

[0027] 4.2 Sort all remote sensing feature frames according to the rarity of the main category and the richness of the category, giving priority to comparing the rarity of the main category, that is, the remote sensing feature frame with a smaller proportion of pixels in the original ultra-high resolution image is ranked higher. When the ranking is the same, compare the richness of the category, that is, the remote sensing feature frame with more categories is ranked higher;

[0028] 4.3 Determine the number of images of the multi-size anchor region cropping group and the remote sensing object frame to be mixed according to a certain ratio, select a corresponding number of remote sensing object frames according to the sorting results, mix the images of these remote sensing object frames with the images in the multi-size anchor region cropping group as training samples of a batch and unify the image sizes through resampling, wherein the higher the sorting of the remote sensing object frame, the greater the probability of being sampled.

[0029] Furthermore, in step (5), according to the category of vocabulary t, t is filled into the template to obtain the corresponding text description T t , the template is represented as a satellite image of{}, and then T t After the word segmentation is performed by a word segmenter (such as CLIP), the text is converted from a string to an integer encoding and a series of special characters and padding masks are added to obtain a text prototype; finally, the text prototype is input into a pre-trained visual language model (such as GeoRSCLIP) to extract the corresponding prototype representation F t .

[0030] Furthermore, the visual encoder in step (6) uses a backbone network (such as Segformer-mit-b5 [XieE, Wang W, Yu Z, et al. SegFormer: Simple and efficient design for semantic segmentation with transformers [J]. Advances in neural information processing systems, 2021, 34: 12077-12090]) to extract features from image information. The input of the backbone network is an image sequence mixed with a multi-scale anchor region cropping group and a remote sensing object frame. After feature extraction and alignment of the channel dimension of the mapping layer, a multi-scale feature map, i.e., the visual feature of the image, is obtained; and then the multi-scale feature map is combined with the prototype representation F t Pair and merge to get feature F it , and then F it After being concatenated with the multi-scale feature map, it is input into the semantic segmentation decoder for semantic segmentation.

[0031] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the ultra-high resolution remote sensing image semantic segmentation method based on balanced anchor point sampling and prototype representation.

[0032] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the ultra-high-resolution remote sensing image semantic segmentation method based on balanced anchor point sampling and prototype representation.

[0033] The present invention provides a new training method for ultra-high resolution image segmentation. In the process of training the semantic segmentation model, the present invention uses multi-size anchor region cropping to provide a multi-size field of view of a region. Furthermore, in order to balance the object categories, we use object detection to provide object areas for category sampling as part of the input for model training.

[0034] In addition, in order to inject prototype information into the semantic segmentation model, the present invention proposes a new information injection strategy, including remote sensing-specific feature vocabulary integration, remote sensing field-specific pre-trained visual language model extraction prototype representation and feature fusion module; in the training process of the semantic segmentation model, the present invention keeps the model training loss as the cross entropy loss, selects the model with the best training score in multiple rounds of training, and experimentally verifies that the model with the best score performs best on ultra-high resolution images. Therefore, based on the remote sensing field-specific pre-trained visual language model, the present invention achieves higher performance through text information injection and reselection of training images, greatly improving the accuracy and robustness of remote sensing ultra-high resolution image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Schematic diagram of the steps of the ultra-high resolution remote sensing image semantic segmentation method of the present invention.

[0036] Figure 2 It is a schematic diagram of a specific system implementation flow in an embodiment of the method of the present invention. DETAILED DESCRIPTION

[0037] In order to describe the present invention more specifically, the technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0038] like Figure 1 and Figure 2 As shown, the ultra-high resolution remote sensing image semantic segmentation method based on balanced anchor point sampling and prototype representation of the present invention includes the following steps:

[0039] (1) Preprocess the original ultra-high resolution image I and retain only the RGB channels to obtain the processed image I rgb .

[0040] Remote sensing satellite images may usually contain multiple bands. For example, Gaofen-2 satellite images contain not only the RGB band of traditional images, but also the near-infrared band; Sentinel-2 satellite images contain not only the RGB band, but also the near-infrared, short-infrared and other bands.

[0041] Since the mainstream processing method of current computer vision networks is to process in RGB format, and adding non-RGB bands to the existing model will seriously affect the generalization ability of the model, this implementation method needs to select the band of the image in the data processing stage and convert the image into RGB format. rgb .

[0042] (2) For image I rgbPerform data enhancement processing. Optional data enhancement strategies include random flipping, rotation, color space jitter, random cropping, image sharpening, random noise transformation, etc. to obtain the enhanced image I aug .

[0043] In order to improve the generalization ability of the model and increase the diversity of input data, we usually need to perform some data enhancement operations on the image. For remote sensing images, random flip (random_flip) and rotation (rotate) operations are of great significance, because remote sensing objects usually require the model to learn rotation invariance properties; common data preprocessing operations also include image sharpening, random noise transformation, etc. In order to verify the versatility of the method, this implementation only uses random_flip and rotate operations. The specific operation results are as follows:

[0044] I aug =rotate(random_flip(I rgb ))

[0045] The expression of random_flip is as follows, where c represents the color channel of the image. First, a random number p between 0 and 1 is generated, corresponding to the random operation.

[0046]

[0047] The expression for rotate is as follows, where c represents the color channel of the image:

[0048]

[0049] Where: (x 0 ,y 0 ) is the coordinate of the rotation origin, and θ is the rotation angle.

[0050] (3) Using the designed multi-size anchor region cropping method, based on the crop_size scale, aug Perform multi-scale cropping to obtain a multi-scale anchor region cropping group, which contains images of k scales. Perform the same operation on the labels.

[0051] The present invention performs a series of optimizations for adapting ultra-high resolution images for random cropping, namely, multi-size anchor point region cropping, specifically:

[0052] Since the annotation cost of ultra-high resolution remote sensing images is very high, expert annotators usually only select some highly concerned objects for pixel-by-pixel annotation, and annotate a large number of unconcerned areas as background categories. Therefore, after determining the size of the cropping area (crop_h, crop_w), it is necessary to carefully select an anchor area. We select our initial anchor area by restricting the cropping area to have at least 2 categories, and the category with the highest proportion cannot be the background class, and the proportion cannot exceed 75% of the cropping area. Since this area maintains the ground sampling example of the image, it is particularly important for the model to capture details. We call it I l1 , other scales are cut by I l1 to obtain.

[0053] This embodiment specifies a series of cropping scales for cropping the image. lk Compared to I l1 The cropped image is k times larger than the original image size. The cropping method for each scale is as follows:

[0054]

[0055] Among them: crop_h, crop_w are I l1 The size of multi_crop_h and multi_crop_w are I lk In this embodiment, crop_h=crop_w=512.

[0056] In order to meet the above conditions, it is necessary to select an appropriate clipping step size for each scale, and the clipping step size should satisfy:

[0057] stride_h∈[crop_h,(k-1)*crop_h]; stride_w∈[crop_w,(k-1)*crop_w]

[0058] In the specific implementation process, we take stride_h = (k-1)*crop_h, stride_w = (k-1)*crop_w, take 4 levels, k∈[1,2,3,4], for I aug Implement cropping to obtain cropped multi-size images [I l1 ,I l2 ,I l3 ,I l4 ].

[0059] (4) Use the target detection model to obtain the complete remote sensing object frame.

[0060] In recent years, with the rapid development of computer vision in the field of remote sensing, many remote sensing object detection models have emerged, such as Rotated Faster RCNN[Yang S, Pei Z, Zhou F, et al. Rotated faster R-CNN for oriented object detection in aerial images[C] / / Proceedings of the 2020 3rdInternational Conference on Robot Systems and Applications.2020:35-39], RVSA[Wang D, Zhang Q, Xu Y, et al. Advancing plain vision transformer toward remotesensing foundation model[J]. IEEE Transactions on Geoscience and RemoteSensing, 2022, 61:1-15], Oriented RCNN[Xie X, Cheng G, Wang J, et al. Oriented R-CNNfor object detection[C] / / Proceedings of the IEEE / CVF international conferenceon computer vision.2021:3520-3529] models, etc. However, most of these models are aimed at aerial images with low shooting altitudes or satellite images with relatively few objects, and are not suitable for detecting complete objects. The target detection model used in the present invention should be able to detect some complete objects in the original ultra-high resolution remote sensing images to a certain extent to ensure the integrity of the training area; with the emergence of the SAM model in the instance segmentation field, it is possible to obtain boxes by relying on masks, and the SAM series models are naturally more suitable for processing high-resolution images. This embodiment uses the SAM-HQ model to generate multiple detection boxes, each box consists of a four-tuple (w, h, c x ,c y ), the specific mathematical expression is:

[0061] (w,h,c x ,c y )=f θ (I aug )

[0062] Where: f θ () is the mapping function of the SAM-HQ neural network, w, h, c in the quaternionx ,c y Respectively represent the width, height and center point coordinates of the detection box.

[0063] (5) Sort the detected complete remote sensing object frames.

[0064] The present invention needs to know the category distribution of the entire data set in advance, and obtain the proportion of the number of pixels of each category in the data set. The smaller the proportion, the higher the category rarity. Since the result of target detection is usually a complete instance, each detection frame first has a main category and several secondary categories. This implementation sets two attributes for each detection frame: main category rarity and category richness (i.e., the number of category types). All detection frames are sorted according to the main category rarity. The rarer the main category, the higher the ranking of the detection frame. When the main category rarity is the same, the higher the category richness, the higher the ranking. In this way, balanced category sampling is achieved. During the training process, the higher the ranking of the detection frame, the greater the probability of it being used as training data. Finally, we obtain a series of images cut out according to the detection frame [I box1 ,I box2 ,…,I boxn ].

[0065] (6) The cropped image and the remote sensing object frame image are mixed in a certain ratio. In this embodiment, we take a set of multi-scale anchor point region cropped images [I l1 ,I l2 ,I l3 ,I l4 ], and then select four detection boxes [I box1 ,I box2 ,I box3 ,I box4 ], in the selection process, samples with high category rarity and high category richness are given priority to ensure that the top-ranked detection boxes are more likely to be selected, thus obtaining a batch of training samples:

[0066] I train =[I l1 ,I l2 ,I l3 ,I l4 ,I box1 ,I box2 ,I box3 ,I box4 ]

[0067] For the scale of the detection box, if the size of the detection box itself is larger than (crop_h, crop_w), crop the image with a scale of (crop_h, crop_w) from the upper left corner as training data; if the size of the detection box itself is less than (crop_h, crop_w), resample the detection box size to (crop_h, crop_w) as training data.

[0068] (7) The training data I train The feature extraction is sent to the visual encoder based on the backbone network to obtain the features of k stages [F o1 ,F o2 ,…F ok ], the above features are mapped to the aligned dimensions through k mapping layers to obtain a multi-scale feature map [F 1 ,F 2 ,…F k ], used for the final classification of land features.

[0069] Commonly used backbone networks include Resnet[He K, Zhang in neural information processingsystems,2021,34:12077-12090], Mobilenet[Howard AG. Mobilenets: Efficientconvolutional neural networks for mobile vision applications[J].arXivpreprint arXiv:1704.04861,2017], SwinTransformer[Liu Z,Lin Y,Cao Y,et al.Swintransformer:Hierarchical vision transformer using shifted windows[C] / / Proceedings of the IEEE / CVF international conference on computer vision.2021:10012-10022], etc. In this implementation, the backbone network is implemented using Segformer-mit-b5, but is not limited to the Segformer-mit-based method. The Segformer-mit backbone network is a semantic segmentation network backbone architecture based on Transformer, which accepts RGB three-channel image input and outputs feature maps of 4 different scales. Compared with the downsampling scale of the original image, In the implementation process of the Segformer-mit architecture, Efficient Self-Attn and Mix-FFN modules are used to extract features, which can reduce the computational complexity. Overlap PatchMerging is used for feature aggregation and downsampling in each stage. After obtaining the feature maps of multiple stages, the model aligns the number of channels through four 1×1 convolution modules.

[0070] This embodiment uses an input image of size (512, 512) to obtain a multi-scale feature map [F 1 ,F 2 ,…F k ].

[0071] (8) Obtain a series of remote sensing-specific vocabulary descriptions of land objects, map these vocabulary to the categories of each data set to enrich the category information, and obtain the remote sensing text vocabulary library T after integration.

[0072] This implementation method makes a detailed text description and division for each subject category, obtains a varying number of subcategories as shown in Table 1, and provides them to a pre-trained visual language model dedicated to the remote sensing field as text prompts.

[0073] Table 1

[0074]

[0075]

[0076] (9) Use text information and pre-trained visual language models to obtain text features.

[0077] The pre-trained visual language model in the remote sensing field selected in this implementation is GeoRSCLIP. GeoRSCLIP uses OpenCLIP (https: / / github.com / mlfoundations / open_clip.git) weights to initialize the weights of the CLIP model. It is trained on the large-scale remote sensing image-text pair dataset RS5M [Zhang Z, Zhao T, Guo Y, et al. RS5M and GeoRSCLIP: A large scale vision-language dataset and a large vision-language model for remote sensing [J]. IEEE Transactions on Geoscience and Remote Sensing, 2024] and can extract text features with a large amount of remote sensing knowledge. After obtaining the text vocabulary library T in the previous step, each vocabulary description in the library is brought into the CLIP [Radford A, Kim JW, Hallacy C, et al. Learning transferable visual models from natural language supervision [C] / / International conference on machine learning. PMLR, 2021: 8748-8763] standard template "asatellite image of {}" to obtain multiple text prototypes. The above N templates are passed through the CLIP tokenizer to obtain the tokens, and the text is converted from a string to an integer encoding and a series of special characters and padding masks are added to obtain the text token F_token and sent to the text branch of GeoRSCLIP to obtain the prototype representation F t .

[0078] (10) Multi-scale image features [F 1 ,F 2 ,F 3 ,F 4 ] and prototype representation F t Send it to the image text feature fusion module to get the fused feature F it , F it and multi-scale image features [F 1 ,F 2 ,F 3 ,F 4 ] are concatenated and sent to the semantic segmentation decoder to obtain the classification result.

[0079] Since we do not have the corresponding text label for a single image, the text features are kept unchanged in this implementation. After aligning the dimensions of the image features and the text features, the cosine similarity is calculated to obtain the pairing probability of each image feature and each text feature.

[0080] In this embodiment, firstly [F 1 ,F 2 ,F 3 ,F 4 ] The features are resampled and aligned to obtain multi-scale fusion features F if , after a layer of MLP and F t Align the dimensions, then normalize the feature vectors to a range of 0 to 1, and finally calculate the cosine similarity of each vector to obtain the final fusion feature F with a shape of [b, n, h, w]. it , where b is the batch size for training, n is the number of texts, h and w represent the height and width of the feature map respectively; finally, F if and F it The final feature F is obtained by splicing along dimension one and input into the semantic segmentation decoder. The whole feature fusion step is formulated as follows:

[0081] F if =concate([F 1 ,resized(F 2 ),resized(F 3 ),resized(F 4 )], dim = 0)

[0082] F it =BMM(F t ,MLP(F if ) T )

[0083] F=concate([F if ,F it ], dim=1)

[0084] Among them: resized represents bilinear interpolation, concate represents concatenation along the specified dimension, MLP is a two-layer fully connected network, and BMM can be expressed by the following formula. These three operations can be implemented by calling the torch library.

[0085] X′=MLP(X)=(X T W 1 +b 1 ) T W 2 +b 2

[0086] Pn×h×w =BMM(M c×h×w ,N c×n )

[0087]

[0088] (11) Use cross entropy loss to train the model and obtain the final semantic segmentation model.

[0089] The training data used in this embodiment are the public dataset of ultra-high resolution remote sensing images from the GF-2 satellite GID and the dataset URUR from multi-source satellite images [Ji D, Zhao F, Lu H, et al. Ultra-highresolution segmentation with ultra-rich context: A novel benchmark [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023: 23621-23630], where GID [Tong XY, Xia GS, Lu Q, et al. Land-cover classification with high-resolution remote sensing images using transferable deep models [J]. Remote Sensing of Environment, 2020, 237:111322] The image size of the dataset is (6800, 7200), the training set has 120 images, and the validation set has 30 images; the image size of the URUR dataset is (5120, 5120), and the training set, validation set, and test set have 2157, 280, and 571 images, respectively. The model stops after training 160,000 iters, and saves a model checkpoint (complete model parameters) every 8,000 iters (a forward propagation and backpropagation process during training); the training optimizer uses the AdamW optimizer, the learning rate is set to 1e-6, betas is [0.9, 0.999], weight_decay is set to 0.01, and the learning rate decay strategy uses linear decay (LinearLR) in the first 1500 iters, and polynomial decay (PolyLR) is used in the subsequent training process.

[0090] The loss function uses K-type cross entropy loss, and the formula is as follows:

[0091]

[0092] Where: p ic is the probability that the i-th sample is classified as the c-th category, K refers to the number of categories, and N refers to the total number of samples.

[0093] Table 2 shows the performance gain of this implementation compared to the basic Segformer-b5 model on the URUR and GID datasets. In the table, mc represents multi_level_crop (mlrc), samhq represents the detection box obtained by using the SAM-HQ model for target detection, all of which are resampled to a fixed scale and sent to the model for training; keepgsd refers to the detection box obtained by using the SAM-HQ model for target detection, and the detection box that does not meet the fixed size is clipped from the upper left corner to the fixed size model for training; clip-text indicates whether to add the text features extracted by GeoRSCLIP, × indicates that this method is not used, and √ indicates that this method is used on the Segformer-b5 infrastructure.

[0094] Table 2

[0095]

[0096] From the above experimental results, it can be seen that this implementation has achieved a significant performance improvement compared to the basic model of Segformer-b5, and after applying the SAMHQ model to obtain the detection frame, the detection frame is clipped and the ground sampling distance is maintained, which can further promote the improvement of model performance; perhaps because as an ultra-high-resolution segmentation method, any form of resampling will cause pixel loss in the input image, resulting in suboptimal segmentation results. After adding the text information extracted by GeoRSCLIP, the performance of the model is greatly improved, indicating that the pre-trained visual language model can inject difficult-to-learn knowledge into the semantic segmentation model, thereby promoting the model to learn richer features.

[0097] Table 3 shows the IoU accuracy of the present embodiment on a single category, which further demonstrates the effectiveness of the present invention, especially the text feature has a great effect on promoting the segmentation accuracy of fewer categories in the data set.

[0098] Table 3

[0099] Data unlabel building framland forest meadow water road greenhouse Bareland URUR 1.06 72.4 75.91 47.77 - 51.51 49.18 45.19 33.15 GID - 78.91 80.87 76.49 54.7 90.62 - - -

[0100] Table 4 shows the comparison of the performance of the present embodiment and other methods in the art. It can be seen from the experimental results that the performance of the method of the present invention on the three data sets exceeds that of the existing methods, demonstrating the superiority of the method of the present invention.

[0101] Table 4

[0102]

[0103] The above description of the embodiments is to facilitate the understanding and application of the present invention by those skilled in the art. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the present invention is not limited to the above embodiments. Improvements and modifications made by those skilled in the art to the present invention based on the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A semantic segmentation method for ultra-high resolution remote sensing images based on balanced anchor point sampling and prototype representation, comprising the following steps: (1) Preprocess and enhance the original ultra-high resolution image to obtain the enhanced image I aug ; (2) For image I aug Multi-scale cropping is performed based on the anchor point area to obtain a multi-scale anchor point area cropping group; (3) Use the pre-trained target detection model to detect image I aug Conduct detection and generate multiple remote sensing object frames; (4) Sort the remote sensing object frames according to the rarity of the main category and the category richness, and mix the multi-scale anchor region cropping group and the remote sensing object frame images as a batch of training samples based on the sorting results; (5) Obtain a large number of remote sensing-specific vocabulary descriptions, and obtain a remote sensing text vocabulary library T through classification and deduplication. Fill the vocabulary t in the library into the pre-prepared template to obtain the corresponding text description T t , and then the text description T t Input to the pre-trained visual language model to extract the corresponding prototype representation F t , t∈T; (6) Constructing a remote sensing image semantic segmentation model, which includes: Visual encoder, used to extract visual features of images; Semantic segmentation decoder, which performs semantic segmentation based on the fusion of visual and text features to obtain the classification result of the image; (7) The above model is trained using training samples combined with prototype representation. Finally, the ultra-high resolution image to be classified is processed and input into the trained model to predict the output and obtain the corresponding classification result.

2. The method for semantic segmentation of ultra-high resolution remote sensing images according to claim 1, characterized in that: The specific implementation method of step (1) is as follows: firstly, the original ultra-high resolution image is preprocessed to retain only the RGB three-channel bands to obtain the processed image I rgb ; Then for image I rgb Data enhancement processing includes random flipping, rotation, color space jitter, image sharpening, and random noise transformation to obtain the enhanced image I aug .

3. The method for semantic segmentation of ultra-high resolution remote sensing images according to claim 1, characterized in that: The specific implementation of step (2) is as follows: 2.1 In Image I aug A point is randomly selected from the image, and a fixed-size area is clipped with this point as the upper left vertex, namely the anchor area. The anchor area must have at least 2 categories, and the pixel category with the highest proportion cannot be the background and its proportion cannot exceed h times the total area of ​​the anchor area, where h is a real number and 0<h<1; 2.2 Set a series of clipping scales [l2,l3,…,l K ], for any scale l k , the size of the cropped image at this scale is k times the size of the anchor point area, k is a natural number and 2≤k≤K, K is a natural number greater than 2; 2.3 According to the scale l k The cropped image size is in image I aug The search is performed with a horizontal step size of (k-1)crop_w and a vertical step size of (k-1)crop_h. The searched image must completely contain the anchor area. crop_h and crop_w are the height and width of the anchor area, respectively. Then, a random image is selected from the searched images as the scale l. k The cropped image I lk ; 2.4 According to steps 2.1 to 2.3, a multi-size anchor region cropping group is obtained, which includes the anchor region image I l1 and cropped images at multiple scales[I l2 ,I l3 ,…,I lK ].

4. The method for semantic segmentation of ultra-high resolution remote sensing images according to claim 1, characterized in that: In step (3), the instance segmentation model is used to detect the image I aug The instances in the example are selected, and the threshold is set to filter out the instances with higher confidence. A series of remote sensing object frames are generated according to the instance mask. Each remote sensing object frame is represented by a four-tuple (w, h, c x ,c y ), where w and h are the width and height of the remote sensing feature frame, respectively, and c x ,c y is the center point coordinate of the remote sensing object frame.

5. The method for semantic segmentation of ultra-high resolution remote sensing images according to claim 1, characterized in that: The specific implementation of step (4) is as follows: 4.1 Count the pixel proportion of each category in the original ultra-high resolution image according to the labels. Each remote sensing feature frame has a main category and several subcategories. The main category is the category with the highest pixel proportion in the remote sensing feature frame. 4.2 Sort all remote sensing feature frames according to the rarity of the main category and the richness of the category, giving priority to comparing the rarity of the main category, that is, the remote sensing feature frame with a smaller proportion of pixels in the original ultra-high resolution image is ranked higher. When the ranking is the same, compare the richness of the category, that is, the remote sensing feature frame with more categories is ranked higher; 4.3 Determine the number of images of the multi-size anchor region cropping group and the remote sensing object frame to be mixed according to a certain ratio, select a corresponding number of remote sensing object frames according to the sorting results, mix the images of these remote sensing object frames with the images in the multi-size anchor region cropping group as training samples of a batch and unify the image sizes through resampling, wherein the higher the sorting of the remote sensing object frame, the greater the probability of being sampled.

6. The method for semantic segmentation of ultra-high resolution remote sensing images according to claim 1, characterized in that: In step (5), according to the category of vocabulary t, t is filled into the template to obtain the corresponding text description T t , the template is represented as a satellite image of {}, and then T t After the word segmentation, the text is converted from a string to an integer encoding and a series of special characters and padding masks are added to obtain the text prototype; finally, the text prototype is input into the pre-trained visual language model to extract the corresponding prototype representation F t .

7. The method for semantic segmentation of ultra-high resolution remote sensing images according to claim 1, characterized in that: The visual encoder in step (6) uses a backbone network to extract features from image information. The input of the backbone network is an image sequence mixed with a multi-scale anchor region cropping group and a remote sensing object frame. After feature extraction and alignment of the channel dimension at the mapping layer, a multi-scale feature map, i.e., the visual features of the image, is obtained. The multi-scale feature map is then combined with the prototype representation F t Pair and merge to get feature F it , and then F it After being concatenated with the multi-scale feature map, it is input into the semantic segmentation decoder for semantic segmentation.

8. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: The processor is used to execute the computer program to implement the ultra-high resolution remote sensing image semantic segmentation method as claimed in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for semantic segmentation of ultra-high resolution remote sensing images as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • High-resolution remote sensing image classification method based on residual network and transfer learning

    CN112836614A

  • Remote sensing scene recognition method and system based on multiple modes and medium

    CN116665114A

  • Open vocabulary semantic segmentation method and device based on three-dimensional Gaussian scene

    CN118887665A

  • Image segmentation method and apparatus, computer device, and storage medium

    US20210166395A1