Remote sensing small sample semantic segmentation method and system based on multi-scale prototype fusion
Through the combination of multi-scale prototype fusion and co-attention module, the problem of insufficient utilization of scale differences and correlation information in semantic segmentation of remote sensing small samples is solved, and the segmentation accuracy and accuracy of different scale targets in remote sensing images are improved.
Patent Information
- Application Number
- CN202510508078.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
AI Technical Summary
When the existing remote sensing small sample semantic segmentation method is too large to extract high-quality prototypes from the support set for segmentation of the query set, and lacks the discovery of correlation information between the support set and the query set, which leads to the inability to fully utilize the geographical knowledge contained in the support set to guide the segmentation of the query set.
A multi-scale prototype fusion network is used to perform weighted average pooling operations on multi-scale support features. The similarity matrix and attention weight between the support set and the query set are calculated through the co-attention module, and the co-attention results are generated, and the semantic segmentation results are generated with the query feature stitching input decoder.
The segmentation accuracy of the small sample semantic segmentation model for different scale targets can be improved, and the target scale of the support set is too large, and the geographical knowledge of the support set is used to guide the segmentation of the query set, which improves the accuracy of the segmentation results.
Smart Images

Figure CN120339627A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of remote sensing sample semantic segmentation, and particularly relates to a remote sensing small sample semantic segmentation method and system based on multi-scale prototype fusion. Background Art
[0002] Semantic segmentation of remote sensing images is an important task in the fields of computer vision and remote sensing technology, aiming to classify pixels in optical remote sensing images, so as to divide different regions in the images into categories with specific semantic meanings. Remote sensing images contain rich details of ground objects, but the pixels in the original images only have spectral information and lack clear semantic interpretations. Semantic segmentation technology assigns corresponding semantic labels to each pixel according to the differences in features such as spectrum, texture, and shape of different ground objects in the image, such as classifying the pixels in the image into different categories such as buildings, roads, vegetation, water bodies, bare land, etc. With the development of artificial intelligence, deep learning methods have entered the field of semantic segmentation of optical remote sensing images. Coupled with the increasing availability of remote sensing data, it has made more effective scene understanding and object recognition possible. The current mainstream remote sensing image semantic segmentation models mainly use CNN or ViT as the main network architecture, and usually have high requirements for the sample size of training data. However, due to the wide coverage of remote sensing images and the involvement of many complex ground object categories, accurate annotation of remote sensing images requires annotators to have certain geoscience professional knowledge, which makes it extremely difficult and expensive to obtain a large number of high-quality remote sensing image annotation samples. At the same time, in the face of a specific scene, since the categories to be segmented change, the training set also needs to be reconstructed, and the categories that have not been annotated before need to be specifically annotated, which greatly increases the cost of applying the model to different tasks.
[0003] In this context, applying the method of few-shot learning to semantic segmentation of remote sensing images has become a key approach to solving the problem of rapid network transfer between different tasks. Few-shot learning aims to enable the model to quickly learn and adapt to new tasks or categories with only a small number of training samples (usually far fewer than those required by traditional methods). Few-shot learning attempts to simulate the rapid learning ability of humans, extract key information from limited examples, and then make accurate predictions for unseen samples. As a specific application of few-shot learning in the field of semantic segmentation, few-shot semantic segmentation has important theoretical and practical values. Theoretically, it promotes the development of machine learning models towards a direction closer to human intelligence, exploring how to achieve rapid knowledge transfer and application with extremely little information. In practice, few-shot semantic segmentation provides possible solutions for many fields where data acquisition is difficult. For example, in the rapid scene assessment after natural disasters, only a small number of pre-disaster or images of the same type of disasters may be available. By implementing few-shot semantic segmentation, an effective segmentation model can be quickly constructed in these scenarios, providing key support for subsequent decision-making and actions, thereby expanding the application scope of semantic segmentation technology and improving the adaptability and flexibility of the system. Therefore, few-shot semantic segmentation has unique applicability and advantages in the field of remote sensing. It is very necessary to conduct research on few-shot semantic segmentation in the field of remote sensing to make full use of a large amount of high-quality, unlabeled remote sensing data.
[0004] As an improvement, in Chinese Patent CN202411429175.4, a small-sample semantic segmentation method and system based on inter-class relationships discloses a method that constructs a cross-support set and query set feature aggregation network to obtain aggregated features, and uses an extended transformer to increase the inter-class gap and reduce the intra-class gap to implement pixel-level dense matching. This method makes full use of the correlation between the support set and the query set to supplement prior knowledge and can effectively improve the segmentation accuracy of the query set. However, since semantic segmentation relying only on aggregated features cannot perceive the scale differences between the support set and the query set, the performance of this method is relatively limited when processing remote sensing images. Chinese Patent CN202411285200.6, a small-sample remote sensing image semantic segmentation system based on multi-level feature aggregation and loss weighting, discloses a small-sample remote sensing image semantic segmentation system based on multi-level feature aggregation and loss weighting. The multi-level feature extraction module generates features of multiple levels and scales for the system, providing more segmentation information to adapt to target objects of different sizes and rich information for the final accurate segmentation. The accuracy loss weighting module uses the segmentation accuracy and intersection over union as criteria to calculate the loss weights to balance the training in the case of multiple segmentation tasks and improve the segmentation performance of the model. Although this system uses multi-level aggregated features for the final segmentation, it fuses the segmentation results of different-level features at the segmentation result level instead of fusing the multi-level features. On the one hand, this greatly increases the computational complexity of the segmentation process and also makes some detailed information easy to be lost. To sum up, in the current remote sensing small-sample semantic segmentation methods, when the scale gap between the objects in the support set and the query set is too large, it is difficult to extract high-quality prototypes from the support set for the segmentation of the query set; moreover, before segmentation, the support set and the query set are usually processed independently, lacking the exploration of the correlation information between the support set and the query set, resulting in the inability to fully utilize the geographical knowledge contained in the support set to guide the segmentation of the query set. Summary of the Invention
[0005] The present invention provides a remote sensing small-sample semantic segmentation method and system based on multi-scale prototype fusion, aiming to solve the problems in the current remote sensing small-sample semantic segmentation methods that when the scale gap between the objects in the support set and the query set is too large, it is difficult to extract high-quality prototypes from the support set for the segmentation of the query set; and the lack of exploration of the correlation information between the support set and the query set, resulting in the inability to fully utilize the geographical knowledge contained in the support set to guide the segmentation of the query set.
[0006] To achieve the above object, the present invention adopts the following technical solutions: The present invention provides a remote sensing small-sample semantic segmentation method based on multi-scale prototype fusion, including the following steps: S1. Obtain target support set data and target query set data from remote sensing images; S2. Extract features from the target support set data and the target query set data through a feature extraction network to obtain multi-scale support features and multi-scale query features; S3. Perform weighted average pooling operations on the multi-scale support features through a multi-scale prototype fusion network to obtain multi-scale prototypes, fuse the multi-scale prototypes, and separate the foreground and background prototypes to obtain foreground prototypes and background prototypes containing multi-scale information; S4. Calculate the similarity matrix and attention weights between the target support set data and the target query set data through a co-attention module to generate co-attention results; S5. Concatenate the multi-scale prototypes, co-attention results, and query features, and input the concatenated results into a decoder to generate semantic segmentation results; Among them, the feature extraction network, the multi-scale prototype fusion network, and the co-attention module are formed in an encoder, and the encoder and the decoder are constructed in a few-shot semantic segmentation model.
[0007] In some embodiments, in S2, ResNet-50 is used as the feature extraction network to extract the output of the first stage, the fusion result of the outputs of the second and third stages, and the output of the fourth stage of the residual network part as the multi-scale support features and multi-scale query features.
[0008] Further, in S2, the following formulas (1) and (2) are used to extract the multi-scale support features and multi-scale query features: (1); (2); Wherein, Support set samples, Query set samples, , , And Are respectively the outputs of the first, second, third, and fourth stages of the ResNet-50 residual network part, Multi-scale support features, Multi-scale query features.
[0009] In some embodiments, in S3, the multi-scale prototypes are fused and the foreground and background prototypes are separated through a masked support set label to obtain foreground prototypes and background prototypes containing multi-scale information.
[0010] Further, in S3, the following formulas (3) and (4) are used to fuse the multi-scale prototypes and separate the foreground and background prototypes: (3) (4); Among them, multi-scale prototype, weighted average pooling, prototype fusion and separation operations, background prototype, foreground prototype.
[0011] In some embodiments, S4 specifically includes: Use the support set labels to mask the multi-scale support features, calculate the query and influence for the masked result and the multi-scale query features respectively to obtain the influence of the support features on the query features and the influence of the query features on the support features, and then calculate the similarity matrix and the attention weights; obtain the co-attention result from the attention weights, the influence of the support features on the query features, and the influence of the query features on the support features.
[0012] Further, in S4, the co-attention result is obtained through the following formula (5); (6); (7); (8); Among them, is the influence of the support features on the query features, influence and the influence of the query features on the support features, matrix multiplication, , are query parameters, and are key parameters, normalization exponential function, is the attention weight, tensor concatenation, is the co-attention result.
[0013] In some embodiments, in S5, the semantic segmentation result is obtained through the following formula (9): (9); Among them, is the semantic segmentation result, and Decoder is the decoder.
[0014] In some embodiments, the method is performed through a constructed and trained few-shot semantic segmentation model. Constructing and training the few-shot semantic segmentation model specifically includes: Divide the remote sensing image dataset into a training set and a validation set according to categories; randomly select pictures of the same ground object category from the training set as the support set and the query set respectively; Construct a few-shot semantic segmentation model including an encoder and a decoder. The encoder includes a feature extraction network, a multi-scale prototype fusion network, and a co-attention module; Input the support set and the query set into the few-shot semantic segmentation model to calculate the semantic segmentation results of the query set, calculate the loss according to the query set labels, and optimize the parameters of the few-shot semantic segmentation model according to the calculated loss; Use the validation set to calculate evaluation indicators of the semantic segmentation ability of the few-shot semantic segmentation model, and evaluate the few-shot semantic segmentation performance of the few-shot semantic segmentation model.
[0015] The present invention also provides a remote sensing few-shot semantic segmentation system based on multi-scale prototype fusion. The system includes a data acquisition module, a feature extraction module, a multi-scale prototype fusion module, a co-attention module, and a semantic segmentation module; where: Data acquisition module: used to obtain target support set data and target query set data in remote sensing images; Feature extraction module: used to extract features from the target support set data and the target query set data through a feature extraction network to obtain multi-scale support features and multi-scale query features; Multi-scale prototype fusion module: used to perform weighted average pooling operations on the multi-scale support features through a multi-scale prototype fusion network to obtain multi-scale prototypes, fuse the multi-scale prototypes and separate the foreground and background prototypes to obtain foreground prototypes and background prototypes containing multi-scale information; Co-attention module: used to calculate the similarity matrix and attention weights between the target support set data and the target query set data through the co-attention module to generate co-attention results; Semantic segmentation module: used to splice the multi-scale prototypes, the co-attention results and the query features, and input the splicing results into the decoder to generate semantic segmentation results.
[0016] Compared with the prior art, the remote sensing few-shot semantic segmentation method and system based on multi-scale prototype fusion of the present invention have the following beneficial effects: The remote sensing few-shot semantic segmentation method based on multi-scale prototype fusion of the present invention, compared with the current prior art that only relies on aggregated features for semantic segmentation and is difficult to perceive the scale differences between the support set and the query set, has limited performance when processing remote sensing images. Through multi-scale prototype fusion, support features are extracted and fused from multiple scales to generate foreground and background prototypes containing rich multi-scale information, which can better handle the situation where the target scale differences between the support set and the query set are too large, and effectively improve the segmentation accuracy of the few-shot semantic segmentation model for targets of different scales.
[0017] Although multi-level feature aggregation is adopted in the current existing technologies, the segmentation results of different-level features are fused at the segmentation result level, and the correlation information between the support set and the query set is not fully explored. The present invention introduces a co-attention module to deeply calculate the similarity matrix and attention weights between the features of the support set and the query set, deeply mine the correlation between the two, so as to more effectively use the geographical knowledge in the support set to guide the segmentation of the query set, and can obtain more accurate segmentation results based on the existing technologies, having better practicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings in the specification are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention.
[0019] Figure 1 It is a schematic flow chart of a remote sensing small sample semantic segmentation method based on multi-scale prototype fusion of the present invention; Figure 2 It is a schematic flow chart of multi-scale prototype fusion in a remote sensing small sample semantic segmentation method based on multi-scale prototype fusion of the present invention; Figure 3 It is a schematic flow chart of co-attention result calculation in a remote sensing small sample semantic segmentation method based on multi-scale prototype fusion of the present invention; Figure 4 It is a schematic flow chart of the construction and training of a small sample semantic segmentation model in a remote sensing small sample semantic segmentation method based on multi-scale prototype fusion of the present invention; Figure 5 It is a schematic diagram of the comparison of segmentation results between the present invention and a conventional method on a validation set in an embodiment of a remote sensing small sample semantic segmentation method based on multi-scale prototype fusion of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0021] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0022] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0023] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0024] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.
[0025] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "installed", "connected", "connected" are understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0026] As Figure 1 shown, the remote sensing small sample semantic segmentation method based on multi-scale prototype fusion of the present invention includes the following steps: S1. Obtain target support set data and target query set data in the remote sensing image; S2. Extract features from the target support set data and the target query set data through a feature extraction network to obtain multi-scale support features and multi-scale query features; S3. Perform weighted average pooling operation on the multi-scale support features through a multi-scale prototype fusion network to obtain multi-scale prototypes, fuse the multi-scale prototypes and separate the foreground and background prototypes to obtain foreground prototypes and background prototypes containing multi-scale information; S4. Calculate the similarity matrix and attention weights between the target support set data and the target query set data through a co-attention module to generate a co-attention result; S5. Concatenate the multi-scale prototype, co-attention result and query feature, and input the concatenated result into the decoder to generate the semantic segmentation result; Among them, the feature extraction network, multi-scale prototype fusion network and co-attention module are formed in the encoder, and the encoder and decoder are constructed in the few-shot semantic segmentation model.
[0027] The present invention adopts multi-scale prototype fusion: breaking through the limitations of traditional single scale, extracting and fusing the features of the support set from multiple scales. Using weighted average pooling operation to process the multi-scale support features, and then separating the foreground and background prototypes, enabling the model to effectively capture information at different scales. When facing the problem of large target scale differences between the support set and the query set, it can extract more representative prototypes from the support set and enhance the adaptability of the model to targets at different scales. The present invention also adopts a co-attention module: by masking the features of the support set and the query set, calculating the query representation, influence representation, similarity matrix and attention weights, it realizes in-depth mining of the correlation information between the support set and the query set. This design allows the model to make full use of the geographical knowledge in the support set and provide strong guidance for the segmentation of the query set.
[0028] In addition, the present invention constructs a few-shot semantic segmentation model with an encoder-decoder structure. Among them, the encoder includes a feature extraction network, a multi-scale prototype fusion network and a co-attention module. The present invention provides a few-shot semantic segmentation model with multi-scale prototype fusion and co-attention for remote sensing few-shot semantic segmentation through a series of complete operation processes from data selection, training set and validation set division, to the construction, training and validation evaluation of the few-shot semantic segmentation model. The algorithms involved in the multi-scale prototype fusion and co-attention module of the present invention, such as weighted average pooling, foreground and background prototype separation, query representation and influence representation calculation, similarity matrix and attention weight calculation, etc., effectively improve the segmentation accuracy of the few-shot semantic segmentation model for targets at different scales and can obtain more accurate segmentation results.
[0029] The present invention also provides a remote sensing few-shot semantic segmentation system based on multi-scale prototype fusion. The system includes a data acquisition module, a feature extraction module, a multi-scale prototype fusion module, a co-attention module and a semantic segmentation module; among them: The data acquisition module: is used to obtain the target support set data and target query set data in the remote sensing image; The feature extraction module: is used to extract features from the target support set data and target query set data through the feature extraction network to obtain multi-scale support features and multi-scale query features; Multi-scale prototype fusion module: It is used to perform weighted average pooling operation on multi-scale support features through a multi-scale prototype fusion network to obtain multi-scale prototypes, fuse the multi-scale prototypes and separate the foreground and background prototypes to obtain foreground prototypes and background prototypes containing multi-scale information; Co-attention module: It is used to calculate the similarity matrix and attention weights between the target support set data and the target query set data through the co-attention module to generate co-attention results; Semantic segmentation module: It is used to splice the multi-scale prototypes, co-attention results and query features, and input the splicing results into the decoder to generate semantic segmentation results.
[0030] The following further details the remote sensing few-shot semantic segmentation method and system based on multi-scale prototype fusion of the present invention through specific embodiments.
[0031] As Figures 1-5 shown, Step 1: Data selection and division of training set and validation set: Divide the remote sensing dataset containing 200 high-resolution remote sensing images covering 15 common ground object categories into a training set and a validation set according to categories. Select 10 ground object categories as the training set for training and validate on the validation set containing the other 5 ground object categories.
[0032] Step 2: Construction of few-shot semantic segmentation model: Construct a few-shot semantic segmentation model with an encoder and decoder structure, where the encoder includes: a feature extraction network, a multi-scale prototype fusion network, and a co-attention module. The model generates the semantic segmentation result of the query set through the prior knowledge extracted from the support set and its own parameters.
[0033] Step 3: Few-shot scenario model training: Use the training set generated in Step 1 to construct a few-shot scenario and train the semantic segmentation model to optimize the segmentation ability of the model in the few-shot scenario.
[0034] Step 4: Validation and evaluation: Use the validation set generated in Step 1 to calculate the evaluation metrics for the semantic segmentation ability of the model in the few-shot scenario, and evaluate the few-shot semantic segmentation performance of the model.
[0035] Specifically, the specific process of Step 1 is as follows: Step 1.1: Dataset selection. Select 200 high-resolution remote sensing images from a publicly available labeled dataset as the dataset. This dataset covers 15 common ground object categories. Non-overlapping cropping is performed on the remote sensing images, and they are cropped into 30553 slices of 512×512 pixels.
[0036] Step 1.2: Divide the dataset. Divide the dataset by category, select 10 ground object categories as the training set categories, and the remaining 5 categories as the validation set categories. Make the training set categories and the validation set categories mutually exclusive to ensure that the model has never learned the ground object category features in the validation set, thus constructing a small-sample scenario for validation.
[0037] Further, the specific process of Step 2 is as follows: Step 2.1: Feature extraction. Input the support set S and the query set Q into the feature extraction network respectively to obtain multi-scale support features and query features. In the present invention, ResNet-50 is used as the feature extraction network, and the outputs of the first stage (Layer1), the fusion of the outputs of the second and third stages (Layer2+Layer3), and the output of the fourth stage (Layer4) of the residual network part are extracted as multi-scale features : (1); (2); Among them, support set samples, query set samples, , , and are the outputs of the first, second, third, and fourth stages of the ResNet-50 residual network part respectively, multi-scale support features, multi-scale query features.
[0038] Step 2.2: Multi-scale prototype fusion. Perform weighted average pooling operation on the multi-scale support features obtained in Step 2.1 to obtain multi-scale prototypes, and then fuse the multi-scale prototypes and separate the foreground and background prototypes to obtain foreground and background prototypes containing multi-scale information: (3) (4); Among them, multi-scale prototypes, weighted average pooling, prototype fusion and separation operation, background prototype, foreground prototype.
[0039] Step 2.3: Co-attention result calculation. Use the support set labels to mask the multi-scale support features obtained in Step 2.1, and mask the masked result and the multi-scale query features Perform query representation and influence representation calculations respectively to obtain the influence of the support features on the query features and the influence of the query features on the support features . Then calculate the similarity matrix and attention weights, and finally calculate the co-attention result from the attention weights and the previously obtained influence: (5); (6); (7); (8); wherein, is the influence of the support features on the query features, the influence and the influence of the query features on the support features, matrix multiplication, , is the query parameter, and are the key parameters, normalized exponential function, is the attention weight, tensor concatenation, is the co-attention result.
[0040] Step 2.4, generation of segmentation results. Concatenate the foreground prototype , background prototype obtained in Step 2.2 and the co-attention result obtained in Step 2.3 with the query features, and input the concatenated result into the decoder to obtain the semantic segmentation result : (9); wherein, is the semantic segmentation result, and Decoder is the decoder.
[0041] Specifically, the specific process of Step 3 is as follows: Step 3.1: Construction of support set and query set. Randomly select an image from the training set as the support set image, and then randomly select an image with the same ground object category as the support set image as the query set image.
[0042] Step 3.2: Model training. Input the support set and query set into the few-shot semantic segmentation model to calculate the semantic segmentation result of the query set, then calculate the loss according to the query set label, and finally optimize the model parameters according to the calculated loss.
[0043] The remote sensing small sample semantic segmentation method and system based on multi-scale prototype fusion aims at the problems that there are differences in the target scales between the support set and the query set and it is difficult to explore relevant information. By extracting prototypes of support features at different scales and introducing a co-attention calculation module for support and query pictures, the accuracy of the segmentation result of the query set is significantly improved. It can be applied to the actual scenario of semantic segmentation of ground object categories with a small number of annotations in remote sensing images and has certain popularization value.
[0044] Finally, it should be noted that the above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Any person skilled in the art can smoothly implement the present invention according to the instructions and the above description. Any equivalent changes made by slightly modifying and evolving the technical content disclosed above are equivalent embodiments of the present invention. At the same time, any equivalent changes, modifications, and evolutions made to the above embodiments based on the essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A remote sensing small sample semantic segmentation method based on multi-scale prototype fusion, characterized in that, It includes the following steps: S1. Obtain target support set data and target query set data from remote sensing images; S2. Extract features from the target support set data and target query set data through a feature extraction network to obtain multi-scale support features and multi-scale query features; S3. Perform weighted average pooling operation on the multi-scale support features through a multi-scale prototype fusion network to obtain multi-scale prototypes, fuse the multi-scale prototypes and separate the foreground and background prototypes to obtain foreground prototypes and background prototypes containing multi-scale information; S4. Calculate the similarity matrix and attention weights between the target support set data and the target query set data through a co-attention module to generate co-attention results; S5. Concatenate the multi-scale prototypes, co-attention results and query features, and input the concatenated results into a decoder to generate semantic segmentation results; Among them, the feature extraction network, multi-scale prototype fusion network and co-attention module are formed in the encoder, and the encoder and decoder are constructed in the few-shot semantic segmentation model.
2. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 1, wherein In the S2, ResNet-50 is used as the feature extraction network to extract the multi-scale support features and multi-scale query features, and the output of the first stage, the fusion result of the outputs of the second and third stages, and the output of the fourth stage of the residual network part are extracted as the multi-scale support features and multi-scale query features.
3. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 2, wherein In the S2, the following formulas (1) and (2) are used to extract multi-scale support features and multi-scale query features: (1); (2); Among them, support set samples, query set samples, 、 、 and are the outputs of the first, second, third, and fourth stages of the ResNet-50 residual network part respectively, multi-scale support features, multi-scale query features.
4. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 1, wherein, In the S3, the multi-scale prototypes are fused and the foreground and background prototypes are separated through the mask support set labels to obtain foreground prototypes and background prototypes containing multi-scale information.
5. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 4, characterized in that, In the S3, the multi-scale prototypes are fused and the foreground and background prototypes are separated through the following formulas (3) and (4): (3) (4); Among them, multi-scale prototype, weighted average pooling, prototype fusion and separation operations, background prototype, foreground prototype.
6. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 1, characterized in that, The S4 specifically includes: Use the support set label to mask the multi-scale support features, calculate the query and influence for the masked result and the multi-scale query features respectively to obtain the influence of the support features on the query features and the influence of the query features on the support features, and then calculate the similarity matrix and attention weights; the co-attention results are obtained from the attention weights, the influence of the support features on the query features and the influence of the query features on the support features.
7. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 6, wherein In the S4, the co-attention results are obtained through the following formula (5); (6); (7); (8); Among them, To support the influence of the feature pair on the query feature, The influence of the influence and query feature pair on the support feature, Matrix multiplication, 、 Is the query parameter, And Is the key parameter, Softmax function, Is the attention weight, Tensor concatenation, Is the co-attention result.
8. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 1, characterized in that In the S5, the semantic segmentation results are obtained through the following formula (9): (9); Among them, is the semantic segmentation result, and Decoder is the decoder.
9. The remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to claim 1, characterized in that, The method is carried out through a constructed and trained few-shot semantic segmentation model. The construction and training of the few-shot semantic segmentation model specifically include: Divide the remote sensing image dataset into a training set and a validation set by category; randomly select pictures of the same ground object category from the training set as the support set and the query set respectively; Construct a few-shot semantic segmentation model including an encoder and a decoder. The encoder includes a feature extraction network, a multi-scale prototype fusion network and a co-attention module; Input the support set and the query set into the few-shot semantic segmentation model to calculate the semantic segmentation results of the query set, calculate the loss according to the query set labels, and optimize the parameters of the few-shot semantic segmentation model according to the calculated loss. The semantic segmentation ability of the small-sample semantic segmentation model is used to calculate evaluation indicators, and the small-sample semantic segmentation performance of the small-sample semantic segmentation model is evaluated.
10. The system on which the remote sensing small sample semantic segmentation method based on multi-scale prototype fusion according to any one of claims 1-9 is based, characterized in that, The system includes a data acquisition module, a feature extraction module, a multi-scale prototype fusion module, a co-attention module, and a semantic segmentation module; among them: Data acquisition module: used to obtain target support set data and target query set data in remote sensing images; Feature extraction module: used to extract features from the target support set data and target query set data through a feature extraction network to obtain multi-scale support features and multi-scale query features; Multi-scale prototype fusion module: used to perform weighted average pooling operations on the multi-scale support features through a multi-scale prototype fusion network to obtain multi-scale prototypes, fuse the multi-scale prototypes and separate the foreground and background prototypes to obtain foreground prototypes and background prototypes containing multi-scale information; Co-attention module: used to calculate the similarity matrix and attention weights between the target support set data and the target query set data through the co-attention module to generate co-attention results; Semantic segmentation module: used to splice the multi-scale prototypes, co-attention results and query features, and input the splicing results into a decoder to generate semantic segmentation results.
Citation Information
Patent Citations
Small sample semantic segmentation method and system based on inter-class relationship
CN119380019A
Small sample remote sensing image semantic segmentation system based on multilevel feature aggregation and loss weighting
CN119399456A
Cited By
Feature pool driven few-sample segmentation method and system for large-size image
CN121053658A
A feature pool driven few-shot segmentation method and system for large size images
CN121053658B
Small sample fine granularity classification method and system based on component perception
CN122090182A