Remote sensing image semantic segmentation method and system based on prototype perception and edge enhancement
By adopting prototype perception and edge enhancement methods in semantic segmentation of remote sensing images, the problem of traditional methods relying on a large amount of labeled data and being difficult to distinguish the edges of objects is solved, and high-precision segmentation and detailed boundary recognition are achieved under the limited labeled data.
Patent Information
- Application Number
- CN202510357351.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Traditional deep learning methods rely on a large amount of labeled data in semantic segmentation of remote sensing images, and it is difficult to distinguish the edges of objects when the number of samples is small, resulting in poor segmentation effect.
Using a method based on prototype perception and edge enhancement, features are extracted through Resnet50, multiple prototypes are extracted from supported features using PMMs, and P-F Matching is performed to achieve pixel-level matching. At the same time, edge information is extracted by Sobel operator for feature enhancement, and finally the features are fused and predicted through spatial pyramid pooling.
Improve segmentation accuracy when the labeling data is limited, reduce dependence on large-scale data sets, reduce data set labeling costs, and achieve more refined segmentation in boundary areas, reducing boundary blur problem.
Smart Images

Figure CN120164220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image semantic segmentation, and particularly to a remote sensing image semantic segmentation method and system based on prototype perception and edge enhancement. Background Art
[0002] Remote sensing image semantic segmentation refers to classifying each pixel in a remote sensing image and labeling it as a specific semantic category, such as forest, road or building, etc., which has important applications in fields such as land use, environmental monitoring, and urban planning. With the rapid development of deep learning technology, the semantic segmentation accuracy of remote sensing images has been significantly improved. However, traditional deep learning methods usually rely on a large amount of labeled data to train models, which is particularly difficult and costly for remote sensing images.
[0003] Few-shot segmentation refers to accurately segmenting objects in an image with only a small amount of labeled data. Specifically, the few-shot segmentation task usually uses a small number of support samples to guide the model to segment the query image. This method not only reduces the dependence on large-scale datasets but also can achieve good segmentation results in data-scarce situations. Especially in the field of remote sensing, the high cost and inefficiency of data annotation make few-shot segmentation a very promising solution.
[0004] However, existing traditional few-shot segmentation methods often have certain limitations. Most methods rely on directly obtaining a single prototype from the support image to segment the query image, which can lead to semantic ambiguity and difficulty in distinguishing its category, especially when the number of samples is small. In addition, most methods also ignore the spatial details in the query image, especially the information of object edges. The edge is an important boundary between an object and the background, and it plays a crucial role in object recognition. However, traditional methods often only rely on the overall category features and lack careful attention to object edges, resulting in poor segmentation effects of the model in complex or blurred boundary regions. Summary of the Invention
[0005] To solve the deficiencies of the prior art, the present invention provides a remote sensing image semantic segmentation method and system based on prototype perception and edge enhancement; first, Resnet50 is used as a feature extractor to generate corresponding query and support features. Secondly, PMMs are used to obtain multiple prototypes from the support features, and global matching features are obtained through the P-F Matching module to achieve pixel-level matching between the query features and the support prototypes. Thirdly, the Sobel operator is used to extract the edge information of the query image to enhance the query features, thereby obtaining edge-enhanced features. Finally, the query features, global matching features, and edge-enhanced features are fused together. After the fused features are subjected to spatial pyramid pooling, they are sent to a predictor to obtain the final segmentation result.
[0006] On the one hand, a semantic segmentation method for remote sensing images based on prototype perception and edge enhancement is provided, including: Obtain support images and query images; randomly sample K + 1 images from the remote sensing dataset, where K images are used as support images and one image is used as a query image; the support images are set with target object category labels; the query image has no label; Input the K support images and one query image into the trained remote sensing image semantic segmentation model to obtain the semantic segmentation result of the query image; among them, the trained remote sensing image semantic segmentation model includes: extracting features from the K support images to obtain support features; extracting several prototypes from the support features; extracting features from the query image to obtain query features; performing pixel-level feature matching on the query features based on the prototypes to obtain global matching features; extracting edge information from the query image and enhancing the query features based on the extracted edge information to obtain edge-enhanced features; fusing the query features, global matching features, and edge-enhanced features to obtain fused features, and predicting the fused features to obtain the image segmentation result.
[0007] On the other hand, a semantic segmentation system for remote sensing images based on prototype perception and edge enhancement is provided, including: An acquisition module, which is configured to: obtain support images and query images; randomly sample K + 1 images from the remote sensing dataset, where K images are used as support images and one image is used as a query image; the support images are set with target object category labels; the query image has no label; A segmentation module, which is configured to: input the K support images and one query image into the trained remote sensing image semantic segmentation model to obtain the semantic segmentation result of the query image; among them, the trained remote sensing image semantic segmentation model includes: extracting features from the K support images to obtain support features; extracting several prototypes from the support features; extracting features from the query image to obtain query features; performing pixel-level feature matching on the query features based on the prototypes to obtain global matching features; extracting edge information from the query image and enhancing the query features based on the extracted edge information to obtain edge-enhanced features; fusing the query features, global matching features, and edge-enhanced features to obtain fused features, and predicting the fused features to obtain the image segmentation result.
[0008] The above technical solutions have the following advantages or beneficial effects: Compared with traditional deep learning methods, it can improve the segmentation accuracy under the condition of limited labeled data, reduce the dependence on large-scale data sets, and lower the data set labeling cost. Aiming at the deficiency of traditional few-shot segmentation methods in target boundary processing, the present invention enhances features by using edge prior information, making the segmentation result more refined in the boundary region, reducing the boundary blur problem, improving the classification accuracy, and having broad application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0010] Figure 1 It is a flowchart of the method for the first embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0012] First Embodiment This embodiment provides a remote sensing image semantic segmentation method based on prototype perception and edge enhancement; As Figure 1 shown, the remote sensing image semantic segmentation method based on prototype perception and edge enhancement includes: S101: Obtain support images and query images; randomly sample K + 1 images from the remote sensing data set, where K are used as support images and one is used as a query image; the support images are set with target object category labels; the query image has no label; S102: Input the K support images and one query image into the trained remote sensing image semantic segmentation model to obtain the semantic segmentation result of the query image; wherein, the trained remote sensing image semantic segmentation model includes: extracting support features from the K support images; extracting several prototypes from the support features; extracting query features from the query image; performing pixel-level feature matching on the query features based on the prototypes to obtain global matching features; extracting edge information from the query image and enhancing the query features based on the extracted edge information to obtain edge-enhanced features; fusing the query features, global matching features, and edge-enhanced features to obtain fused features, and predicting the fused features to obtain the image segmentation result.
[0013] Further, the remote sensing data set includes: First, the original remote sensing image is cropped into images with a size of in the way of a sliding window, and divided into a training set and a test set according to a ratio of 7:3. The training set consists of image-mask pairs, and the test set consists of image-mask pairs, where is the RGB image, and is the corresponding segmentation label GroundTruth. The label annotates both the category of the target object and the area where the target object is located.
[0014] Subsequently, all the land cover types in the dataset are divided into two non-overlapping class subsets and . Taking the iSAID dataset as an example, this dataset has a total of 15 categories. 10 of these categories are classified as : ship, storage tank, baseball field, tennis court, basketball court, track and field stadium, bridge, large vehicle, small vehicle, helicopter, and the remaining 5 categories are classified as : swimming pool, roundabout, football field, airplane, port.
[0015] In the training stage of the remote sensing image semantic segmentation model, the categories in are set to be invisible, that is, the pixel points with the category of in the training images are regarded as the background, do not participate in the calculation of the subsequent loss function, and do not trigger any parameter updates or gradient backpropagation regarding these categories.
[0016] In the test stage of the remote sensing image semantic segmentation model, the categories in are set to be visible and participate in the calculation of the final loss function.
[0017] In the training stage of the remote sensing image semantic segmentation model, randomly sample image-mask pairs from the training set to form a set of training data , and each set of training data contains K support image-mask pairs and 1 query image-mask pair .
[0018] In the training stage of the remote sensing image semantic segmentation model, the model learns the mapping relationship of the support image-mask pairs, predicts the segmentation label of the query image , and calculates the cross-entropy loss function Update the weights of the model. After the model training is completed, evaluate the model on the test set.
[0019] In the test stage of the remote sensing image semantic segmentation model, randomly sample from the test set image-label pairs to form a set of test data. Each set of test data contains K support image-label pairs and 1 query image-label pair. The model finally predicts the segmentation label of the query image and compares it with to evaluate the segmentation performance of the model.
[0020] Furthermore, the feature extraction of the K support images to obtain support features is performed using the Resnet50 network as the feature extractor.
[0021] Furthermore, the feature extraction of the query image to obtain query features is performed using the Resnet50 network as the feature extractor.
[0022] It should be understood that using a pre-trained convolutional neural network (CNN) as the feature extractor is specifically manifested as using ResNet50 as the backbone network to extract the features of the query image (Query Image) and the support image (Support Image). The ResNet50 network includes: the first residual block conv2_x, the second residual block conv3_x, the third residual block conv4_x, and the fourth residual block conv5_x connected in sequence. Each residual block contains multiple residual units. Assume the input image size is , where H and W represent the height and width of the image respectively, and 3 is the number of channels. After the input image passes through the ResNet50 network, the output feature of the fourth residual block conv5_x of the ResNet50 network is .
[0023] In order to be able to extract global context information more comprehensively, use the Pyramid Pooling Module (PPM) to perform multi-scale feature integration on the output of the fourth residual block of the ResNet50 network. Specifically, use the pyramid pooling module to perform average pooling operations on the output of the fourth residual block conv5_x at scales of , , respectively, and reduce the channel dimension to one-third through a convolution of . Subsequently, use bilinear interpolation to upsample the pooled feature maps back to the original size, and finally perform a concatenation operation on the feature maps from different scales to obtain the support features of the support image .
[0024] ; Among them, is a splicing operation, is an upsampling operation, is the output of the fourth residual block, represents the pooling output of the corresponding scale.
[0025] Among them, the query feature of the query image is obtained in the same way as the support feature of the support image.
[0026] Furthermore, several prototypes are extracted from the support features, and several prototypes are generated from the support features by using Prototype Mixture Models (PMMs).
[0027] A prototype refers to a representative vector calculated from the features of the support image, and its role is to provide a representative feature of a category for the segmentation model. Associating a category with multiple prototypes can more effectively capture the details in the support image and utilize the support features, enabling the model to compare the similarity with the query image based on these prototypes, thereby identifying and segmenting the target object.
[0028] Furthermore, the process of extracting several prototypes from the support features specifically includes: Assume that the feature map of the support image is , and it is split according to the channel dimension to obtain feature samples with a length of , .
[0029] PMMs is a probability mixture model, and the probability mixture model combines the probabilities of the basic distributions: ; ; ; ; Among them, is the mixing weight, satisfying and , is a parameter that learns with the model, represents the i-th feature sample, represents the k-th probability model based on the kernel distance function, is a parameter that is updated with model training, is the normalization factor, is the Bessel function, is a fixed parameter, set to 20, denotes the mean vector of, that is, the prototype, is the kernel distance function, is the L2 vector norm. By calculating the mean vectors of K probability models based on the kernel distance function, K prototypes are obtained.
[0030] During the training process of the Probabilistic Mixture Models (PMMs), it is optimized according to the Expectation-Maximization (EM) algorithm. Each iteration consists of the following steps: (1-1): Calculate the expected value of each sample belonging to each prototype , according to the sample calculate its probability of belonging to each prototype : ; where, is a fixed parameter, set to 20. (1-2): Update the prototype according to the calculated expected value : ; Iterate repeatedly through steps (1-1) and (1-2) to continuously optimize the estimation of the prototype until convergence. Furthermore, the pixel-level feature matching of the query feature based on the prototype to obtain the global matching feature includes: Assume the query image feature map , for each pixel position (x, y), calculate its feature vector and the cosine similarity between the prototype : ; For each prototype, a total of cosine similarity values are obtained, and a similarity map is constructed accordingly, ; one similarity value is calculated for one pixel point, pixel points form the similarity map; By calculating the similarities between the K prototypes and the query image features, K pixel-level similarity maps are obtained.
[0031] For each pixel position in the query feature, select the prototype with the maximum similarity to match it, and replace each pixel position in the query feature with the prototype with the maximum similarity. Specifically, for each pixel position (x, y), select the prototype with the maximum similarity: ; represents the index of the most similar prototype at the pixel position (x, y); Subsequently, each pixel of the query feature is replaced with its prototype with the highest similarity , thereby constructing a globally matched feature , the globally matched feature has the same size as the query feature, and the feature at each position of the globally matched feature comes from the prototype with the highest similarity .
[0032] The advantages of the above technical solution are: The generated prototypes can be pixel-level matched with the query features to improve the segmentation accuracy.
[0033] Furthermore, the extraction of edge information from the query image includes: Using the Sobel operator to extract the edge information: First, use the Sobel operator to perform edge detection on the query image , and the Sobel operator highlights the edge regions in the image by calculating the gradient of each pixel point; After being processed by the Sobel operator, the obtained processing result is normalized by the Sigmoid activation function to obtain an edge map ; ; where S is the sigmoid activation function.
[0034] Furthermore, the feature enhancement of the query feature based on the extracted edge information to obtain an edge-enhanced feature includes: Performing an element-wise multiplication operation on the edge map and the feature map of the query image to highlight the edge regions and at the same time suppress the features of the background regions: ; where represents the element-wise multiplication operation, is a learnable weight coefficient used to control the influence of the edge features on , is the edge-enhanced feature.
[0035] Furthermore, the feature fusion of the query feature, the globally matched feature, and the edge-enhanced feature to obtain a fused feature includes: Performing fusion on , , in the channel dimension, and passing through The convolution reduces the number of channels, and then the fused features are passed to the Atrous Spatial Pyramid Pooling (ASPP) module to obtain the final feature map. The ASPP module obtains multi-scale information through convolutional layers with different dilation rates, which helps to capture the features of objects at different scales existing in remote sensing images. : ; Further, predicting the fused features to obtain the image segmentation result includes: Sending the feature map into the predictor to obtain the final prediction output; The predictor includes: an upsampling layer and a convolutional layer connected in sequence.
[0036] Further, in S102: inputting both the support image and the query image into the trained remote sensing image semantic segmentation model to obtain the image segmentation result; wherein, the training process of the trained remote sensing image semantic segmentation model includes: Constructing a remote sensing data set, where the remote sensing data set is K support images with known segmentation results and a query image; dividing the remote sensing data set into a training set and a test set according to a ratio; Inputting the training set into the model to train the model. During the training process, using the support image and the query image as the input values of the model, and using the image segmentation result as the output value of the model; When the total loss function value of the model no longer decreases, or when the number of iterations reaches the set number of times, stop training to obtain the trained remote sensing image semantic segmentation model.
[0037] Further, the total loss function of the model includes: ; ; ; ; Among them, is an adjustable constant used to control the influence of the auxiliary loss, represents the total loss function, represents the cross-entropy loss function, represents the auxiliary loss function; represents the total number of N pixels; represents the probability distribution; is the pixel prediction value at position (i, j), and are the height and width of the image, respectively, is a hyperparameter, is a constant.
[0038] Use the cross-entropy loss function to measure the difference between the predicted label and the ground truth label. Assume that the input image has N pixels, and the output of the model is a probability distribution for each pixel belonging to one of K classes .
[0039] To alleviate the problem of scale imbalance in remote sensing images, an auxiliary loss function is used to make the model focus on small objects that are difficult to segment. It is stipulated that the size of the target object in the image is less than pixels is a small object. First, calculate the scale score by computing the probability of each pixel in the prediction mask , and this score is obtained by weighted summation of the predicted probabilities of each pixel. When is small, it indicates that the size of the target object is small, or there is a large deviation between the prediction of the model in this area and the ground truth label. Dynamically adjust the loss weight of each pixel according to the value.
[0040] Among them, is a hyperparameter used to control the degree of attention to small objects. is a very small constant used to avoid numerical instability in numerical calculations when Score is very small. When Score is close to 0 (i.e., indicating a small object), the loss value will increase significantly, promoting the model to pay attention to small objects. Embodiment 2 This embodiment provides a remote sensing image semantic segmentation system based on prototype perception and edge enhancement, including: An acquisition module, which is configured to: acquire support images and query images; randomly sample K + 1 images from a remote sensing dataset, where K are used as support images and one is used as a query image; the support images are set with target object category labels; the query image has no label; A segmentation module, which is configured to: input K support images and a query image into a trained remote sensing image semantic segmentation model to obtain the semantic segmentation result of the query image; wherein, the trained remote sensing image semantic segmentation model includes: extracting features from the K support images to obtain support features; extracting a plurality of prototypes from the support features; extracting features from the query image to obtain query features; performing pixel-level feature matching on the query features based on the prototypes to obtain global matching features; extracting edge information from the query image and enhancing the query features based on the extracted edge information to obtain edge-enhanced features; fusing the query features, the global matching features and the edge-enhanced features to obtain fused features, and predicting the fused features to obtain an image segmentation result.
[0041] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A remote sensing image semantic segmentation method based on prototype perception and edge enhancement, characterized by: include: Obtain support images and query images; Randomly sample K+1 images from a remote sensing dataset, wherein K images are used as support images and one image is used as a query image; the support images are provided with target object category labels; and the query image has no label; K supporting images and one query image are input into a trained remote sensing image semantic segmentation model to obtain a semantic segmentation result of the query image; wherein the trained remote sensing image semantic segmentation model includes: extracting features from the K supporting images to obtain supporting features; extracting a number of prototypes from the supporting features; extracting features from the query image to obtain query features; performing pixel-level feature matching on the query features based on the prototypes to obtain global matching features; extracting edge information from the query image, and enhancing the query features based on the extracted edge information to obtain edge enhanced features; performing feature fusion on the query features, the global matching features and the edge enhanced features to obtain fused features, predicting the fused features to obtain an image segmentation result.
2. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 1, characterized in that: Feature extraction is performed on K support images to obtain support features, using the Resnet50 network as the feature extractor; Feature extraction is performed on the query image to obtain query features, using the Resnet50 network as the feature extractor; Use the pyramid pooling module to perform the output of the fourth residual block conv5_x separately , , The average pooling operation of scale is performed, and then The convolution reduces the channel dimension to one third, and then uses bilinear interpolation to upsample the pooled feature map to restore it to its original size. Finally, the feature maps from different scales are spliced to obtain the support features of the support image. ; ; in, For splicing operation, For the upsampling operation, is the output of the fourth residual block, Represents the pooling output of the corresponding scale; query features of the query image The acquisition process of is the same as that of the supporting features of the supporting image.
3. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 1, characterized in that: Several prototypes are extracted from the supporting features. The specific process includes: Assume that the feature map of the support image is , split it according to the channel dimension to obtain The length is Feature samples , ; PMMs are probabilistic mixture models that combine the probabilities of the underlying distributions: ; ; ; ; in, is a mixing weight that satisfies and , are the parameters learned with the model, represents the i-th feature sample, represents the kth probability model based on the kernel distance function, are parameters that are updated as the model is trained, is the normalization factor, is the Bessel function, is a fixed parameter, express The mean vector of , that is, the prototype, is the kernel distance function, is the L2 vector norm; K prototypes are obtained by calculating the mean vector of K probability models based on the kernel distance function.
4. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 3, characterized in that: During the training process of the probability mixture model PMMs, optimization is performed based on the expectation maximization algorithm, and the following steps are performed in each iteration: (1-1): Calculate the expected value of each sample belonging to each prototype , according to the sample Calculate the number of Probability: ; in, is a fixed parameter; (1-2): Update the prototype according to the calculated expected value : ; Repeat steps (1-1) and (1-2) to continuously optimize the prototype estimate until convergence.
5. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 1, characterized in that: The pixel-level feature matching of the query feature based on the prototype to obtain the global matching feature includes: Assume that the query image feature map ,for For each pixel position (x, y), calculate its feature vector and prototype The cosine similarity between : ; For each prototype, we get cosine similarity values and use them to construct a similarity graph , ; One pixel calculates a similarity value, Pixels constitute a similarity graph; By calculating the similarity between K prototypes and query image features, K pixel-level similarity graphs are obtained; For each pixel position in the query feature, select the prototype with the greatest similarity To match it, replace each pixel position in the query feature with the prototype with the greatest similarity Specifically, for each pixel position (x, y), select the prototype with the greatest similarity : ; Represents the index of the most similar prototype at pixel position (x, y); Each pixel of the query feature is then replaced with its most similar prototype , thereby constructing a global matching feature , global matching features The size of is the same as the query feature and the global matching feature The feature of each position comes from the prototype with the greatest similarity .
6. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 1, characterized in that: Extract edge information from the query image, including: Using the Sobel operator to extract edge information: First, use the Sobel operator to extract edge information from the query image. For edge detection, the Sobel operator calculates The gradient of each pixel highlights the edge area in the image; After being processed by the Sobel operator, the processed result is normalized by the Sigmoid activation function to obtain the edge map ; ; Among them, S is the sigmoid activation function.
7. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 1, characterized in that: Based on the extracted edge information, the query feature is enhanced to obtain edge enhancement features, including: The edge graph The feature map of the query image Perform element-wise multiplication to highlight edge regions while suppressing features in background regions: ; in, represents an element-wise multiplication operation, is a learnable weight coefficient used to control the edge feature pair The impact of It is an edge enhancement feature.
8. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 1, characterized in that: The query feature, the global matching feature and the edge enhancement feature are fused to obtain the fused feature, including: Will , , Fusion is performed in the channel dimension and through The convolution reduces the number of channels, and then the fused features are passed to the hollow space pyramid pooling ASPP module to obtain the final feature map The ASPP module obtains multi-scale information through convolutional layers with different expansion rates, which helps to capture the features of objects of different scales in remote sensing images. : 。 9. The remote sensing image semantic segmentation method based on prototype perception and edge enhancement as claimed in claim 1, characterized in that: The trained remote sensing image semantic segmentation model. The training process includes: Constructing a remote sensing data set, wherein the remote sensing data set is K support images with known segmentation results and one query image; dividing the remote sensing data set into a training set and a test set in proportion; Input the training set into the model and train the model. During the training process, the support image and the query image are used as the input values of the model, and the image segmentation result is used as the output value of the model. When the total loss function value of the model no longer decreases, or the number of iterations reaches the set number, the training is stopped to obtain the trained remote sensing image semantic segmentation model; The total loss function of the model includes: ; ; ; ; in, is an adjustable constant used to control the effect of auxiliary loss, represents the total loss function, represents the cross entropy loss function, represents the auxiliary loss function; Represents the total number of N pixels; represents a probability distribution; is the predicted value of the pixel at position (i, j), and are the height and width of the image, is a hyperparameter, is a constant; using the cross entropy loss function Used to measure the difference between the predicted label and the actual label. Assuming that the input image has N pixels, the output of the model is a probability distribution of each pixel belonging to one of the K categories. .
10. A remote sensing image semantic segmentation system based on prototype perception and edge enhancement, characterized by: include: An acquisition module is configured to: acquire a support image and a query image; Randomly sample K+1 images from a remote sensing dataset, wherein K images are used as support images and one image is used as a query image; the support images are provided with target object category labels; and the query image has no label; The segmentation module is configured as follows: K supporting images and one query image are input into a trained remote sensing image semantic segmentation model to obtain a semantic segmentation result of the query image; wherein the trained remote sensing image semantic segmentation model includes: extracting features from the K supporting images to obtain supporting features; extracting several prototypes from the supporting features; extracting features from the query image to obtain query features; performing pixel-level feature matching on the query features based on the prototypes to obtain global matching features; extracting edge information from the query image, and enhancing the query features based on the extracted edge information to obtain edge enhanced features; performing feature fusion on the query features, the global matching features and the edge enhanced features to obtain fused features, predicting the fused features to obtain an image segmentation result.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on shared convolution kernel and boundary loss function
CN115035295A
Intestinal polyp segmentation method and system based on small sample learning
CN115049603A
Image semantic segmentation model training method, image semantic segmentation method and device
CN116310311A
Small sample semantic segmentation method and system for prototype evolution
CN117422879A
Few-sample semantic segmentation method based on prototype perception
CN119273915A
Cited By
Intelligent station inspection image and customs declaration data association method based on edge calculation
CN121858775A