A feature matching remote sensing image small sample semantic segmentation model evolution method
By constructing a feature-matching semantic segmentation model for remote sensing images with few samples, and utilizing the prior knowledge of the basic model and feature fusion technology, the problem of insufficient semantic segmentation capability for remote sensing images with few samples is solved, and more efficient semantic segmentation results are achieved.
Patent Information
- Application Number
- CN202511207081.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing methods for few-sample semantic segmentation of remote sensing images ignore the strong prior knowledge of the base model, resulting in insufficient few-sample semantic segmentation capabilities.
A few-sample semantic segmentation model for remote sensing images based on feature matching is constructed. Features are extracted through the basic model backbone module, features are filtered through the foreground learning module, matching and mapping modules generate matching features, and feature fusion is performed using the decoder module. Finally, the model is optimized through a loss function.
It enhances the semantic segmentation capability of remote sensing images and improves the model's generalization ability and segmentation accuracy.
Smart Images

Figure CN120747517B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of pattern recognition, and particularly relates to a feature matching remote sensing image small sample semantic segmentation model evolution method. BACKGROUND
[0002] Semantic segmentation is a basic task in computer vision, which has a wide range of applications in autonomous driving, medical image analysis. Semantic segmentation needs to distinguish the categories of a given picture and accurately predict the segmentation mask of each category. However, the dense pixel-level annotation data leads to high cost, and the generalization ability of the model is still a bottleneck when facing new categories. Therefore, the small sample semantic segmentation task is proposed.
[0003] Few-Shot Segmentation (FSS) is an extension of few-shot learning technology in the field of semantic segmentation, and its basic goal is to give an image with pixel-level segmentation mask annotation as a support image, and to segment the potential target from the image to be measured. For a small number of labeled samples (support set) from a new category, the small sample semantic segmentation model can accurately segment other samples (query set) in the same category. Compared with natural images, remote sensing images have large resolution difference, complex texture, and fuzzy semantic boundary, which makes the remote sensing image FSS task more challenging.
[0004] In recent years, the method for small sample semantic segmentation of remote sensing images has developed rapidly. Previous methods can be divided into three categories, the first category is to construct a deep feature pyramid comparison module, or to propose an attention aggregation network to adaptively fuse multi-scale features. The second method enhances feature information through mutual reinforcement of global semantics and spatial density. The third method integrates text modal information in class description, and increases support information to improve semantic segmentation performance. However, these methods ignore the strong prior knowledge of the basic model. SUMMARY
[0005] The purpose of the present application is to solve the problem that the prior art method for small sample semantic segmentation of remote sensing images ignores the strong prior knowledge of the basic model, resulting in insufficient small sample semantic segmentation capability, and a feature matching remote sensing image small sample semantic segmentation model evolution method is proposed, which models the correlation of the support image, the support image foreground region and the target image of each module of the remote sensing image small sample semantic segmentation model, and enhances the semantic segmentation capability for remote sensing images.
[0006] The technical scheme of the present application is: a feature matching remote sensing image small sample semantic segmentation model evolution method, comprising the following steps:
[0007] The basic model backbone module of the remote sensing image small sample semantic segmentation model is constructed to extract query features of a query image and support features of a support image.
[0008] The foreground learning module of the remote sensing image small sample semantic segmentation model is constructed to receive support masks and support features of the support image, to obtain spatial features and semantic features, and to screen the spatial features and the semantic features through cosine similarity to obtain foreground features.
[0009] The matching mapping module of the remote sensing image small sample semantic segmentation model is constructed to multiply the support masks and the support features of the support image pixel by pixel to obtain mask foreground features, and to calculate the similarity of the mask foreground features and the query features to generate matching features.
[0010] The decoder module of the remote sensing image small sample semantic segmentation model is constructed to input the query features, the foreground features and the matching features after being spliced in the channel dimension into a multi-layer self-attention network to output a semantic segmentation mask of the query image.
[0011] Based on the semantic segmentation mask, a loss function is constructed to optimize and train the remote sensing image small sample semantic segmentation model to complete the evolution of the remote sensing image small sample semantic segmentation model.
[0012] Preferably, the basic model backbone module is a pre-trained basic diffusion model Dinov2, and the parameters of a visual encoder ViT in the pre-trained basic diffusion model Dinov2 are all frozen in the training stage.
[0013] Preferably, the foreground learning module receives support masks and support features of the support image to obtain spatial features and semantic features, and screens the spatial features and the semantic features through cosine similarity to obtain foreground features, which specifically includes the following steps:
[0014] The support masks of the support image are input into the foreground learning module to learn spatial features of the support image about a target region through GAP;
[0015] The support features are input into the foreground learning module to predict the prediction categories of the support features through a linear layer;
[0016] Based on the spatial features of the target region and the prediction categories of the support features, the semantic features of the support image about the target region are learned through GAP;
[0017] The cosine similarity of the spatial features of the target region and the query features, and the cosine similarity of the semantic features of the target region and the query features are calculated, and the feature with the largest cosine similarity is selected from the spatial features and the semantic features as the foreground feature.
[0018] As preferred, the support mask of the support image is multiplied with the support feature pixel by pixel by the matching mapping module to obtain a mask foreground feature, and similarity of the mask foreground feature and the query feature is calculated to generate a matching feature, specifically:
[0019] The support feature and the support mask of the support image are input into the matching mapping module, and the region information of the support image is extracted by pixel by pixel multiplication to obtain a mask foreground feature;
[0020] The similarity of the mask foreground feature and the query feature is calculated to further generate a matching feature.
[0021] As preferred, the specific formula for calculating the similarity of the mask foreground feature and the query feature is:
[0022]
[0023] wherein, the mask foreground feature is represented as, , the similarity of the mask foreground feature and the query feature is represented as, and the L2 norm normalization is represented as.
[0024] As preferred, the calculation formula of the matching feature is:
[0025]
[0026] wherein, the max-min normalization is represented as, and the max operation is represented as.
[0027] As preferred, the decoder module includes four layers of stacked sub-modules and one linear layer connected in sequence, and each layer of stacked sub-module is composed of a multi-head self-attention unit and a feedforward neural network unit connected in sequence;
[0028] The multi-head self-attention unit is used to promote the interaction between features and fuse the input features;
[0029] The feedforward neural network unit is used to enhance the non-linear expression ability of the features;
[0030] The linear layer is used to map the feature channel to 1.
[0031] As preferred, the query feature, the foreground feature and the matching feature are spliced in the channel dimension and input into the multi-layer self-attention network, and the semantic segmentation mask of the query image is output, specifically:
[0032] The query feature, the foreground feature and the matching feature are superimposed in the channel dimension to obtain a feature ; wherein represents a real number field, represents the number of channels of a feature, represents the height of a feature, represents the width of a feature;
[0033] The scale of the feature is changed through a reshape operation to obtain a feature , and the feature is input to a four-layer stacked sub-module for feature interaction and fusion to obtain a fused feature;
[0034] The channel latitude of the fused feature is mapped to 1 through a linear layer to obtain a prediction mask;
[0035] The prediction mask is restored to the original size through a reshape operation and bilinear interpolation to obtain the final prediction mask, i.e., the semantic segmentation mask of the query image.
[0036] As a preferred, the loss function is:
[0037]
[0038] wherein, represents a loss function, represents a binary cross-entropy loss, used to evaluate the difference between the prediction mask of the query image and the real mask, the prediction mask being the semantic segmentation mask represents a Dice loss, represents the prediction mask output by the decoder module, represents the query mask, i.e., the real mask.
[0039] The present application has the following advantages:
[0040] 1. The present application proposes a small sample semantic segmentation method for remote sensing images based on Dinov2, which utilizes the powerful prior knowledge of the basic model Dinov2 to model the relationship between the target image and the support image in the remote sensing image. Therefore, the model of the present application is more direct and efficient.
[0041] 2. The present application introduces a foreground learning module and a matching mapping module, which respectively model the foreground semantic information and the matching feature information, and utilize the decoder to model the relationship between the features to predict the segmentation of the query image. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 The flowchart of a feature matching-based remote sensing image small sample semantic segmentation model evolution method is shown.
[0043] Figure 2A flow chart of a feature matching-based remote sensing image small sample semantic segmentation model evolution method is shown. DETAILED DESCRIPTION
[0044] Exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments illustrated and described herein are merely exemplary and are not intended to limit the scope of the present application, which is defined by the appended claims.
[0045] Embodiment 1:
[0046] The present application proposes a Dinov2-oriented remote sensing image small sample semantic segmentation architecture. Small sample semantic segmentation usually has very little data, and the model needs to learn the same target in the query image from very few support samples, and the training classes and test classes are completely different. The architecture uses the powerful visual feature matching capability of the basic visual model Dinov2 to design a lightweight feature interaction decoder for small sample semantic segmentation of remote sensing images. The visual model Dinov2 can obtain query image features and support image features with strong matching capability, and the feature interaction decoder can further establish the correlation between the query image and the support image, and use the foreground region obtained from the support image and its mask image to further establish the relationship with the query image, thereby enhancing the small sample semantic segmentation capability.
[0047] The present application uses the powerful visual feature matching capability of the basic visual model Dinov2 to introduce a lightweight feature interaction decoder, which is more direct and efficient. The present application introduces a new type of feature interaction decoder, which enhances the semantic segmentation capability for remote sensing images by modeling the correlation between the support image, the support image foreground region and the target image.
[0048] As shown in Figure 1 and Figure 2 A feature matching-based remote sensing image small sample semantic segmentation model evolution method includes the following steps:
[0049] S1. Constructing a basic model backbone module of a remote sensing image small sample semantic segmentation model for extracting query features of a query image and support features of a support image;
[0050] S2. Constructing a foreground learning module of a remote sensing image small sample semantic segmentation model for receiving support masks and support features of a support image, then obtaining spatial features and semantic features, and screening the spatial features and the semantic features through cosine similarity to obtain foreground features;
[0051] S3. Construct a matching mapping module for a remote sensing image small sample semantic segmentation model, which is used to multiply the support mask of the supporting image with the support features pixel by pixel to obtain the mask foreground features, and calculate the similarity between the mask foreground features and the query features to generate matching features;
[0052] S4. Construct a decoder module for a remote sensing image few-sample semantic segmentation model, which is used to concatenate query features, foreground features and matching features in the channel dimension and input them into a multi-layer self-attention network to output the semantic segmentation mask of the query image;
[0053] S5. Based on the semantic segmentation mask, a loss function is constructed to optimize and train the remote sensing image few-sample semantic segmentation model, thus completing the evolution of the remote sensing image few-sample semantic segmentation model.
[0054] In this embodiment, the core module of the remote sensing image few-sample semantic segmentation model is the pre-trained Dinov2 base diffusion model. The powerful prior knowledge of the Dinov2 model is used to extract corresponding features from the image. Specifically, the feature encoder ViT in Dinov2 is used to obtain the original representation of the pre-trained model. Given a query image... and supporting images First, the image encoder receives the query image and support images as input, respectively, and extracts the query features using the visual encoder ViT in the basic vision model Dinov2. and supporting features .in, Represents the real number field. Indicates the image height. Indicates the image width. The number of channels representing the feature. Indicates feature height, This represents the feature width. During visual feature extraction, to preserve the prior knowledge of the base visual model Dinov2—that is, to retain the feature matching ability of the base model to a greater extent—this invention freezes all parameters of the image encoder ViT. To further align with the input of Dinov2, the input images are uniformly scaled to 518×518, and data augmentation operations such as random cropping are performed during training to improve the model's generalization ability.
[0055] In this embodiment, the foreground learning module of the present invention receives supporting features. and support mask As input, spatial features of the supporting image about the target region are learned through GAP. The introduction of support masks helps to accurately filter out target regions from support features, thus facilitating the acquisition of spatial-level features of the target. Simultaneously receiving support features... Using linear layers to predict support features From the category information, the predicted category is obtained. Similarly, supporting features... The predicted categories are learned through GAP to support the semantic features of the image about the target region. Supports spatial features of images. and semantic features Respectively with query image features Calculate the cosine similarity, which represents the degree of similarity between image features. The higher the similarity value, the more likely the two regions are to be the same target. Therefore, select the feature with the highest similarity as the final output foreground feature. .
[0056] In this embodiment, the present invention constructs a matching mapping module to model the similarity between query features and supporting features. The matching mapping module in this invention receives multi-scale supporting features. and support mask As input, region information supporting features is extracted through pixel-by-pixel multiplication. Subsequently, the similarity between the mask foreground features and the query image features is calculated, and matching features are further generated. The specific calculation formula is as follows:
[0057]
[0058]
[0059] in, This represents the foreground features of the mask, i.e., the regional information supporting the features. , Represents the foreground features of the mask and query features similarity, This indicates L2 norm normalization. This represents max-min normalization. This indicates the operation of retrieving the maximum value.
[0060] In this embodiment, to predict the segmentation of the query image, the present invention proposes a decoder module, the input of which is the query features. Foreground characteristics and matching features Features are obtained by overlaying along the channel dimension. The decoder module comprises four stacked sub-modules, each consisting of a multi-head self-attention module and a feedforward neural network module. The multi-layered stacked self-attention modules facilitate the interaction of information between different features, further strengthening the correlation between query features and supporting features, and then outputting a predicted segment of the query image. The self-attention module promotes the interaction between object features and can further fuse input features, thereby improving the accuracy of the predicted target. The size was changed after the reshape operation. To match the input of the self-attention mechanism, where , Indicates a query. Indicates key, The values are represented. Finally, the outputs of the four stacked submodules are passed through a linear layer to map the channel dimension to 1, resulting in the prediction mask. The predicted mask is further restored to its original size through a reshape operation and bilinear interpolation to obtain the final predicted mask. .
[0061] In this embodiment, the present invention uses binary cross-entropy (BCE) as the main loss function to evaluate the difference between the predicted mask and the true mask of the query image. Considering the complex characteristics of remote sensing images, the present invention further introduces the Dice Loss function to further evaluate the loss. By calculating the loss between the predicted mask and the query mask, the few-sample semantic segmentation model is optimized. The loss function is:
[0062]
[0063] in, Represents the loss function. This represents the binary cross-entropy loss, used to evaluate the difference between the predicted mask and the ground truth mask of the query image. The predicted mask is the semantic segmentation mask. This indicates Dice's loss. This represents the prediction mask output by the decoder module. This indicates that the query mask is the same as the actual mask.
[0064] Example 2:
[0065] Based on Example 1, this embodiment of the invention conducted experiments on the widely used small-sample semantic segmentation dataset, i.e., the iSAID dataset, to illustrate the effectiveness of the method proposed in this invention.
[0066] The iSAID dataset is a large-scale, finely labeled segmentation dataset containing 2806 high-resolution images. This dataset includes 15 categories: ships, storage tanks, baseball fields, tennis courts, basketball courts, athletic fields, bridges, large vehicles, small vehicles, helicopters, swimming pools, roundabouts, football fields, airplanes, and ports. These categories are evenly divided into three classes. The few-shot model is tested on the target class and trained on the remaining classes. Experimental results are shown in Table 1. It can be seen that the proposed method achieves 27.81% Mean IoU on the iSAID dataset, surpassing the performance of current state-of-the-art few-shot semantic segmentation methods.
[0067] Table 1. Comparison of the present invention with state-of-the-art methods on the few-sample semantic segmentation task.
[0068]
[0069] Example 3:
[0070] Based on Example 1, this embodiment of the invention provides an evolution system for a remote sensing image few-sample semantic segmentation model based on feature matching, which can be used to implement the remote sensing image few-sample semantic segmentation model evolution method based on feature matching as described in the foregoing embodiments.
[0071] In this embodiment, the system may be an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. The processor executes the program to implement some or all of the steps of the evolution of the feature matching-based remote sensing image small sample semantic segmentation model as described in Embodiment 1.
[0072] In an exemplary embodiment, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the feature matching-based remote sensing image few-sample semantic segmentation model evolution method as described in Embodiment 1 above.
[0073] In an exemplary embodiment, the readable storage medium may be a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the feature matching-based remote sensing image small sample semantic segmentation model evolution method according to Embodiment 1 above.
[0074] In an exemplary embodiment, the computer program product includes a computer program that, when executed by a processor, implements the feature matching-based remote sensing image few-sample semantic segmentation model evolution method according to Embodiment 1 above.
[0075] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0076] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0077] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0078] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0079] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0080] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. An evolutionary method for feature-matching remote sensing image few-sample semantic segmentation model, characterized in that, Includes the following steps: The basic model backbone module for constructing a few-sample semantic segmentation model for remote sensing images is used to extract query features from query images and support features from supporting images. A foreground learning module is constructed for a remote sensing image few-sample semantic segmentation model. This module receives the support mask and support features of the supporting image, and then obtains spatial and semantic features. The spatial and semantic features are then filtered by cosine similarity to obtain foreground features. A matching mapping module is constructed for a remote sensing image few-sample semantic segmentation model. This module is used to multiply the support mask of the supporting image with the support features pixel by pixel to obtain the mask foreground features, and calculate the similarity between the mask foreground features and the query features to generate matching features. A decoder module for constructing a few-sample semantic segmentation model for remote sensing images is used to concatenate query features, foreground features, and matching features in the channel dimension and input them into a multi-layer self-attention network to output a semantic segmentation mask of the query image. Based on the semantic segmentation mask, a loss function is constructed to optimize and train the remote sensing image few-sample semantic segmentation model, thus completing the evolution of the remote sensing image few-sample semantic segmentation model. The core module of the basic model is the pre-trained basic diffusion model Dinov2; the parameters of the visual encoder ViT in the pre-trained basic diffusion model Dinov2 are all frozen during the training phase. The decoder module consists of four stacked sub-modules connected in sequence and one linear layer. Each stacked sub-module is composed of a multi-head self-attention unit and a feedforward neural network unit connected in sequence. Multi-head self-attention units are used to facilitate interactions between features and to fuse input features; Feedforward neural network units are used to enhance the nonlinear representation of features; Linear layers are used to map feature channels to 1.
2. The feature-matching remote sensing image few-sample semantic segmentation model evolution method according to claim 1, characterized in that, The foreground learning module receives the support mask and support features of the supporting image, thereby obtaining spatial and semantic features. Foreground features are then obtained by filtering the spatial and semantic features using cosine similarity. The specific steps include: The support mask of the support image is input into the foreground learning module, and the spatial features of the support image about the target region are learned through GAP. The supporting features are input into the foreground learning module, and the predicted categories of the supporting features are obtained through linear layer prediction. Based on the predicted categories of spatial features and supporting features of the target region, semantic features of the supporting image about the target region are learned through GAP; Calculate the cosine similarity between the spatial features of the target region and the query features, as well as the cosine similarity between the semantic features of the target region and the query features, and select the feature with the highest cosine similarity from the spatial features and semantic features as the foreground feature.
3. The feature-matching remote sensing image few-sample semantic segmentation model evolution method according to claim 1, characterized in that, The matching mapping module multiplies the support mask and support features of the supporting image pixel by pixel to obtain the mask foreground features, and calculates the similarity between the mask foreground features and the query features to generate matching features. Specifically: The supporting features and supporting mask of the supporting image are input into the matching mapping module, and the region information of the supporting image is extracted by pixel-by-pixel multiplication to obtain the mask foreground features; The similarity between the mask foreground features and the query features is calculated, and then matching features are generated.
4. The method for evolving a feature-matching remote sensing image few-sample semantic segmentation model according to claim 3, characterized in that, The specific formula for calculating the similarity between the mask foreground features and the query features is as follows: in, Representing the foreground features of the mask, having , Represents the foreground features of the mask and query features similarity, This indicates L2 norm normalization.
5. The feature-matching remote sensing image few-sample semantic segmentation model evolution method according to claim 4, characterized in that, Matching features The calculation formula is: in, This represents max-min normalization. This indicates the operation of retrieving the maximum value.
6. The method for evolving a feature-matching remote sensing image few-sample semantic segmentation model according to claim 1, characterized in that, The query features, foreground features, and matching features are concatenated along the channel dimension and then input into a multi-layer self-attention network to output a semantic segmentation mask for the query image, specifically: The query features, foreground features, and matching features are superimposed along the channel dimension to obtain the features. ;in, Represents the real number field. The number of channels representing the feature. Indicates the height of the feature. Indicates the width of the feature; Change features through the reshape operation The scale is used to obtain features. , will feature The input is fed into a four-layer stacked submodule for feature interaction and fusion to obtain fused features; The channel dimensions of the fused features are mapped to 1 by a linear layer to obtain the prediction mask; The predicted mask is restored to its original size by reshape operation and bilinear interpolation to obtain the final predicted mask, which is the semantic segmentation mask of the query image.
7. The method for evolving a feature-matching remote sensing image few-sample semantic segmentation model according to claim 1, characterized in that, The loss function is: in, Represents the loss function. This represents the binary cross-entropy loss, used to evaluate the difference between the predicted mask and the ground truth mask of the query image. The predicted mask is the semantic segmentation mask. This indicates Dice's loss. This represents the prediction mask output by the decoder module. This indicates that the query mask is the same as the actual mask.
Citation Information
Patent Citations
Small sample segmentation method and system based on semantic transfer and context distribution modeling
CN120182605A
Multi-modal large model image segmentation method and device based on hierarchical lexical representation
CN120298683A