SAM-based remote sensing image farmland segmentation model, method, device and medium
By adding Adapter-Feature layer and automatically generating Prompter prompt embedding in the Block block of the SAM model, the problem of poor generalization and automatic Prompter prompt capability in the remote sensing image farmland segmentation method is solved, and efficient and accurate farmland segmentation mask generation is achieved.
Patent Information
- Application Number
- CN202510149972.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing remote sensing image farmland segmentation methods have problems such as poor generalization ability and poor automatic Prompter prompting ability, which makes it difficult to achieve efficient and accurate farmland segmentation in the absence of data sets and poor network generalization.
Using the SAM-based remote sensing image farmland segmentation model, the Adapter-Feature layer is added to the Block block of the SAM model, high-frequency information of the picture is extracted, and the Prompter prompt embedding is automatically generated, and the prompt embedding is automatically generated for the mask decoder module to achieve higher quality mask segmentation.
The generalization capability of the model and the ability to automatically Prompter prompts are improved, and the accurate interpretation of remote sensing images and high-quality farmland segmentation mask generation are realized.
Smart Images

Figure CN120164093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of obtaining farmland segmentation masks using remote sensing images, and in particular, to a remote sensing image farmland segmentation model, method, device, and medium based on SAM. Background Art
[0002] Accurately dividing and positioning farmland plots is of great significance for dynamically monitoring land resources and efficiently managing agricultural production. Traditional methods are mostly based on edge detection models and region segmentation algorithms, which require a lot of manpower, have a long processing cycle, and low accuracy. With the rapid development of remote sensing technology and the application of neural networks in the field of image segmentation, methods for segmenting plots by combining remote sensing images and neural networks have gradually emerged.
[0003] Currently, publicly available remote sensing plot datasets are scarce, and there are also few cases of segmenting remote sensing image plots based on neural networks. To improve the efficiency and accuracy of remote sensing image plot segmentation, it is necessary to effectively collect remote sensing satellite images to establish a dataset and build a neural network that is efficient and practical for remote sensing plot segmentation.
[0004] Existing neural networks for segmentation such as UNet, SegNet, DeeplabV3+, and TransUNet have poor generalization performance in remote sensing field segmentation. Although the CNN-based network segmentation model performs well on the training set, its generalization performance is also poor; while using neural network models such as ViT, Swin-transformer, and mask2former, the ability of the transformer cannot be fully utilized when training with remote sensing images; for the ViT-based network segmentation model, it is difficult to utilize the performance of ViT with the information of remote sensing images; the dataset for instance segmentation of remote sensing images is scarce.
[0005] For example: The invention application with the application number 202411406225.7 discloses a fully automatic segmentation method for cultivated land plots based on SAM, belonging to the technical field of image recognition. The method used in this application scheme does not require any training data, and can well complete the extraction task for remote sensing images of all terrains and landforms, with high automation and universality. However, it also has problems such as weak generalization ability and poor automatic Prompter prompting ability.
[0006] For the above reasons, a method for segmenting farmland in remote sensing images based on the SAM model is needed to effectively solve the problems of lack of dataset and poor network generalization by using the powerful generalization and zero-shot capabilities of the SAM model. Summary of the Invention
[0007] The purpose of the present invention is to provide a farmland segmentation model, method, device and medium for remote sensing images based on SAM, improve the generalization ability of the model, automatically generate Prompter prompting ability, and achieve accurate interpretation of remote sensing images.
[0008] An embodiment of the present invention provides a farmland segmentation model, method, device and medium for remote sensing images based on SAM.
[0009] First aspect: A farmland segmentation model for remote sensing images based on SAM, comprising:
[0010] An image encoder module, which adds an Adapter-Feature layer to the Block block of the SAM model, extracts high-frequency information of the picture, and encodes the image;
[0011] An automatic Prompter module, which automatically generates prompt embeddings for the mask decoder module of the SAM model;
[0012] A mask decoder module, which uses the Prompter prompt embedding to perform mask segmentation decoding, and after splicing with the Block block features extracted by MSAF, generates a higher-quality mask segmentation.
[0013] Further, the structure of the Adapter-Feature layer includes:
[0014] The HFC obtained by unbiased Fourier convolution is combined with the Embedding obtained by the Block block through MLP, and then after extracting features as prompts and Adapter fine-tuning through MLPtune, GELU function calculation is performed, and after adjusting the feature size through MLPup, it is output to the next Block block.
[0015] Further, the structure of the MSAF includes:
[0016] Extract the features in the first and last Block blocks in the encoder, and jointly perform multi-scale feature fusion with the MaskFeature generated by the decoder through MSAF to obtain the fused Block block features.
[0017] Further, the structure of the automatic Prompter module is:
[0018] Extract the features in each block through the Feature Map to generate a FeatureAggregator. The FeatureAggregator generates candidate object boxes based on the RPN and performs position encoding through the PE map. Then, obtain the visual features of the object from the candidate object boxes based on the PE Map through ROIPooling, and use the visual features to derive three perception heads: a semantic head, a localization head, and a prompt head. Among them, the prompt head generates prompt embeddings for the mask decoder.
[0019] Further, the structure of the Feature Map is as follows:
[0020] Downsample each Block to the same channel, gradually merge the features of each layer through skip connections, obtain the final fusion convolutional layer FusionConv, and use FusionConv to obtain the FeatureAggregator;
[0021] The formula is expressed as:
[0022] F i ′ = Φ DownConv (F i )
[0023] m1 = F1′
[0024] m i = m i-1 + Φ Conv2D (m i-1 ) + F i ′
[0025] F agg = Φ FusionConv (m k )
[0026] Among them, Φ DownConv is downsampling, Φ Conv2D is gradual merging, and Φ FusionConv is fusion convolution.
[0027] Further, when performing backpropagation training on the model, the loss function includes a boundary map loss, a distance map loss, and a mask map loss, and the model is trained in reverse by jointly constraining the boundary map loss, the distance map loss, and the mask map loss.
[0028] Further, for the mask segmentation output by the model, post-processing is performed through MPRNet to remove the noise in the mask segmentation.
[0029] Second aspect: A method for segmenting farmland in remote sensing images based on SAM, including:
[0030] S1. Improve the image encoder module and mask decoder module based on the SAM model, and add an automatic Prompter module;
[0031] S2. Train the improved SAM model to obtain the trained improved SAM model;
[0032] S3. Perform remote sensing image farmland segmentation based on the trained improved SAM model to obtain high-quality mask segmentation.
[0033] In a third aspect: An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method provided in the second aspect are implemented.
[0034] In a fourth aspect: A non-transitory computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the method provided in the second aspect are implemented.
[0035] Advantages of the present invention:
[0036] 1. The present invention automatically generates Prompter prompts based on SAM, and automatically generates prompt embeddings for the mask decoder module of the SAM model, which can achieve accurate interpretation of remote sensing images.
[0037] 2. The present invention improves the SAM model, adds an Adapter-Featur block in the original SAM model Block block, pays more attention to the high-frequency information in the picture, and extracts the high-frequency information in the picture by adding the Adapter-Feature block.
[0038] 3. When the present invention performs reverse training on the model, the loss function considers the boundary map loss, distance map loss, and mask map loss. The model is reversely trained by jointly constraining the boundary map loss, distance map loss, and mask map loss, and the obtained SAM model can achieve more accurate mask segmentation. Description of the drawings
[0039] Figure 1 It is a schematic structural diagram of the remote sensing image farmland segmentation model based on SAM of the present invention;
[0040] Figure 2 It is a schematic structural diagram of the Adapter-Feature of the present invention;
[0041] Figure 3 It is a schematic structural diagram of the Block block of the present invention;
[0042] Figure 4 It is a schematic structural diagram of the FeatureAggregator of the present invention;
[0043] Figure 5 The input map of farmland for the farmland segmentation model of remote sensing images based on SAM of the present invention;
[0044] Figure 6 The segmentation mask map output by the farmland segmentation model of remote sensing images based on SAM of the present invention;
[0045] Figure 7 The schematic flow chart of the segmentation method of the present invention;
[0046] Figure 8 The schematic structural diagram of the electronic device of the present invention.
[0047] The glossary of Chinese and English terms involved in the illustrations of the present invention:
[0048] SAM (segment anything model) model;
[0049] Prompter Prompt;
[0050] Image Encoder Image Encoding;
[0051] Mask Decoder Mask Decoding;
[0052] Mask featuer Mask Feature;
[0053] Adapter-Feature Fine-tuning Feature Module;
[0054] Feature Map Feature Map;
[0055] Objectproposals Candidate Object Bounding Boxes;
[0056] PE Map Position Encoding Map;
[0057] ROIPooling Region of Interest Pooling;
[0058] FeatureAggregator Feature Aggregation;
[0059] MLP Multilayer Perceptron;
[0060] RPN (Region ProposalNetwork) Region Proposal Network
[0061] Multi-HeadAttention Multi-Head Attention;
[0062] MPRNet Multi-Stage Progressive Image Restoration Network;
[0063] Deblurring Deblurring;
[0064] MSAF multi-scale attention feature fusion DETAILED DESCRIPTION
[0065] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar symbols throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0066] Existing network models have poor generalization performance in remote sensing segmentation, and there are relatively few datasets for instance segmentation, which is not conducive to obtaining high-quality segmentation masks.
[0067] In view of the above problems, the present invention provides a remote sensing image farmland segmentation model based on SAM. Figure 1 A schematic diagram of the structure of a segmentation model provided in an embodiment of the present invention, the model includes:
[0068] Image encoder module,The image encoder module adds an Adapter-Feature layer to the Block block of the SAM model, extracts the high-frequency information of the image, and encodes the image.
[0069] Specifically, the encoder of the original SAM consists of several ViT-based blocks, which are used to extract the features of the input image. The original SAM is trained based on tens of millions of high-quality annotated images, which are rich in useful information and can make full use of the features of each ViT block.
[0070] Add an Adapter-Feature block to the original Block block to pay more attention to the high-frequency information in the image and extract the high-frequency information in the image by adding an Adapter-Feature block.
[0071] The specific Adapter-Feature block structure is as follows Figure 2 As shown in the figure, the Adapter-Feature block combines the HFC obtained by unbiased Fourier convolution with the Embedding obtained by the Block block through MLP, then extracts features as prompts through MLPtune and fine-tunes the Adapter, calculates the GELU function, adjusts the feature size through MLPup, and outputs it to the next Block block.
[0072] The structure of the Block is as follows Figure 3 As shown: The image is processed by Multi-Head Attention and Multi-Layer Perceptron (MLP).
[0073] The Automatic Prompter module automatically generates prompt embeddings for the SAM model's mask decoder module; the Automatic Prompter module automatically generates Prompter prompt embeddings and is the core module.
[0074] The Automatic Prompter module extracts features from each block through the Feature Map, generating a Feature Aggregator. The Feature Aggregator generates candidate object boxes based on the RPN and performs position encoding through the PE map; then, through ROIPooling, it obtains the visual features of the object from the candidate object boxes based on the PE Map, and uses the visual features to derive three perception heads: the semantic head, the localization head, and the prompt head. The role of the semantic head is to identify a specific object category, while the localization head is responsible for establishing the matching criterion between the generated prompt representation and the target instance mask, that is, greedy matching based on localization (Intersection over Union, IoU). The prompt head generates prompt embeddings for the mask decoder.
[0075] Specifically, the Feature Map structure is as Figure 4 shown.
[0076] The Feature Map downsamples each block to the same channel, gradually merges the features of each layer through skip connections, obtains the final fused convolutional layer FusionConv, and uses the FusionConv to obtain the Feature Aggregator;
[0077] The formula is expressed as:
[0078] F i ′ = Φ DownConv (F i )
[0079] m1 = F1′
[0080] m i = m i-1 + Φ Conv2D (m i-1 ) + F i ′
[0081] F agg = Φ FusionConv (m k )
[0082] Among them, Φ DownConv is downsampling, Φ Conv2D is gradual merging, and Φ FusionConv is fused convolution.
[0083] The mask decoder module adds HQ-Output Token and MSAF. Among them, MSAF extracts the features in the first and last Block blocks in the encoder, and together with the MaskFeature generated by the decoder, performs multi-scale feature fusion through MSAF to obtain the fused Block block features. Then, using the Prompter prompt embedding, mask segmentation decoding is performed, and after splicing with the Block block features extracted by MSAF, higher-quality mask segmentation is generated.
[0084] Furthermore, when the model is trained in reverse, the loss function includes the contour map loss, the dist_contour map loss, and the mask map loss. The model is trained in reverse by jointly constraining with the contour map loss, the dist_contour map loss, and the mask map loss.
[0085] Specifically, the BBCEWithLogitLoss function, the MSELoss function, and the RCFLoss function are jointly used for constraint training.
[0086] Furthermore, the mask segmentation output by the model is post-processed by MPRNet to remove the noise in the mask segmentation.
[0087] Using the improved SAM model of the present invention, input the remote sensing image farmland map as shown in the figure, and the output mask segmentation result map is as Figure 6 shown. Without manual Prompter prompts, the segmentation accuracy is high, and a high-quality mask segmentation map can be obtained.
[0088] The present invention also discloses a method for segmenting remote sensing image farmland based on SAM, as Figure 7 shown, including the steps:
[0089] S1. Improve the image encoder module and the mask decoder module based on the SAM model, and add an automatic Prompter module;
[0090] S2. Train the improved SAM model to obtain the trained improved SAM model;
[0091] S3. Perform remote sensing image farmland segmentation based on the trained improved SAM model to obtain high-quality mask segmentation.
[0092] Using the method of the present invention, based on SAM, Prompter prompts are automatically generated, and Adapter-Featur blocks are added to the original SAM model Block blocks, paying more attention to the high-frequency information in the picture, and extracting the high-frequency information in the picture by adding Adapter-Feature blocks.
[0093] The present invention also provides an electronic deviceFigure 8 The structural schematic diagram of the electronic device provided by the embodiment of the present invention is as follows. Figure 8 As shown, the electronic device may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory to execute the following methods, for example.
[0094] S1. Improve the image encoder module and the mask decoder module based on the SAM model, and add an automatic Prompter module.
[0095] S2. Train the improved SAM model to obtain the trained improved SAM model.
[0096] S3. Perform remote sensing image farmland segmentation based on the trained improved SAM model to obtain high-quality mask segmentation.
[0097] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0098] The embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the methods provided in the above-mentioned embodiments, for example, including:
[0099] S1. Improve the image encoder module and the mask decoder module based on the SAM model, and add an automatic Prompter module.
[0100] S2. Train the improved SAM model to obtain the trained improved SAM model.
[0101] S3. Perform remote sensing image farmland segmentation based on the trained improved SAM model to obtain high-quality mask segmentation.
[0102] The model embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A remote sensing image farmland segmentation model based on SAM, characterized in that: include: Image encoder module, adds Adapter-Feature layer to the Block of SAM model, extracts high-frequency information of the image, and encodes the image; Automatic Prompter module, which automatically generates prompt embeddings for the SAM model mask decoder module; The mask decoder module uses Prompter embedding to perform mask segmentation decoding and concatenates it with the Block features extracted by MSAF to generate higher quality mask segmentation.
2. The segmentation model according to claim 1, characterized in that The Adapter-Feature layer structure includes: The HFC obtained by unbiased Fourier convolution is combined with the Embedding obtained by the Block through MLP. Then, the features are extracted by MLPtune as prompts and fine-tuned by Adapter. The GELU function is calculated and the feature size is adjusted by MLPup before being output to the next Block.
3. The segmentation model according to claim 1, characterized in that The MSAF structure includes: The features of the first and last blocks in the encoder are extracted and fused with the mask features generated by the decoder through MSAF to obtain the fused block features.
4. The segmentation model according to claim 1, characterized in that The automatic prompter module structure is: The features in each Block are extracted through Feature Map to generate a Feature Aggregator. Feature Aggregator generates candidate object boxes based on RPN and encodes the positions through PE map. Then, ROIPooling is used to obtain the visual features of the object from the candidate object boxes based on PE Map. The visual features are used to derive three perception heads: semantic head, positioning head and hint head. The hint head generates hint embedding for the mask decoder.
5. The segmentation model according to claim 4, characterized in that The structure of the Feature Map is: Each Block is downsampled to the same channel, and the features of each layer are gradually merged through skip connections to obtain the final fusion convolution layer FusionConv, and FusionConv is used to obtain FeatureAggregator; The formula is: F i ′=Φ DownConv (F i ) m1=F1′ m i =m i-1 +Φ Conv2D (m i-1 )+F i ′ F agg =Φ FusionConv (m k ) Among them, Φ DownConv is downsampling, Φ Conv2D To gradually merge, Φ FusionConv It is a fused convolution.
6. The segmentation model according to claim 1, characterized in that When the model is reversely trained, the loss function includes boundary map loss, distance map loss and mask map loss. The model is reversely trained by jointly constraining the boundary map loss, distance map loss and mask map loss.
7. The segmentation model according to claim 1, characterized in that The mask segmentation output by the model is post-processed through MPRNet to remove noise in the mask segmentation.
8. A remote sensing image farmland segmentation method based on SAM, characterized in that: Includes steps: S1. Improve the image encoder module and mask decoder module based on the SAM model, and add an automatic prompter module; S2, training the improved SAM model to obtain the trained improved SAM model; S3. Based on the trained improved SAM model, remote sensing image farmland segmentation is performed to obtain high-quality mask segmentation.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the segmentation method according to claim 8 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the segmentation method according to claim 8 are implemented.
Citation Information
Patent Citations
Full-automatic segmentation method for cultivated land parcels based on SAM
CN118918332A
Image denoising method, system and device and storage medium
CN116012266A
Remote sensing image building segmentation method based on adaptive feedback SAM and related device
CN118505715A
Techniques for weakly supervised referring image segmentation
US20240013504A1
Cited By
Farmland parcel extraction method and system based on large model fine tuning
CN121305378A