Cultivated Land Extraction Method, Device, Equipment and Medium Based on Remote Sensing Images

By combining semantic segmentation model and visual segmentation model, extracting and segmenting cultivated land information in remote sensing images, the problem of inaccurate cultivated land boundaries in the existing technology is solved, and more efficient cultivated land identification and agricultural management are achieved.

CN116994140BActive Publication Date: 2025-06-24BEIJING AEROSPACE HONGTU INFORMATION TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311021776.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-14
Publication Date
2025-06-24
Estimated Expiration
2043-08-14

AI Technical Summary

Technical Problem

The prior art is difficult to extend to different spatial and temporal conditions in arable land extraction, and it is difficult to utilize advanced image features, resulting in inaccurate arable land boundaries.

Method used

Using a method based on remote sensing image, combined with a pre-trained semantic segmentation model and a visual segmentation model, the cultivated map spots are extracted through the semantic segmentation model, and image segmentation is performed through the visual segmentation model, superimposed calculations are used to determine the semantic category of the segmented object, and finally the cultivated land distribution information is extracted.

Benefits of technology

It has improved the accuracy of arable land identification, improved the boundaries of arable land, improved the efficiency of agricultural management and land use planning, and provided important support for the sustainable development of precision agriculture and agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994140B_ABST
    Figure CN116994140B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, equipment and medium for extracting cultivated land based on remote sensing images, which relates to the technical field of remote sensing interpretation. The method includes: obtaining a remote sensing image to be predicted; performing cultivated land prediction on the remote sensing image to be predicted through a pre-trained semantic segmentation model to obtain target patches corresponding to the remote sensing image to be predicted; performing image segmentation processing on the remote sensing image to be predicted through a pre-trained visual segmentation model to obtain a plurality of different segmentation objects; superimposing the target patches and the segmentation objects, and for each segmentation object, calculating the semantic category in the cultivated land patches at the spatial position corresponding to the target patches; extracting the cultivated land distribution information of each segmentation object based on the semantic category. The present application can improve the accuracy of cultivated land recognition and perfect the cultivated land boundary, which helps to improve the efficiency of agricultural management and land use planning, and provides important support for precision agriculture and agricultural sustainable development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of remote sensing interpretation, and in particular, to a method, device, equipment and medium for extracting cultivated land based on remote sensing images. Background Art

[0002] As the basic unit of agricultural management, extracting information about cultivated land is very important for providing key information on agricultural productivity and environmental footprint to meet food security and sustainability needs, such as crop mapping, crop yield estimation, and sustainable agricultural planning. In related technologies, the cultivated land extraction algorithm uses machine learning algorithms with medium and low-level features, which cannot be well extended to different spatial and temporal conditions, limits the utilization of high-level image features, and it is also difficult to obtain accurate closed boundaries. Summary of the Invention

[0003] The purpose of this application is to provide a method, device, equipment and medium for extracting cultivated land based on remote sensing images, which can improve the accuracy of cultivated land recognition and perfect the cultivated land boundary, contribute to improving the efficiency of agricultural management and land use planning, and provide important support for precision agriculture and sustainable agricultural development.

[0004] In a first aspect, the present invention provides a method for extracting cultivated land based on remote sensing images, the method comprising: obtaining a to-be-predicted remote sensing image; performing cultivated land prediction on the to-be-predicted remote sensing image through a pre-trained semantic segmentation model to obtain a target patch corresponding to the to-be-predicted remote sensing image, the target patch including cultivated land patches; performing image segmentation processing on the to-be-predicted remote sensing image through a pre-trained visual segmentation model to obtain a plurality of different segmentation objects; wherein, the spatial positions of the segmentation objects correspond to those of the target patch; superimposing the target patch and the segmentation objects, and for each segmentation object, calculating the semantic category in the cultivated land patches at the spatial position corresponding to the target patch; extracting the cultivated land distribution information of each segmentation object based on the semantic category.

[0005] In an optional embodiment, performing cultivated land prediction on the to-be-predicted remote sensing image through a pre-trained semantic segmentation model to obtain a target patch corresponding to the to-be-predicted remote sensing image includes: extracting a feature map of the to-be-predicted remote sensing image through a pre-trained semantic segmentation model based on multi-layer convolution and pooling operations; performing upsampling and convolution operations on the feature map to restore the feature map to the resolution of the original to-be-predicted remote sensing image, and predicting the semantic category to which each pixel in the to-be-predicted remote sensing image belongs, and each semantic category is used to be represented in the form of a patch.

[0006] In an alternative embodiment, the visual segmentation model is a SAM segmentation model, and the SAM segmentation model includes an image encoder module, a prompt encoder module, and a mask decoder; wherein, the image encoder module uses a vision transformer pre-trained by MAE for feature dimensionality reduction of the remote sensing image to be predicted; the prompt encoder module includes a sparse prompt unit and a dense prompt unit, where the sparse prompt unit represents the prompt points and prompt boxes through positional encoding and adds them to the learned embeddings of the prompt types and free-form texts obtained from the multi-modal neural network; the dense prompt unit uses convolutional embeddings and performs element-wise summation with the result of the feature dimensionality reduction; the mask decoder is used to map the outputs of the image encoder module and the prompt encoder module to masks.

[0007] In an alternative embodiment, the image segmentation process is performed on the remote sensing image to be predicted through a pre-trained visual segmentation model, and multiple different segmentation objects are obtained, including: performing single-point sampling on several points in the grid above the remote sensing image to be predicted through the pre-trained visual segmentation model, generating multiple masks from each single-point prompt, and then performing quality screening and duplicate data deletion on the generated masks through non-maximum suppression; setting parameters based on preset adjustable parameters to perform image segmentation processing based on the set parameters to obtain multiple different segmentation objects.

[0008] In an alternative embodiment, the preset adjustable parameters include sampling point density, quality screening threshold, duplicate mask removal threshold, image cropping parameters, and post-processing parameters.

[0009] In an alternative embodiment, the target patch further includes a background patch; superimposing the target patch and the segmentation object, and for each segmentation object, calculating the semantic category in the spatial position corresponding to the target patch, including: superimposing the target patch and the segmentation object, and for each segmentation object, counting the number of pixels belonging to the cultivated land patch and the background patch corresponding to the superimposed segmentation object, and respectively calculating the first pixel ratio of the cultivated land patch corresponding to the segmentation object and the second pixel ratio of the background patch; based on the relationship between the first pixel ratio of the cultivated land patch and the second pixel ratio of the background patch, determining the target pixel ratio of the cultivated land patch in the spatial position corresponding to the target patch for each segmentation object, so as to determine the semantic category corresponding to the segmentation object based on the target pixel ratio.

[0010] In an alternative embodiment, the method further includes: splicing the segmentation objects according to the geographical location information and converting the spliced target object into a vector patch to obtain the distribution range of the cultivated land.

[0011] In a second aspect, the present invention provides a cultivated land extraction device based on remote sensing images, the device comprising: an image acquisition module for acquiring the to-be-predicted remote sensing image; a cultivated land prediction module for performing cultivated land prediction on the to-be-predicted remote sensing image through a pre-trained semantic segmentation model to obtain target patches corresponding to the to-be-predicted remote sensing image, the target patches including cultivated land patches; an image segmentation module for performing image segmentation processing on the to-be-predicted remote sensing image through a pre-trained visual segmentation model to obtain a plurality of different segmentation objects; wherein the segmentation objects correspond to the spatial positions of the target patches; an overlay calculation module for overlaying the target patches and the segmentation objects, and for each segmentation object, calculating the semantic category in the cultivated land patches at the spatial position corresponding to the target patch; a cultivated land extraction module for extracting the cultivated land distribution information of each segmentation object based on the semantic category.

[0012] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, the memory storing computer-executable instructions capable of being executed by the processor, and the processor executing the computer-executable instructions to implement the cultivated land extraction method based on remote sensing images according to any one of the foregoing embodiments.

[0013] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the cultivated land extraction method based on remote sensing images according to any one of the foregoing embodiments.

[0014] The cultivated land extraction method, device, equipment and medium provided by the present application can improve the accuracy of cultivated land recognition and perfect the cultivated land boundary by combining the semantic annotation of the semantic segmentation model and the segmentation objects of SAM, which helps to improve the efficiency of agricultural management and land use planning, and provides important support for precision agriculture and sustainable agricultural development. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a flowchart of a cultivated land extraction method based on remote sensing images provided by an embodiment of the present application;

[0017] Figure 2Structural diagram of a SAM segmentation model provided by an embodiment of the present application;

[0018] Figure 3 Flowchart of a specific cultivated land extraction based on remote sensing images provided by an embodiment of the present application;

[0019] Figure 4 Schematic diagram of remote sensing images in an example provided by an embodiment of the present application;

[0020] Figure 5 Schematic diagram of a sample data set provided by an embodiment of the present application;

[0021] Figure 6 Cultivated land patches extracted by a semantic segmentation model provided by an embodiment of the present application;

[0022] Figure 7 Schematic diagram of a SAM segmentation result provided by an embodiment of the present application;

[0023] Figure 8 Schematic diagram of the voting result of the pixel category proportion provided by an embodiment of the present application;

[0024] Figure 9 Schematic diagram of the vector result of cultivated land extraction provided by an embodiment of the present application;

[0025] Figure 10 Structural diagram of a cultivated land extraction device based on remote sensing images provided by an embodiment of the present application;

[0026] Figure 11 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0028] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0029] It should be noted that similar reference numerals and letters refer to similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0030] The growing population and consumption pose global challenges to agriculture and natural resources. Improving food production and sustainability is the solution to this major challenge. As the basic unit of agricultural management, detailed spatial information is crucial for providing key information on agricultural productivity and environmental footprint to meet food security and sustainability needs, such as crop mapping, crop yield estimation, and sustainable agricultural planning.

[0031] Remote sensing earth observation technology has the advantages of large-scale, long-term, and periodic monitoring. Using deep learning to obtain surface cultivated land information from remote sensing images is one of the most widely studied fields currently.

[0032] Currently, most existing cultivated land extraction algorithms use machine learning algorithms with medium and low-level features, which cannot be well extended to different spatial and temporal conditions. At the same time, it limits the utilization of high-level image features and is also difficult to obtain accurate closed boundaries. Conventional deep learning-based semantic segmentation models have relatively ideal effects in cultivated land recognition. These models can provide semantic information for each pixel of the image and achieve a rough extraction of cultivated land. However, due to the fuzzy and inaccurate boundaries predicted by the models, accurate masks cannot be obtained, that is, the cultivated land areas cannot be accurately marked.

[0033] Based on this, the embodiments of the present application provide a method for extracting cultivated land from remote sensing images. See Figure 1 As shown, the method mainly includes the following steps:

[0034] Step S110, obtain the remote sensing image to be predicted.

[0035] In one implementation, the bit depth of the remote sensing image is usually 16 bits, and it needs to be converted to 8 bits, which is suitable for the deep learning semantic segmentation model. When performing bit depth processing, linear stretching can be used, that is, the pixel values distributed in [value min , value max in each band are normalized, and then uniformly linearly stretched to [out min , out max , and the formula is:

[0036]

[0037] In the formula, value(x,y) is each pixel in the band; when value min , value maxWhen calculating the minimum and maximum values of pixels, it is called minimum-maximum stretching; when value min , value max is the percentile pixel value of the histogram cumulative, it is called percentile stretching, and the common one is 2-98% stretching; out min , out max is the pixel range after stretching. When it is 0 and 255, it is the pixel range of an 8-bit depth image, and result1 is the stretching result.

[0038] Step S120, use a pre-trained semantic segmentation model to perform cultivated land prediction on the remote sensing image to be predicted, and obtain the target patches corresponding to the remote sensing image to be predicted. The target patches include cultivated land patches.

[0039] To perform target extraction on the remote sensing image to be predicted, a semantic segmentation model in deep learning can be used. By assigning each pixel in the remote sensing image to be predicted to a specific semantic category, the pixel-level segmentation and classification of the remote sensing image to be predicted can be realized, so as to obtain the category corresponding to the cultivated land after segmentation.

[0040] The semantic segmentation model can learn the features of different target objects in the image, classify each pixel into the corresponding semantic category, such as cultivated land, buildings, water bodies, etc., and can obtain a target extraction result with high accuracy, as well as relatively accurate position and boundary information of the target in the image.

[0041] In one implementation, the semantic segmentation model can adopt a semantic segmentation model in deep learning, such as adopting a Convolutional Neural Network (CNN) architecture, and optimizing it through a large amount of remote sensing image data with pixel-level annotations during the training process.

[0042] In specific implementation, the following steps 1.1) and 1.2) can be adopted for semantic segmentation processing:

[0043] Step 1.1), use a pre-trained semantic segmentation model to extract the feature map of the remote sensing image to be predicted based on multi-layer convolution and pooling operations;

[0044] Step 1.2), perform upsampling and convolution operations on the feature map to restore the feature map to the resolution of the original remote sensing image to be predicted, and predict the semantic category to which each pixel in the remote sensing image to be predicted belongs. Each semantic category is used to be represented in the form of patches.

[0045] Specifically, the model extracts the features of the image through multiple convolutional and pooling operations, uses upsampling and convolutional operations to restore the feature map to the resolution of the original image, and predicts the semantic category to which each pixel belongs. During training, the model learns pixel-level semantic information by minimizing the difference between the predicted label and the true label.

[0046] Optionally, semantic segmentation models such as UNet, DeepLab series, SegFormer, etc. can also be used to train the model on the sample dataset of the target, obtain the model weights, and use these weights to extract the target objects from the image to be predicted, obtaining a rough class classification, that is, the semantic annotation of each pixel.

[0047] Step S130, perform image segmentation processing on the remote sensing image to be predicted through a pre-trained visual segmentation model, obtaining multiple different segmentation objects; among them, the segmentation objects correspond to the spatial positions of the target patches.

[0048] In one implementation, the visual segmentation model can be the SAM (Segment Anything Model) segmentation model. The SAM segmentation model includes an Image Encoder module, a Prompt Encoder module, and a Mask Decoder, as shown in Figure 2 shown.

[0049] Among them, the Image Encoder module uses a MAE-pre-trained vision transformer to perform feature dimensionality reduction on the remote sensing image to be predicted. The Image Encoder module uses a MAE-pre-trained Vision Transformer (ViT), which can minimally adapt to processing high-resolution input images. Among them, ViT is used to perform image embedding on the image.

[0050] The Prompt Encoder module includes a sparse prompt unit and a dense prompt unit. Among them, the sparse prompt unit represents the prompt points and prompt boxes through positional encoding, and adds them to the learned embeddings of the prompt types and free-form texts obtained from the multi-modal neural network; the dense prompt unit uses convolutional embeddings and performs element-wise summation with the result of the feature dimensionality reduction. That is, the Prompt Encoder module has two prompt methods: sparse (prompt points, boxes, text) and dense (mask). The sparse representation represents the prompt points and prompt boxes through positional encoding, and adds them to the learned embeddings of the prompt types and free-form texts obtained from CLIP. The dense prompt uses convolutional embeddings and performs element-wise summation with the image embedding.

[0051] The mask decoder is used to map the outputs of the image encoder module and the prompt encoder module to masks. The MaskDecoder module maps the image embedding, prompt embedding, and output tokens to masks. First, a modified Transformer decoder module is adopted, followed by a dynamic mask prediction head. The Transformer decoder module contains self-attention and cross-attention modules to update all embeddings. Subsequently, the image embedding is upsampled, and the MLP maps the output tokens to a dynamic linear classifier to calculate the mask probabilities for each image location, and selects the mask with a high confidence score (IoU).

[0052] During the training of the SAM segmentation model, a linear combination of focal loss and dice loss can be used as the loss function for backpropagation. The AdamW optimizer and a linear learning rate (initially 8e-4) are used for model iteration and learning rate decay. The training input includes the image and the target points or target boxes converted from the ground truth masks, and random noise is added to the coordinates of the points and boxes.

[0053] In one implementation, the remote sensing image to be predicted is processed by a pre-trained visual segmentation model for image segmentation to obtain multiple different segmentation objects. Specifically, it may include the following steps 2.1) and 2.2):

[0054] Step 2.1), perform single-point sampling on several points in the grid above the remote sensing image to be predicted by a pre-trained visual segmentation model, generate multiple masks from each single-point prompt, and then perform quality screening and duplicate data deletion on the generated masks through non-maximum suppression;

[0055] Step 2.2), set parameters based on preset adjustable parameters, and perform image segmentation processing based on the set parameters to obtain multiple different segmentation objects.

[0056] Specifically, the object segmentation mode of the SAM model is to perform full-automatic segmentation of the image using the method of grid sampling points and non-maximum suppression based on the trained model weights and model structure. By performing input prompt sampling on several points in the grid above the image, multiple masks are generated from each single-point prompt, and then these masks are subjected to quality screening and duplicate data deletion using non-maximum suppression. There are several adjustable parameters in the automatic segmentation generation for controlling the density of sampling points and the threshold for removing low-quality or duplicate masks. In addition, the generation can be automatically run on the cropped image to improve the performance of smaller objects, and the post-processing can remove stray pixels and holes.

[0057] The above-mentioned preset adjustable parameters include sampling point density, quality screening threshold, repeat masking removal threshold, image cropping parameters, and post-processing parameters. In one implementation, the adjustable parameters can be set as follows:

[0058] 1. Sampling point density: The number of points sampled along one side of the image, which controls the density of single-point input prompts selected in the grid above the image. A higher density can better capture details in the image.

[0059] 2. Quality screening threshold: Used to screen out low-quality masks. A [0,1] threshold can be defined to measure the quality of the mask, based on the IoU (Intersection over Union) overlap metric. Masks below the threshold will be discarded.

[0060] 3. Repeat masking removal threshold: Used to remove duplicate masks to avoid redundant data. The repeatability between masks can be judged according to the IoU metric of the overlap measure.

[0061] 4. Image cropping: The image can be automatically cropped to improve the segmentation performance for smaller objects. By cropping the image, smaller objects can be enlarged, making them easier to detect and segment in the model.

[0062] 5. Post-processing: After segmentation generation, post-processing steps can be applied to improve the segmentation results. Post-processing includes operations such as removing discontinuous pixels and filling holes to make the segmentation results more accurate and continuous.

[0063] After the parameter settings are completed, the RGB three-band image to be segmented with 8-bit depth is sliced and input into SAM to obtain the segmented objects.

[0064] Step S140, overlay the target patch and the segmented object, and for each segmented object, calculate the semantic category in the cultivated land patch corresponding to its spatial position with the target patch.

[0065] In one implementation, the following steps 3.1) and 3.2) can be executed:

[0066] Step 3.1), overlay the target patch and the segmented object. For each segmented object, count the number of pixels belonging to the cultivated land patch and the background patch corresponding to the overlaid segmented object, and calculate the first pixel ratio of the cultivated land patch and the second pixel ratio of the background patch corresponding to the segmented object respectively;

[0067] Step 3.2), based on the relationship between the first pixel ratio of the cultivated land patch and the second pixel ratio of the background patch, determine the target pixel ratio of the cultivated land patch corresponding to the spatial position of each segmented object with the target patch, so as to determine the semantic category corresponding to the segmented object based on the target pixel ratio.

[0068] Specifically, when using a semantic segmentation model for image segmentation, each pixel is assigned a semantic category. However, since semantic segmentation is performed at the pixel level, there may be some inaccurate predictions or blurred boundaries. The SAM segmentation model has the ability to provide accurate and fine-grained boundaries for different objects, but it cannot determine the specific semantic category to which an object belongs.

[0069] To assign semantic information to the segmented objects, the voting module combines the predicted categories of the semantic segmentation model and the spatial information of the segmented objects to determine the final semantic category of each segmented object. This approach can improve the accuracy and reliability of classification, making the target extraction results more precise and credible.

[0070] The specific implementation steps are as follows:

[0071] 1. Use the semantic segmentation model to perform semantic segmentation on the image, generating the predicted results of the semantic category for each pixel. These predicted results represent the rough classification information of different regions in the image.

[0072] 2. Use the SAM segmentation model to segment the image, dividing the image into different object regions (segmented objects). Each segmented object corresponds to the spatial position of the predicted results of the semantic segmentation model.

[0073] 3. For each segmented object, count the proportion of pixels of each category in the predicted results of the corresponding position semantic segmentation. Traverse the pixels in the segmented object, calculate the number of pixels of each category, and calculate the proportion of this category in the entire segmented object.

[0074] 4. Select the category with the highest proportion as the final semantic category of the segmented object.

[0075] The formula is as follows:

[0076]

[0077] Among them, result2 represents the semantic category to which the current segmented object belongs, pixel_i represents the semantic category i to which the pixel belongs, count(pixel_i) represents the number of pixels of category i under the current segmented object, and sum(pixel) represents the total number of pixels. This method can better adapt to the sizes and shapes of different segmented objects and provide more reliable target extraction results.

[0078] Considering that the SAM (Segment Anything Model) has powerful segmentation capabilities, it can automatically segment images in all directions, obtain different masked objects, and has accurate and fine boundary effects. However, the SAM model obtained through multiple experiments cannot determine the semantic category information of the objects. Therefore, in the embodiments of this application, by superimposing the segmented objects on the target patches, the superimposed segmented objects can determine the corresponding semantic information, which is convenient for subsequent cultivated land identification.

[0079] Step S150, extract the cultivated land distribution information of each segmented object based on the semantic category.

[0080] Since the segmented objects obtained after segmentation by the SAM segmentation model are fragmented partial objects, in order to obtain the cultivated land information of the entire remotely sensed image to be predicted, the segmented objects can be stitched according to the geographical location information, and the stitched target objects can be converted into vector patches to obtain the distribution range of the cultivated land.

[0081] The embodiments of this application provide a specific process for cultivated land extraction based on remotely sensed images. See Figure 3 As shown, the method includes, after training the semantic segmentation model, performing semantic classification prediction on the remotely sensed image to be predicted through the semantic segmentation model, configuring the SAM parameters, and segmenting the remotely sensed image to be predicted through the SAM segmentation model after parameter configuration. After obtaining the results of the two model processes, they are superimposed, and pixel analogy voting is performed on the superimposed result to determine various types of patches in the remotely sensed image to be predicted. By stitching various types of patches, the final cultivated land information is obtained.

[0082] Furthermore, this application also provides a specific example, including the following steps:

[0083] 1. Image preprocessing

[0084] Select the high-resolution 2 satellite remotely sensed image of a certain city in the first half of 2022. See Figure 4 As shown, after preprocessing such as radiometric correction, orthorectification, and fusion, a cloud-free composite image with a spatial resolution of 0.8 meters is obtained, with bands of the blue band, green band, and red band, and the data type is 8-bit.

[0085] 2. Sample dataset production

[0086] Overlay the image with the labels, crop it with a 256*256 pixel window size at an overlap rate of 30%, and then perform data augmentation operations such as rotation and flipping to obtain the sample dataset. See Figure 5 As shown.

[0087] 3. Semantic segmentation model training and prediction

[0088] In the embodiments of the present application, the semantic segmentation model SegFormer is used to obtain semantic annotations of pixels. The sample data set is input into the semantic segmentation model, and the optimal training weight file is obtained through multiple rounds of iteration. Then, the weights are loaded into the network model, and the preprocessed image to be predicted is cropped into image patches of 256*256 pixel size according to an overlapping rate of 50% and input into the model for cultivated land prediction to obtain cultivated land patches, as Figure 6 shown, where the white in the figure represents the cultivated land patches and the black represents the background patches.

[0089] 4. SAM Fully Automatic Object Segmentation

[0090] The fully automatic segmentation mode of the SAM model is adopted. The sampling point density parameter is set to 64, the quality screening threshold is 0.86, the repeated mask removal threshold is set to 0.92, the image cropping layer is set to 1, and unconnected regions and holes with an area less than 50 pixels are removed during post-processing. The preprocessed image to be predicted is cropped into image patches of 256*256 pixel size according to an overlapping rate of 50% for object segmentation to obtain highly accurate and detailed boundary objects, as Figure 7 shown.

[0091] 5. Pixel Category Proportion Voting

[0092] The prediction results of the semantic segmentation model and the SAM segmentation objects are overlaid. For each segmentation object, the number of pixels belonging to cultivated land and background is counted, and their proportions are calculated. According to the pixel category with the largest proportion, that is, cultivated land or background, the final semantic label of the object is determined. See the voting results in Figure 8 shown.

[0093] 6. Stitching Prediction Results and Converting to Vector Patches

[0094] All prediction results of 256*256 pixel size are stitched according to the geospatial position relationship and then converted into vector patches. See Figure 9 shown to obtain the distribution range of cultivated land.

[0095] In summary, the embodiments of the present application are applicable to most medium and high-resolution multispectral remote sensing image data at home and abroad. This method can make full use of the semantic annotation ability of the semantic segmentation model and the segmentation ability of SAM (Segment Anything Model). Combining the two can effectively solve the problems of category and boundary extraction caused by complex and diverse cultivated land characteristics in cultivated land extraction. It has advantages such as strong applicability, high accuracy, and high automation, and can provide basic technical support for fields such as land survey, ecological change, disaster detection, and assessment.

[0096] Based on the above method embodiments, the embodiments of the present application also provide a cultivated land extraction device based on remote sensing images. SeeFigure 10 As shown, the device mainly includes the following parts:

[0097] An image acquisition module 101, configured to acquire the remote sensing image to be predicted;

[0098] A cultivated land prediction module 102, configured to perform cultivated land prediction on the remote sensing image to be predicted through a pre-trained semantic segmentation model, and obtain the target patches corresponding to the remote sensing image to be predicted, where the target patches include cultivated land patches;

[0099] An image segmentation module 103, configured to perform image segmentation processing on the remote sensing image to be predicted through a pre-trained visual segmentation model, and obtain a plurality of different segmentation objects; wherein, the spatial positions of the segmentation objects correspond to those of the target patches;

[0100] An overlay calculation module 104, configured to overlay the target patches and the segmentation objects, and for each segmentation object, calculate the semantic category in the cultivated land patches at the spatial position corresponding to the target patches;

[0101] A cultivated land extraction module 105, configured to extract the cultivated land distribution information of each segmentation object based on the semantic category.

[0102] In an optional implementation manner, the above-mentioned cultivated land prediction module 102 is further configured to: extract the feature map of the remote sensing image to be predicted through a pre-trained semantic segmentation model based on multi-layer convolution and pooling operations; perform upsampling and convolution operations on the feature map to restore the feature map to the resolution of the original remote sensing image to be predicted, and predict the semantic category to which each pixel in the remote sensing image to be predicted belongs, and each semantic category is used to be represented in the form of patches.

[0103] In an optional implementation manner, the visual segmentation model is a SAM segmentation model, and the SAM segmentation model includes an image encoder module, a prompt encoder module, and a mask decoder; wherein, the image encoder module uses a vision transformer pre-trained by MAE to perform feature dimensionality reduction on the remote sensing image to be predicted; the prompt encoder module includes a sparse prompt unit and a dense prompt unit, wherein the sparse prompt unit represents the prompt points and prompt boxes through position encoding, and adds them to the learned embeddings of the prompt types and free-form texts obtained from the multi-modal neural network; the dense prompt unit uses convolutional embeddings and performs element-wise summation with the result of the feature dimensionality reduction; the mask decoder is used to map the outputs of the image encoder module and the prompt encoder module to masks.

[0104] In an alternative embodiment, the above-mentioned image segmentation module 103 is configured to: perform single-point sampling processing on several points in the grid above the remote sensing image to be predicted through a pre-trained visual segmentation model, generate multiple masks from each single-point prompt, and then perform quality screening and duplicate data deletion on the generated masks through non-maximum suppression; perform parameter setting based on preset adjustable parameters, so as to perform image segmentation processing based on the set parameters to obtain multiple different segmentation objects.

[0105] In an alternative embodiment, the preset adjustable parameters include sampling point density, quality screening threshold, duplicate mask removal threshold, image cropping parameters, and post-processing parameters.

[0106] In an alternative embodiment, the above-mentioned overlay calculation module 104 is further configured to: overlay the target patch and the segmentation object, and for each segmentation object, count the number of pixels belonging to the cultivated land patch and the background patch corresponding to the overlaid segmentation object, and calculate the first pixel ratio of the cultivated land patch and the second pixel ratio of the background patch corresponding to the segmentation object respectively; based on the relationship between the first pixel ratio of the cultivated land patch and the second pixel ratio of the background patch, determine the target pixel ratio of the cultivated land patch in the spatial position corresponding to the target patch for each segmentation object, so as to determine the semantic category corresponding to the segmentation object based on the target pixel ratio.

[0107] In an alternative embodiment, the above-mentioned device further includes: a splicing module, configured to splice the segmentation objects according to the geographical location information, and convert the spliced target object into a vector patch to obtain the distribution range of cultivated land.

[0108] The device for extracting cultivated land based on remote sensing images provided by the embodiments of the present application has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the embodiments of the device for extracting cultivated land based on remote sensing images, reference may be made to the corresponding content in the foregoing method embodiments of the device for extracting cultivated land based on remote sensing images.

[0109] Embodiments of the present application further provide an electronic device, as Figure 11 shown, which is a schematic structural diagram of the electronic device. Among them, the electronic device 100 includes a processor 111 and a memory 110. The memory 110 stores computer-executable instructions that can be executed by the processor 111, and the processor 111 executes the computer-executable instructions to implement any one of the foregoing methods for extracting cultivated land based on remote sensing images.

[0110] In Figure 11 the shown embodiment, the electronic device further includes a bus 112 and a communication interface 113. Among them, the processor 111, the communication interface 113, and the memory 110 are connected through the bus 112.

[0111] Among them, the memory 110 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 113 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 112 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 112 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a bidirectional arrow is used in Figure 11 , but it does not mean that there is only one bus or one type of bus.

[0112] The processor 111 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 111. The above-mentioned processor 111 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor 111 reads the information in the memory and combines its hardware to complete the steps of the above-mentioned method for extracting cultivated land based on remote sensing images in the foregoing embodiments.

[0113] The embodiment of the present application also provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the above-mentioned cultivated land extraction method based on remote sensing images. For specific implementation, reference can be made to the foregoing method embodiments and will not be elaborated herein.

[0114] The computer program product of the cultivated land extraction method, device, equipment and medium based on remote sensing images provided by the embodiments of the present application includes a computer-readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated herein.

[0115] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0116] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0117] In the description of the present application, it should be noted that the terms "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0118] In the description of the present application, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "installed", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for extracting cultivated land based on remote sensing images, characterized in that, The method includes: Obtaining the remote sensing image to be predicted; Performing cultivated land prediction on the remote sensing image to be predicted through a pre-trained semantic segmentation model to obtain target patches corresponding to the remote sensing image to be predicted, where the target patches include cultivated land patches; Performing image segmentation processing on the remote sensing image to be predicted through a pre-trained visual segmentation model to obtain multiple different segmentation objects; wherein, the spatial positions of the segmentation objects correspond to those of the target patches; Overlaying the target patches and the segmentation objects, and for each segmentation object, calculating the semantic category in the cultivated land patches at the spatial position corresponding to the target patch; Extracting the cultivated land distribution information of each segmentation object based on the semantic category; The visual segmentation model is a SAM segmentation model, and the SAM segmentation model includes an image encoder module, a prompt encoder module, and a mask decoder; Among them, the image encoder module uses a vision transformer pre-trained by MAE to perform feature dimensionality reduction on the remote sensing image to be predicted; The prompt encoder module includes a sparse prompt unit and a dense prompt unit. The sparse prompt unit represents the prompt points and prompt boxes through positional encoding and adds them to the learned embeddings of the prompt types and free-form texts obtained from the multi-modal neural network; the dense prompt unit uses convolutional embeddings and performs element-wise summation with the result of the feature dimensionality reduction; The mask decoder is used to map the outputs of the image encoder module and the prompt encoder module to masks.

2. The method for extracting cultivated land based on remote sensing images according to claim 1, wherein Performing cultivated land prediction on the remote sensing image to be predicted through a pre-trained semantic segmentation model to obtain target patches corresponding to the remote sensing image to be predicted, including: Extracting the feature map of the remote sensing image to be predicted through a pre-trained semantic segmentation model based on multi-layer convolution and pooling operations; Performing upsampling and convolution operations on the feature map to restore the feature map to the resolution of the original remote sensing image to be predicted, and predicting the semantic category to which each pixel in the remote sensing image to be predicted belongs, and each semantic category is used to be characterized in the form of patches.

3. The cultivated land extraction method based on remote sensing images according to claim 1, wherein Performing image segmentation processing on the remote sensing image to be predicted through a pre-trained visual segmentation model to obtain multiple different segmentation objects, including: Performing single-point sampling processing on several points in the grid above the remote sensing image to be predicted through a pre-trained visual segmentation model, generating multiple masks from each single-point prompt, and then performing quality screening and duplicate data deletion on the generated masks through non-maximum suppression; Performing parameter setting based on preset adjustable parameters, and performing image segmentation processing based on the set parameters to obtain multiple different segmentation objects.

4. The method for extracting cultivated land based on remote sensing images according to claim 3, wherein The preset adjustable parameters include sampling point density, quality screening threshold, duplicate mask removal threshold, image cropping parameters, and post-processing parameters.

5. The cultivated land extraction method based on remote sensing images according to claim 1, wherein The target patches further include background patches; overlaying the target patches and the segmentation objects, and for each segmentation object, calculating the semantic category in the spatial position corresponding to the target patch, including: Overlay the target patch and the segmentation object. For each segmentation object, count the number of pixels belonging to the cultivated land patch and the background patch corresponding to the overlaid segmentation object, and calculate the first pixel ratio of the cultivated land patch and the second pixel ratio of the background patch corresponding to the segmentation object respectively; Based on the relationship between the first pixel ratio of the cultivated land patch and the second pixel ratio of the background patch, determine the target pixel ratio of the cultivated land patch in the spatial position corresponding to each segmentation object and the target patch, so as to determine the semantic category corresponding to the segmentation object based on the target pixel ratio.

6. The method for extracting cultivated land based on remote sensing images according to claim 1, wherein The method further includes: Stitch the segmentation objects according to the geographical location information, and convert the stitched target object into a vector patch to obtain the distribution range of cultivated land.

7. An arable land extraction device based on remote sensing images, characterized in that, The device includes: An image acquisition module for acquiring the remote sensing image to be predicted; A cultivated land prediction module for predicting the cultivated land of the remote sensing image to be predicted through a pre-trained semantic segmentation model to obtain the target patch corresponding to the remote sensing image to be predicted, where the target patch includes the cultivated land patch; An image segmentation module for performing image segmentation processing on the remote sensing image to be predicted through a pre-trained visual segmentation model to obtain a plurality of different segmentation objects; wherein, the segmentation objects correspond to the spatial positions of the target patches; An overlay calculation module for overlaying the target patch and the segmentation object, and calculating the semantic category in the cultivated land patch in the spatial position corresponding to the target patch for each segmentation object; A cultivated land extraction module for extracting the cultivated land distribution information of each segmentation object based on the semantic category; The visual segmentation model is a SAM segmentation model, and the SAM segmentation model includes an image encoder module, a prompt encoder module, and a mask decoder; Among them, the image encoder module uses a vision transformer pre-trained by MAE to perform feature dimensionality reduction on the remote sensing image to be predicted; The prompt encoder module includes a sparse prompt unit and a dense prompt unit. Among them, the sparse prompt unit represents the prompt point and the prompt box through position encoding, and adds them to the learned embeddings of the prompt type and free-form text obtained from the multi-modal neural network; the dense prompt unit uses convolutional embeddings and performs element-wise summation with the result of the feature dimensionality reduction; The mask decoder is used to map the outputs of the image encoder module and the prompt encoder module to masks.

8. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores computer executable instructions that can be executed by the processor. The processor executes the computer executable instructions to implement the method for extracting cultivated land based on remote sensing images according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer executable instructions. When the computer executable instructions are called and executed by the processor, the computer executable instructions cause the processor to implement the method for extracting cultivated land based on remote sensing images according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Remote sensing image deep learning segmentation method and system based on superpixels and watershed

    CN114898216A