Method and device for extracting key points of plot polygon based on edge and edge direction
Through deep learning technology, network structures such as VisionTransformer and directional convolution blocks are used to extract the edge and edge direction of the cultivated land plot in remote sensing images, and combined with the connectivity attention module, the problems of edge fracture and key point redundancy are solved, and more complete and accurate cultivated land polygons are generated.
Patent Information
- Application Number
- CN202510153879.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems of edge fracture and key points redundancy in vectorization in edge extraction of cultivated land plots in remote sensing images, resulting in incomplete or distorted farmland polygons extracted.
Using a deep learning-based method, the edge and edge directions of agricultural plots are extracted through network structures such as VisionTransformer and directional convolution blocks, and combined with the connectivity attention module, edge fractures are repaired and key point extraction is optimized to generate a complete agricultural plot polygon.
It effectively solves the problems of edge fracture and key point redundancy, and the generated cultivated land polygons are more complete and accurate, meeting application requirements.
Smart Images

Figure CN120070913A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a method and device for extracting key points of plot polygons based on edges and edge directions. Background Art
[0002] With the continuous improvement of the spatial resolution of satellite remote sensing images, the visual features of ground objects presented by them are more refined. Sub-meter high-resolution remote sensing images can basically define the forms and types of ground objects, making it possible for fine-grained classification and precise monitoring of remote sensing ground objects. The object of cultivated land extraction in agricultural remote sensing has changed from large-scale and large-area intensive cultivated land areas in the past to more precise cultivated land plots.
[0003] At present, the extraction of cultivated land has shifted from traditional machine learning to methods based on Convolutional Neural Networks (CNNs). The deep convolutional neural network can automatically learn representative features from the sample set, which provides a new idea for the extraction of plots in high-resolution remote sensing images. Currently, the automatic extraction methods for remote sensing cultivated land plots include methods based on edge detection and methods based on semantic segmentation. Among them, the methods based on edge detection divide cultivated land plots by locating plot boundaries. For example, methods such as BDCN and DexiNed divide image pixels into boundaries and non-boundaries. However, since edge detection only considers the gradient information within a local small window, the detection ability is limited, which may lead to the situation where edges cannot be connected. The methods based on semantic segmentation utilize the texture and spectral information of the image, and based on the homogeneity of pixels, aggregate from bottom to top to form image objects to be extracted. Such as methods in the DeepLab series and U-Net series. However, since these methods will lose some image details when the network performs downsampling operations, a large number of internal boundaries of the plots will be lost in the extracted semantic results, resulting in the problem of adhesion between adjacent plots.
[0004] In addition, in applications such as remote sensing mapping, it is often necessary to convert raster images into vectors for effective use. Existing vectorization algorithms, such as those provided by open-source geospatial data abstraction libraries (GDAL) and commercial software such as ArcGIS, rely on boundary tracing for vectorization. These methods trace and connect edges pixel by pixel in raster data to generate polygon results. There are often a large number of redundant nodes and connecting edges in the vectorization results of semantic segmentation methods. Although simplification algorithms (such as Douglas-Peucker) can reduce redundancy, they may over-simplify the boundaries, resulting in distortion of the plot results. In some cases, these algorithms cannot fully remove redundant points. Therefore, the cultivated land polygons generated by current methods cannot meet the application requirements. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to make full use of edge semantics, direction, and geometric information to solve the edge break problem in the extraction of closed edges of cultivated land in remote sensing images and the problem of redundant key points in the vectorization process, and a method for extracting key points of plot polygons and constructing plot polygon vectors based on edges and edge directions is proposed.
[0006] The present invention focuses on extracting the boundaries and key nodes of cultivated land through the edges and edge directions of cultivated land based on deep learning, and achieving the goal of generating plot polygon vectors.
[0007] The present invention provides the following technical solutions.
[0008] Based on the starting point of strengthening context information learning, the present invention selects the Vision Transformer (ViT) as the feature extraction network. The output result after passing through the Transformer encoder is used to obtain the edges and edge directions of agricultural plots through the edge branch and the edge direction branch. A Directional Convolution Block (DCB) and a Connectivity Attention Module (CAM) are designed to extract complete edges and edge directions. For the edge direction branch, a step-by-step upsampling method is adopted, and the size of the feature map is gradually restored through four directional convolution blocks, and the linear features of the edges are extracted during this process. For the edge branch, we adopt the bi-directional feature aggregation strategy of the Bi-directional Multi Level Aggregation (BiMLA) decoder using EDTER, and introduce a connectivity attention mechanism CAM. The present invention combines the model output with post-processing, aiming to convert the results of edge detection and direction prediction into complete plot polygons. First, according to the direction information predicted by the edge direction branch, the broken edge points are extended according to the edge direction to repair the broken edges and form a more complete boundary. After the edge connection is completed, through key point extraction and edge connectivity information, the edge lines are converted into polygon representations while retaining the geometric features of the boundary, generating complete agricultural plot polygons. The design of the above network and post-processing methods aims to simultaneously learn the semantic features and geometric features of farmland edges and obtain complete plot polygon results.
[0009] The first aspect of the present invention relates to a method for extracting key points of plot polygons based on edges and edge directions, comprising the following steps:
[0010] Step 1, Prepare the remote sensing image plot dataset: The dataset includes the original remote sensing image and the corresponding label image of the image. The label image is in shp format. Convert the label image. The remote sensing image contains 3 bands, and the label image includes 1 type. Divide the dataset into a training set, a validation set, and a test set, and preprocess the training set and the validation set. Thus, a training set TR and a validation set VA containing images of size H×W can be obtained, where H and W represent the height and width of the image respectively;
[0011] Step 2, Construct a multi-task extraction network based on edges and edge directions: This network consists of a Vision Transformer, a Bidirectional Multi-level Aggregation Decoder BiMLA, a Direction Convolution Block DCB, and a Connected Attention Module CAM;
[0012] Step 3, Model training and selection of the optimal model: Use the training set TR and the validation set VA obtained in Step 1 as training data and validation data respectively to train the network constructed in Step 2, and select the model with the highest validation accuracy as the output model M;
[0013] Step 4, Edge and edge direction prediction: Input the image I of the test set into the model M obtained in Step 3 for prediction to obtain the plot edge result E and the plot edge direction result D;
[0014] Step 5, Post-processing: Extract the direction turning points from the plot edge direction result obtained in Step 4, and use the direction turning points to optimize the plot edge result E obtained in Step 4 to obtain the plot polygon R;
[0015] Furthermore, the specific implementation of Step 1 includes the following sub-steps.
[0016] Step 1.1, Collect the cultivated land extraction dataset: The dataset contains the original remote sensing image with 3 bands and the corresponding label image of 1 type. Divide the dataset into a training set, a validation set, and a test set;
[0017] Step 1.2, Data cropping: Select the red, green, and blue bands of the images in the training set and the validation set, and crop them into images of size H×W, where H and W represent the height and width of the image respectively;
[0018] Step 1.3, Label conversion: Select the labels in the training set and the validation set, and convert them into 0-1 edge labels and 0-180 / A edge direction labels. H and W represent the height and width of the image respectively, and A represents the interval of the edge direction angle. Represent the category of the edge to which a certain pixel position belongs in the form of OneHot encoding. If the value of the m-th channel at a certain position of the label is 1, it means that the edge direction at this position belongs to the m-th category, indicating that the angle between this edge and the horizontal direction is within the range of 0 - A*m.
[0019] Step 1.4, Data Augmentation: The training set obtained in Step 1.2 is augmented by rotation, Gaussian noise, mirroring, and translation.
[0020] Step 1.5, Normalize the validation set obtained in Step 1.2 and the training set obtained in Step 1.3 to obtain the training set TR and the validation set VA.
[0021] Furthermore, the specific implementation of Step 2 includes the following sub-steps.
[0022] Step 2.1, Construct the network input layer: The number of bands of the input image is 3, and the edge labels with values from 0 to 1 and the labels from 0 to 180 / A have dimensions of H×W.
[0023] Step 2.2, Construct the feature extraction network: Use Vision Transformer (ViT) for feature extraction. The standard Transformer encoder consists of L Transformer blocks. Each block has a multi-head self-attention operation (MSA), a multi-layer perceptron (MLP), and two layer normalization steps (LN). The query elements of the multi-scale deformable self-attention layer are the pixels of the multi-scale features, and the key elements are the LK points sampled from the multi-scale features instead of all spatial positions, so the computational complexity can be reduced and the convergence speed can be accelerated. In addition, a residual connection layer is applied after each block.
[0024] Step 2.3, Construct the direction convolution block DCB: The direction convolution module contains four direction convolutions in different directions (horizontal, vertical, left diagonal, and right diagonal). Given the input tensor X ∈ R H×W×C , where H, W, and C represent the height, width, and number of channels of the image respectively, X first passes through a 1×1 convolutional layer and then is fed into four parallel paths. Each path contains a direction convolution in a specific direction to extract the edge features in the corresponding direction. Subsequently, the output feature maps of these four direction convolutions are concatenated, and after an upsampling operation and an additional 1×1 convolutional layer, the output of the direction convolution module is finally obtained.
[0025]
[0026] Let w ∈ R 2k+1 be the direction convolution filter of size 2k + 1, D = (D h , D w ) be the direction of the filter w, and Z D ∈ R H ×W×CRepresents the result of the directional convolution. The directional convolution can be defined as: where X*w represents the convolution operation. D is the direction vector of the directional convolution, taking values (0,1), (1,0), (1,1), and (-1,1), which are used for horizontal, vertical, left diagonal, and right diagonal convolutions respectively. For the filter w, we set k = 4 so that each directional convolution has 9 parameters, the same as a 3×3 convolution kernel.
[0027] Step 2.4, construct the Connected Attention Module CAM: For the edge labels, generate the connectivity cube where C o represents the number of samples of a given pixel and its adjacent pixels. For a two-dimensional edge image, C o is denoted as 8, and O in the connectivity graph i,j,c represents the connectivity between a pixel and a specific pixel, where i,j represent the spatial positions of the pixels. c represents the position of its adjacent pixel. If two pixels are connected, then O i,j,c = 1, indicating that these two pixels are connected farmland boundaries. Finally, calculate and splice the connectivity information of the eight neighborhoods to obtain the final connectivity graph.
[0028] As Figure 4 shown, the input tensor is sequentially input into a 3×3 convolution with an expansion rate of 1, and then refer to the Efficient Channel Attention (ECA) block to make full use of the connectivity, and recalibrate the predicted connectivity cube by using the channel attention mechanism. The input feature map is fed into this block after global average pooling, and we can obtain a vector ranging from (0,1), where each factor is multiplied by the corresponding channel in the input feature map. The final output of CAM is a connectivity cube of H×W×C o Finally, supervised learning is carried out using the connectivity graph ground truth generated by the label.
[0029] Furthermore, the specific implementation of step 3 includes the following sub-steps,
[0030] Step 3.1, take the training set TR and validation set VA obtained in step 1 as training data and validation data respectively and input them into the network constructed in step 2;
[0031] Step 3.2, set the hyperparameters of the model: Epochs, BatchSize, LearningRate, and Optimizer;
[0032] Step 3.3, calculate the loss between the edge result and the edge direction result of step 2 and the true label category respectively. The loss function of this model is the joint loss.
[0033] L = L seg + αL con + βLclass (2)
[0034]
[0035] where α and β are two constants. N represents the number of elements in the H×W slice, and y i is the ground truth of the drawing edge of the given pixel at position i, is the predicted probability of the corresponding segmentation branch. C o represents the number of sampled neighboring pixels of the given pixel, is the ground truth indicating whether there is connectivity between the given pixel at position i and the neighboring pixel at position c, is the predicted connectivity corresponding to the connectivity branch. C represents the number of categories in the edge direction, and z i,c is the ground truth label of the sample for category c, is the predicted probability of the sample for category c.
[0036] Step 3.4: Through iterative training, the model selects the model with the highest F1 value as the final prediction model M. The calculation method of F1 is shown in the formula:
[0037]
[0038] Precision is the precision rate, which represents the proportion of actual positive samples among the samples predicted as positive. The calculation method is shown in the formula:
[0039]
[0040] where TP is the number of pixels of the positive samples predicted correctly, FP is the number of pixels of the positive samples predicted incorrectly, and FP is the number of pixels of the negative samples predicted incorrectly.
[0041] Recall is the recall rate, which represents the proportion of the actual positive samples among the samples predicted as positive in the whole sample. The calculation method is shown in the formula:
[0042]
[0043] where FN is the number of pixels of the negative samples predicted incorrectly.
[0044] Furthermore, the specific implementation of Step 5 includes the following sub-steps,
[0045] Step 5.1: The plot edge output in Step 4 is subjected to non-maximum suppression, binarization, and skeleton extraction to obtain a refined plot edge result.
[0046] Step 5.2: Using the refined plot edge result in Step 5.1 to mask the plot edge direction output in Step 4 to obtain a refined edge direction.
[0047] Step 5.3: Define a pixel point with a pixel value of 1 in the eight-connected domain as a break point, and find the break points on the refined plot edge result in Step 5.1.
[0048] Step 5.5: Traverse the break points, obtain the edge direction at the break point position through the refined edge direction output in Step 5.2, and extend the edge according to the edge direction until the extension distance exceeds the threshold th or another edge is encountered, thereby obtaining an improved closed edge.
[0049] Step 5.6: Find the direction mutation points on the refined edge direction output in Step 2 and define them as key points.
[0050] Step 5.7: Regard the closed edge output in Step 5.5 as the connectivity information of the key points, and connect the key points output in Step 5.6 to obtain the plot polygon.
[0051] Thus, the predicted result R of the plot polygon of the image can be obtained. The second aspect of the present invention relates to a device for extracting key points of a plot polygon based on edges and edge directions, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the method for extracting key points of a plot polygon based on edges and edge directions of the present invention.
[0052] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for extracting key points of a plot polygon based on edges and edge directions of the present invention.
[0053] The advantages of the present invention are:
[0054] 1) The present invention proposes a method for plot extraction, which transforms it into tasks of extracting edges, edge directions, and key points. Using the idea of multi-task learning, it simultaneously identifies the edge information and edge direction information of the plot, identifies the vector key points of the plot based on the edge direction, and connects the key points under the guidance of the edge to form the final vector. On the basis of limited edge extraction results, high-quality plot polygons can also be obtained. It improves the problem of unconnected edges that are prone to occur in cultivated land extraction and provides an idea for extracting key points in remote sensing images;
[0055] 2) The present invention designs a direction convolution block (DCB) and a connectivity attention module (CAM). The direction convolution block improves the model's perception ability in terms of direction and can effectively reduce the noise in edge direction classification. The connectivity attention module enhances the connectivity of edges. Finally, through the directionality of the edges, the broken edges are connected, alleviating the problem of edge breakage and improving the quality of plot extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flowchart of a method for extracting key points of plot polygons based on edges and edge directions;
[0057] Figure 2 is a network structure diagram of the plot edge and edge direction in the embodiment of the present invention;
[0058] Figure 3 is a flowchart of the direction convolution block in the embodiment of the present invention;
[0059] Figure 4 is a flowchart of the connectivity attention module in the embodiment of the present invention;
[0060] Figure 5 is the input remote sensing image in the embodiment of the present invention;
[0061] Figure 6 is the preliminary edge result map of the plot in the remote sensing image in the embodiment of the present invention;
[0062] Figure 7 is the edge direction result map of the plot in the remote sensing image in the embodiment of the present invention;
[0063] Figure 8 is the closed edge extraction result map of the plot in the remote sensing image in the embodiment of the present invention;
[0064] Figure 9 is the plot polygon extraction result map of the remote sensing image in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] The present invention will be further described below through embodiments with reference to the accompanying drawings, which are only for explaining the present invention and in no way limit the present invention.
[0066] Embodiment 1
[0067] This embodiment provides a method for extracting key points of plot polygons based on edges and edge directions, including the following steps:
[0068] Step 1, Prepare the iFlytek Remote Sensing Image Cultivated Land Dataset: The dataset includes the original remote sensing images and the corresponding label images for the images. The label images are in shp format. Convert the label images. The remote sensing images contain 3 bands, and the label images include 1 type. Divide the dataset into a training set, a validation set, and a test set, and preprocess the training set and the validation set. Thus, a training set TR and a validation set VA containing images of size 512×512 can be obtained;
[0069] Step 2, Construct a Multi-Task Extraction Network Based on Edges and Edge Directions: This network is composed of a Vision Transformer, a Bidirectional Multi-Level Aggregation Decoder BiMLA, a Direction Convolution Block DCB, and a Connected Attention Module CAM;
[0070] Step 3, Model Training and Optimal Model Selection: Use the training set TR and the validation set VA obtained in Step 1 as training data and validation data respectively to train the network constructed in Step 2, and select the model with the highest validation accuracy as the output model M;
[0071] Step 4, Edge and Edge Direction Prediction: Input the image I of the test set into the model M obtained in Step 3 for prediction to obtain the plot edge result E and the plot edge direction result D; Figure 5 is one of the original images of the test set in the embodiments of the present invention; The obtained preliminary edge result is as Figure 6 shown; The obtained edge direction result is as Figure 7 shown.
[0072] The specific process of the implementation steps of the embodiments of the present invention is as Figure 1 shown.
[0073] Furthermore, the specific implementation of Step 1 includes the following sub-steps,
[0074] Step 1.1, Use the iFlytek Cultivated Land Dataset: The remote sensing images in this dataset contain 3 bands, and the labels are in shpfile format. Select 1120 sampling points and divide them into a training set, a validation set, and a test set according to 8:1:1.
[0075] Step 1.2, Data Cropping: Crop it into images and shpfiles of 512×512, and the images retain the three RGB bands; Thus, a training set containing 896 images of size 512×512, a validation set containing 112 images of size 512×512, and a test set containing 112 images of size 512×512 can be obtained;
[0076] Step 1.3, Label Conversion: Select the labels in the training set and validation set, and convert the shpfile into edge labels from 0 to 1 and edge direction labels from 0 to 180 / A. Here, A is set to 30, indicating that the 180 degrees are divided into 6 categories with an interval of 30 degrees. The corresponding value represents the interval to which the angle between the edge where the point is located and the horizontal direction belongs, and 0 represents the background class. Represent the category of the edge to which a certain pixel position belongs in the form of OneHot encoding. If the value of the m-th channel at a certain position of the label is 1, it means that the edge direction at this position belongs to the m-th category, indicating that the angle between this edge and the horizontal direction is within the range of 0 - A*m.
[0077] Step 1.4, Data Augmentation: Perform data augmentation on the training set obtained in Step 1.3 by means of rotation, Gaussian noise, mirroring, and translation;
[0078] Step 1.5, Normalize the validation set obtained in Step 1.2 and the training set obtained in Step 1.3, and thus the edge dataset and edge direction dataset can be obtained respectively.
[0079] Furthermore, the specific process of Step 2 is as Figure 2 shown, and the specific implementation includes the following sub-steps,
[0080] Step 2.1, Construct the network input layer: The number of bands of the input image is 3, the edge label is from 0 to 1, the number of categories of the edge direction label is 7, and the image size is 512×512;
[0081] Step 2.2, Construct the feature extraction network: Use Vision Transformer for feature extraction. The standard Transformer encoder consists of L transformer blocks. Each block has a multi-head self-attention operation (MSA), a multi-layer perceptron (MLP), and two layer normalization steps (LN). The query elements of the multi-scale deformable self-attention layer are the pixels of the multi-scale features, and the key elements are the LK points sampled from the multi-scale features, rather than all spatial positions, so the computational complexity can be reduced and the convergence speed can be accelerated. In addition, a residual connection layer is applied after each block.
[0082] Step 2.3, Construct the Direction Convolution Block DCB: The direction convolution module contains direction convolutions in 4 different directions (horizontal, vertical, left diagonal, right diagonal). Given the input tensor X ∈ R H×W×C , where H, W, and C represent the height, width, and number of channels of the image respectively. X first passes through a 1×1 convolutional layer and then is sent into 4 parallel paths. Each path contains a direction convolution in a specific direction for extracting edge features in the corresponding direction. Subsequently, the output feature maps of these 4 direction convolutions are concatenated, and after an upsampling operation and an additional 1×1 convolutional layer, the output of the direction convolution module is finally obtained.
[0083]
[0084] Let \(w\in\mathbb{R}^{2k + 1}\) be a directional convolution filter of size \(2k + 1\), \(X\in\mathbb{R}\) H×W×C be the direction of the filter \(w\), \(Z\) D \(\in\mathbb{R}\) H×W×C represent the result of the directional convolution. The directional convolution can be defined as: where \(X*w\) represents the convolution operation. \(D\) is the direction vector of the directional convolution, taking values \((0,1)\), \((1,0)\), \((1,1)\) and \((-1,1)\), which are used for horizontal, vertical, left diagonal and right diagonal convolutions respectively. For the filter \(w\), we set \(k = 4\) so that each directional convolution has 9 parameters, the same as a \(3\times3\) convolution kernel.
[0085] Step 2.4, construct the connected attention module CAM: For the edge labels, generate the connectivity cube where \(C\) o represents the number of samples of a given pixel and its adjacent pixels. For a two-dimensional edge image, \(C\) o is denoted as 8, and \(O\) in the connectivity graph i,j,c represents the connectivity between a pixel and a specific pixel, where \(i,j\) represent the spatial positions of the pixels. \(c\) represents the positions of its adjacent pixels. If two pixels are connected, then \(O\) i,j,c \(= 1\), indicating that these two pixels are connected farmland boundaries. Finally, the connectivity information of the eight neighborhoods is calculated and concatenated to obtain the final connectivity graph.
[0086] As Figure 4 shown, the input tensor is sequentially input into a \(3\times3\) convolution with an expansion rate of 1, and then refer to the efficient channel attention (ECA) block in the reference to make full use of the connectivity, and recalibrate the predicted connectivity cube by using the channel attention mechanism. The input feature map is fed into this block after global average pooling, and we can obtain a vector in the range \((0,1)\), where each factor is multiplied by the corresponding channel in the input feature map. The final output of CAM is a connectivity cube of \(H\times W\times C\) o Finally, supervised learning is carried out using the ground truth of the connectivity graph generated by the label.
[0087] Furthermore, the specific implementation of step 3 includes the following sub-steps
[0088] Step 3.1, take the training set TR and the validation set VA obtained in step 1 as the training data and the validation data respectively, and input them into the network constructed in step 2;
[0089] Step 3.2, set the hyperparameters of the model: where Epochs is 200, BatchSize is 4, LearningRate is 0.001, and Optimizer is Adam;
[0090] Step 3.3, calculate the losses of the edge result and the edge direction result in Step 2 respectively with the true label categories. The loss function of this model is the joint loss.
[0091] L = L seg + αL con + βL class (2)
[0092]
[0093] where α and β are two constants, which are set to 1000 here to balance the loss function. N represents the number of elements in the H×W slice, y i is the true value of the drawing edge of the given pixel at position i, is the predicted probability of the corresponding segmentation branch. C o represents the number of sampled neighboring pixels of the given pixel, is the true value indicating whether the given pixel at position i is connected to the neighboring pixel at position c, is the predicted connectivity corresponding to the connectivity branch. C represents the number of categories in the edge direction, z i,c is the true label of the sample on category c, is the predicted probability of the sample on category c.
[0094] Step 3.4, the model is trained iteratively, and the model with the highest F1 value is selected as the final prediction model M. The calculation method of F1 is as shown in the formula:
[0095]
[0096] Precision is the precision rate, which represents the proportion of samples that are actually positive samples among the samples with positive prediction results. The calculation method is as shown in the formula:
[0097]
[0098] where TP is the number of pixels of positive samples with correct prediction, FP is the number of pixels of positive samples with wrong prediction, and FP is the number of pixels of negative samples with wrong prediction results.
[0099] Recall is the recall rate, which represents the proportion of the number of actual positive samples among the positive samples in the prediction results to the total positive samples in the whole sample. The calculation method is as shown in the formula:
[0100]
[0101] Among them, FN is the number of pixels of negative samples with incorrect prediction results.
[0102] Furthermore, the specific implementation of step 5 includes the following sub-steps
[0103] Step 5.1, perform non-maximum suppression, binarization, and skeleton extraction on the plot edge output in step 4 to obtain a refined plot edge result.
[0104] Step 5.2, use the refined plot edge result in step 5.1 to mask the plot edge direction output in step 4 to obtain a refined edge direction.
[0105] Step 5.3, define the pixel point with only one pixel value of 1 in the eight-connected domain as a break point, and find the break points on the refined plot edge result in step 5.1.
[0106] Step 5.5, traverse the break points, obtain the edge direction at the break point position through the refined edge direction output in step 5.2, and extend the edge according to the edge direction until the extension distance exceeds the threshold th or another edge is encountered, so as to obtain an improved closed edge, as Figure 8 shown.
[0107] Step 5.6, find the direction mutation points on the refined edge direction output in step 2 and define them as key points.
[0108] Step 5.7, regard the closed edge output in step 5.5 as the connectivity information of the key points, and connect the key points output in step 5.6 to obtain a plot polygon.
[0109] Thus, the plot polygon prediction result R of the image can be obtained. As Figure 9 shown.
[0110] The present invention proposes an innovative plot extraction method, aiming to directly generate plot polygons from edges and edge directions. Our method integrates the learning of edge semantics and edge directions within a multi-task framework. By identifying the change points of edge directions and enhancing edge connectivity, we can generate concise and regular plot polygons. Aiming at the problem of edge breakage in the extraction of cultivated land plots by edge detection, using the direction information of the edge, the key nodes of direction change are extracted, and the direction information is used for extension in the domain of the break nodes, effectively improving the integrity of edge detection. Aiming at the problem of redundant key points in the vectorization process, the key point extraction used in the present invention effectively alleviates this problem. Therefore, the present invention has high application value in the field of cultivated land identification in remote sensing images.
[0111] Embodiment 2
[0112] This embodiment provides a device for extracting key points of a plot polygon based on edges and edge directions, which is characterized by including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the method for extracting key points of a plot polygon based on edges and edge directions in Embodiment 1.
[0113] Embodiment 3
[0114] This embodiment relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for extracting key points of a plot polygon based on edges and edge directions described in Embodiment 1.
[0115] The above is only a description of the embodiments of the present invention, but the protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the concept of the present invention.
Claims
1. A method for extracting key points of land polygons based on edges and edge directions, characterized in that: The following steps are involved: Step 1, prepare the remote sensing image plot data set: the data set includes the original remote sensing image and the label image corresponding to the image, and the label image is in shp format; convert the label image; the remote sensing image contains 3 bands, and the label image includes 1 type. The data set is divided into training set, validation set and test set, and the training set and validation set are preprocessed, thereby obtaining a training set TR and a validation set VA containing images of size H×W, where H and W represent the height and width of the image respectively; Step 2: Construct a multi-task extraction network based on edge and edge direction: the network consists of a visual transformer ViT, a bidirectional multi-level aggregation decoder BiMLA, a directional convolution block DCB, and a connected attention module CAM; Step 3, model training and optimal model selection: The training set TR and the validation set VA are used as training data and validation data training input to the edge and edge direction based multi-task extraction network respectively, and the model with the highest validation accuracy is selected as the output model M; Step 4, edge and edge direction prediction: input the image I of the test set into the output model M for prediction, and obtain the plot edge result E and the plot edge direction result D; Step 5, post-processing: extract the direction turning point from the plot edge direction result, use the direction turning point to optimize the plot edge result E, and obtain the closed plot edge result R.
2. The method for extracting key points of land polygons based on edges and edge directions according to claim 1, characterized in that: Step 1 specifically includes: Step 1.1, collect the cultivated land extraction data set: the data set contains three bands of original remote sensing images and one type of corresponding label images, and the data set is divided into training set, validation set and test set; Step 1.2, data cropping: select the red, green, and blue bands of the images in the training set and validation set, and crop them into images of size H×W, where H and W represent the height and width of the image, respectively; Step 1.3, label conversion: select the labels in the training set and the validation set, and convert them into 0-1 edge labels and 0-180 / A edge direction labels. H and W represent the height and width of the image respectively, and A represents the interval of the edge direction angle. The edge category of a certain pixel position is represented in the form of OneHot encoding. If the value of the mth channel of a certain position of the label is 1, it means that the edge direction of the position belongs to the mth category, indicating that the angle between the edge and the horizontal direction is in the range of 0-A*m; Step 1.4, data enhancement: The training set is enhanced by rotation, Gaussian noise, mirroring and translation; Step 1.5, normalize the validation set and training set, thus obtaining the training set TR and validation set VA.
3. The method for extracting key points of land polygons based on edges and edge directions according to claim 1, characterized in that: The specific implementation of step 2 includes the following sub-steps: Step 2.1, construct the network input layer: the number of bands of the input image is 3, the labels are 0-1 edge labels and 0-180 / A, and the size is H×W; Step 2.2, build a feature extraction network: use Vision Transformer for feature extraction, the standard Transformer encoder consists of L transformer blocks; each block has a multi-head self-attention operation MSA, a multi-layer perceptron MLP and two layer normalization steps LN; the query element of the multi-scale deformable self-attention layer is the pixel of the multi-scale feature, and the key element is the LK point sampled from the multi-scale feature instead of all spatial positions; in addition, a residual connection layer is applied after each block; Step 2.3, construct the directional convolution block DCB: The directional convolution module contains directional convolutions in four different directions: horizontal, vertical, left diagonal, and diagonal. Given an input tensor X∈R H×W×C , where H, W, and C represent the height, width, and number of channels of the image, respectively. X first passes through a 1×1 convolution layer and is then sent to four parallel paths; each path contains a directional convolution in a specific direction to extract edge features in the corresponding direction; Subsequently, the output feature maps of the four directional convolutions are concatenated, and after an upsampling operation and an additional 1×1 convolution layer, the output of the directional convolution module is finally obtained; Let w∈R 2k+1 is a directional convolution filter of size 2k+1, D = (D h ,D w ) is the direction of the filter w, Z D ∈R H×W×C Represents the result of directional convolution; directional convolution can be defined as: where X*w represents the convolution operation; D is the direction vector of the directional convolution, with values of (0,1), (1,0), (1,1) and (-1,1), which are used for horizontal, vertical, left diagonal and right diagonal convolutions respectively; for filter w, set k=4 so that each directional convolution has 9 parameters, the same as the 3×3 convolution kernel; Step 2.4, build a connected attention module CAM: for edge labels, generate a connected cube , where C o Represents the number of samples for a given pixel and its neighboring pixels; for a two-dimensional edge image, C o Recorded as 8, O in the connectivity graph i,h,c Represents the connectivity between a pixel and a specific pixel, where i, j represent the spatial position of the pixel; c represents the position of its adjacent pixel. If two pixels are connected, then O i,j,c =1, indicating that the two pixels are connected farmland boundaries; finally, the connectivity information of the eight neighborhoods is calculated and spliced to obtain the final connectivity map; The input tensor is sequentially fed into a 3×3 convolution with an expansion rate of 1, and then the Efficient Channel Attention ECA block is used to make full use of connectivity by recalibrating the predicted connectivity cube using the channel attention mechanism; the input feature map is fed into this block after global average pooling to obtain a vector in the range (0,1), where each factor is multiplied by the corresponding channel in the input feature map; the final output of CAM is a H×W×C o Finally, the ground truth of the connectivity graph generated by the labels is used for supervised learning.
4. The method for extracting key points of land polygons based on edges and edge directions according to claim 1, characterized in that: Step 3 specifically includes: Step 3.1, the training set TR and the validation set VA are input into the edge and edge direction based multi-task extraction network as training data and validation data respectively; Step 3.2, set the model's hyperparameters: Epochs, BatchSize, LearningRate, and Optimizer; Step 3.3, calculate the loss of the edge result and edge direction result with the true label category respectively, and the loss function of this model is the joint loss; L=L seg +αL con +βL class (2) where α and β are two constants; N represents the number of elements in an H×W slice, and y i is the ground truth value representing the edge of the drawing for a given pixel in position i, is the predicted probability of the corresponding segmentation branch; C o represents the number of sampled neighboring pixels for a given pixel, is a true value indicating whether a given pixel at position i is connected to a neighboring pixel at position c. is the predicted connectivity corresponding to the connectivity branch; C represents the number of categories in the edge direction, z i,c is the true label of the sample in category c, is the predicted probability of the sample in category c; Step 3.4: The model is iteratively trained and the model with the highest F1 value is selected as the final prediction model M. The calculation method of F1 is shown in the formula: Precision is the accuracy rate, which indicates the proportion of samples predicted to be positive that are actually positive. The calculation method is shown in the formula: Among them, TP is the number of pixels of the positive sample whose prediction is correct, FP is the number of pixels of the positive sample whose prediction is wrong, and FP is the number of pixels of the negative sample whose prediction result is wrong; Recall is the recall rate, which means the ratio of the actual number of positive samples in the predicted positive samples to the total number of positive samples in the samples. The calculation method is shown in the formula: Where FN is the number of pixels whose prediction results are wrong negative samples.
5. The method for extracting key points of land polygons based on edges and edge directions according to claim 1, characterized in that: The specific implementation of step 5 includes the following sub-steps: Step 5.1, the output plot edge is subjected to non-maximum suppression, binarization and skeleton extraction to obtain a refined plot edge result; Step 5.2, using the refined plot edge result to mask the plot edge direction output in step 4 to obtain a refined edge direction; Step 5.3, set the pixel point with only one pixel value of 1 in the eight-connected domain as the breakpoint, and find the breakpoint on the refined land edge result; Step 5.5, traverse the breakpoints, obtain the edge direction of the breakpoint position through the refined edge direction, extend the edge according to the edge direction, until the extension distance exceeds the threshold th or encounters another edge, thereby obtaining an improved closed edge; Step 5.6, find the direction mutation point in the refined edge direction and define it as the key point; Step 5.7, consider the closed edge as the connectivity information of the key points, connect the key points, and obtain the land polygon; Thus, the prediction result R of the land polygon of the image can be obtained.
6. A device for extracting key points of land polygons based on edges and edge directions, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the method for extracting key points of land polygons based on edges and edge directions as described in any one of claims 1-5.
7. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the method for extracting key points of land polygons based on edges and edge directions described in any one of claims 1 to 5 is implemented.