Regularized land parcel extraction method from high-resolution remote sensing images with multi-task line segment constraints
Through the multi-task line segment constraint, the regular plot extraction method is used to process high-resolution remote sensing images, and the multi-task convolutional neural network model is used to process high-resolution remote sensing images to generate regular plot polygons, which solves the problem of time-consuming and labor-intensive and irregular extraction results in traditional methods, and improves the accuracy and efficiency of plot extraction.
Patent Information
- Application Number
- CN202311195263.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-16
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-09-16
AI Technical Summary
Traditional land extraction methods rely on field investigations or manual processing of high-resolution remote sensing images, which is time-consuming and labor-intensive, difficult to meet the needs of large-scale agricultural applications, and it is difficult to automatically extract high-precision land.
A high-resolution remote sensing image regularized plot extraction method is adopted with multi-task line segment constraints. By constructing a multi-task convolutional neural network model, image feature extraction and line segment constraint processing are performed to generate regular plot polygons.
The accuracy and utilization efficiency of arable land plot extraction are improved, the problems of fuzzy edges and irregular plot extraction results are overcome, and more efficient acquisition of agricultural land plot information is achieved.
Smart Images

Figure CN117173574B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of farmland plot extraction research, and in particular to a method for regularized plot extraction from high-resolution remote sensing images with multi-task line segment constraints. Background Art
[0002] Arable land is the basis of crop planting, and accurate arable land plot information is the basis of refined agricultural research. Traditional plot extraction methods mostly rely on field surveys or artificial digitized high spatial resolution (abbreviated as high resolution) remote sensing images, which are time-consuming and labor-intensive, with a long production cycle, and are difficult to meet the needs of large-scale agricultural applications. As an effective means of obtaining geographic spatial information, remote sensing has the characteristics of wide observation range, strong timeliness, and short periodicity, and is widely used in the extraction and classification of various types of thematic information. Among them, high-resolution remote sensing images can clearly express the spatial structure and surface texture characteristics of objects, highlight the edge information of objects, and provide the basis and conditions for refined plot extraction. However, the current conventional plot extraction methods based on high-resolution remote sensing images mostly rely on the spectral, texture, shape and other features of the image, and do not make full use of the semantic features with a high level of abstraction, making it difficult to automatically extract high-precision plots. Arable land plots are small and fragmented, with complex structures and scattered distribution. In addition, the crop planting structure is diverse, so it is very difficult to extract accurate arable land plot information.
[0003] Deep learning is an advanced pattern recognition technology in the field of artificial intelligence. It has attracted extensive attention from researchers because of its significant effect in image feature learning. Compared with traditional image segmentation methods, deep learning methods do not need to rely on expert knowledge for manual feature extraction, but autonomously learn relevant contextual semantics and other high-level features through convolution filters. These features are closely linked to classifiers, solving the problem of manual selection of classifiers and classification features in traditional machine learning, and realizing end-to-end image processing and analysis. Deep learning technology has developed rapidly, especially deep convolutional neural networks have been widely used in remote sensing image target recognition, image classification and semantic segmentation. In recent years, DCNN (Deep Convolutional Neural Network) has gradually been applied to the study of cultivated land extraction in farmland. However, the output results of the model usually pay more attention to the overlap between the predicted mask and the real land, and rarely care about its geometric information. The method based on semantic annotation cannot solve the problem of blurred edges and irregular land extraction results. The land extraction results generated by the semantic segmentation method are usually pixel grids, which are difficult to use later. Even with complex post-processing methods, these results are difficult to produce ideal regular polygon vectors. Farmland plots vary in shape and size, and planting leads to greater diversity within the plots. Plots are closely grouped, have different features, and usually do not have regular geometric features. Due to factors such as tree occlusion and terrain, the boundary information of the plots is more ambiguous. Previous studies usually have difficulty solving the problem of unclosed or incomplete boundaries. Therefore, extracting geometric polygons from irregular plots is a challenging task. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a regularized plot extraction method for high-resolution remote sensing images with multi-task line segment constraints, which overcomes the problems of difficulty in utilizing extraction results, blurred edges and irregular plot extraction results existing in traditional convolutional neural networks in cultivated land plot identification, and effectively improves the accuracy and utilization efficiency of cultivated land plot extraction.
[0005] To achieve the above purpose, the present invention adopts the following technical scheme: a regularized land parcel extraction method for high-resolution remote sensing images with multi-task line segment constraints, a multi-task neural network construction and land parcel shape regularization, comprising the following steps:
[0006] Step S1: Select high-resolution remote sensing images covering the study area and preprocess the images;
[0007] Step S2: construct a multi-task convolutional neural network model that integrates the spatial information of bottom-level vertices, middle-level line segments, and high-level masks to extract image features at different levels;
[0008] Step S3: Based on step S2, the line segment attraction direction map representation method is used to constrain the middle-level line segment features to improve the problems of irregular geometric shapes, discontinuous boundaries, and non-closure in the land parcel extraction; specifically, this is achieved by calculating the projection distance from each pixel on the map to the nearest line segment;
[0009] Step S4: Establish a joint loss function, select the optimizer of the model and perform parameter fine-tuning;
[0010] Step S5: training the multi-task network model of step S2, and using the trained model to perform regularized land extraction on the remote sensing image obtained in step S1;
[0011] Step S6: Regularization of model output results; specifically including polygon generation and vertex optimization.
[0012] In a preferred embodiment, step S1 specifically includes the following steps:
[0013] Step S11: Obtain high-resolution remote sensing images of the study area and perform preprocessing operations on the images, including radiation correction, geometric correction, image cropping, and multispectral-panchromatic fusion operations.
[0014] In a preferred embodiment, step S2 specifically includes the following steps:
[0015] Step S21: construct a convolutional neural network encoder, wherein the encoder consists of four parts, each of which includes a convolution layer, a pooling layer, and an activation function;
[0016] Step S22: construct a multi-task convolutional neural network decoder, wherein the decoder consists of two parts, the first half includes four upsampling, feature channel concatenation operations and a spatial self-attention module, and the second half includes three branches, one of which is used to output the middle-layer line segment enhancement mask, and the other two branches are used to output the vertex heat map and the high-level mask. Each branch includes a convolution layer, a pooling layer, an activation function, and a task output module. The line segment and vertex, and the mask branches all include a spatial information fusion module, and the middle-layer line segment branch also has a channel enhancement module.
[0017] Step S23: Use the multi-task convolutional neural network model for different levels of feature detection constructed in step S2 to automatically learn different tasks Task1, Task2, ....Task n The spatial information between the two tasks is extracted and features are shared to extract multi-task features.
[0018] In a preferred embodiment, step S3 specifically includes the following steps:
[0019] Step S31: Calculate the projection distance from each pixel to all line segments on the graph, and assign it to the line segment with the shortest projection distance; Calculate the projection distance D from the pixel to the line segment i Use the following formula:
[0020]
[0021] In the formula, w i and h j Represents the horizontal and vertical coordinates of the pixel point, a i , b i 、c i is the normal equation of the projection point of the pixel on the line segment (a i x+b i y+c i =0) three parameters;
[0022] Each pixel x in the graph can be expressed as:
[0023]
[0024] That is, the pixel is assigned to the line segment i whose projection distance to the line segment is the shortest;
[0025] The line segment attraction direction graph can be expressed as
[0026] S = x' i -x i , x i ∈δ i
[0027] Where x′ i belongs to line segment l i Pixel x i On line segment l i Projection on.
[0028] In a preferred embodiment, step S4 specifically includes the following steps:
[0029] Step S41: Based on the multi-task network model constructed in step S2, a multi-task loss function is designed to update the network weights to reduce the error between the actual and predicted results; specifically, a negative log-likelihood function is used as the land parcel prediction mask loss function, and the calculation formula is:
[0030] l Mask =w1l Mseg +w2l Mafm
[0031] In the formula, l Mask It represents the total loss value of the mask prediction part, w1 and w2 are the weight parameters of the loss of each part; the specific single loss calculation formula is:
[0032]
[0033]
[0034] In the formula, l Mseg , l Mafm They represent the classification errors calculated by comparing the masks generated by fusion of different levels of features and the masks generated by fusion of line segment gravity maps with the label pixels; n represents the number of pixels, p m. (x; f m. ) represents the predicted probability of the model after the Softmax function, f m. is the true label;
[0035] Step S42: The line segment attraction graph uses the L1 loss function:
[0036]
[0037] Where l Afm represents the error calculated using L1 loss, where x, f afm Represent the predicted value and the true label respectively;
[0038] Step S43: L1 loss function and cross entropy loss function are used as the point prediction offset and vertex heat map estimation task loss functions respectively:
[0039] l Ver =λ1l Mseg +λ2l Mafm
[0040] In the formula, l Ver Represents the total loss value of the vertex prediction part, λ1 and λ2 are the weight parameters of the loss of each part; the specific single loss calculation formula is:
[0041]
[0042]
[0043] Where f ver is the true label of the sample, the positive class is 1 and the negative class is 0; p is the probability that the sample is predicted to be positive; f voff represents the short-distance offset of positive pixels in the range of [-0.5, 0.5), and e is the predicted value of the offset;
[0044] Step S44: Setting the loss function l of the multi-task neural network Total for
[0045] l Total =θ1l Mask +θ2l Afm +θ3l Ver
[0046] In the formula, θ1, θ2, θ3 represent the weight ratio of each part, l Total It is the loss function of the feature extraction part;
[0047] Step S45: Adopt the Adam optimizer and set the initial learning rate Lr as the optimizer for model prediction, set the training batch of the model to B, and set the number of model iterations to E;
[0048] Step S45: Use the weights pre-trained on the large image processing dataset ImageNet to initialize the network model to speed up the model convergence, improve the model training efficiency, and obtain a better model weight file.
[0049] In a preferred embodiment, step S5 specifically includes the following steps:
[0050] Step S51: using the remote sensing image obtained in step S1, making a sample binary mask and performing format conversion, training a multi-task neural network model with the converted labels, and saving the model weights;
[0051] Step S52: Use the model to predict the target image obtained in step S1 to obtain the vertex heat map and mask detection results of the cultivated land.
[0052] In a preferred embodiment, step S6 specifically includes the following steps:
[0053] Step S61: Use the vertex heat map of cultivated land and the mask detection results obtained in S5; first, process the mask into n instances, each instance is represented by the coordinates converted from the mask boundary pixels
[0054] S i ={x i 1,x i 2,...,x i m}, where i∈{1,2,3...n};
[0055] Step S62: Using the vertex heat map result of the cultivated land block obtained in step S5, perform the local non-maximum suppression operation NMS and retain the first n points with the highest prediction scores to obtain the vertex sparse matrix, and further refine it with the position offset vector. The obtained sparse vertex matrix is expressed as V = {v1,v2,...,v h};
[0056] Step S63: Using the sparse vertex matrix obtained in step S61, for the polygon boundary coordinates S generated in step S61 i ={x i1 ,x i2 ,...,xim}, calculate the Euclidean distance between the boundary coordinates of each instance and the sparse vertex matrix, and assign a set of indices containing the shortest vertex and distance information to each boundary coordinate. The process can be expressed as:
[0057]
[0058] Where represents the jth polygon, k represents the kth boundary coordinate point of the polygon, and i represents the ith vertex of the sparse matrix; the boundary points with non-minimum distance and minimum distance greater than a certain pixel with the same index information are deleted to obtain preliminary vertices for constructing polygons;
[0059] Step S64: Using the preliminary polygon vertex results obtained in step S63, the Douglas-Peucker algorithm is used to simplify the rules. First, a straight line is connected between the starting point and the end point of the curve or trajectory to divide the curve or trajectory into two parts; then, the distance from each intermediate point to the straight line is calculated, and the point with the largest distance is selected as the key point; if the maximum distance is less than the set threshold, the starting point and the end point are retained to form a line segment; if the maximum distance is greater than the threshold, the key point is retained and the curve or trajectory is divided into two parts; for the two divided parts, the above algorithm is continued until all key points are determined; finally, all key points are connected to form a simplified regular polygon representation.
[0060] Compared with the prior art, the present invention has the following beneficial effects: the present invention not only overcomes the problems of blurred edges and irregular plot extraction results in traditional convolutional neural networks in the identification of cultivated land plots, but also overcomes the problems of difficulty in using the extraction results of traditional prediction methods and complex post-processing. The method of the present invention further improves the geometric accuracy of cultivated land plot extraction by combining vertex detection and multi-task learning, providing important data support and technical support for production management and macro decision-making of agricultural management departments. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 A method flow chart of a preferred embodiment of the present invention;
[0062] Figure 2 This is a structural diagram of a spatial self-attention CA module in a preferred embodiment of the present invention;
[0063] Figure 3 This is a structural diagram of a spatial information fusion SCSE module according to a preferred embodiment of the present invention;
[0064] Figure 4 Output LAM module structure diagram for tasks in the preferred embodiment of the present invention;
[0065] Figure 5A structural diagram of a multi-task network model according to a preferred embodiment of the present invention;
[0066] Figure 6 A flowchart of extracting regular polygonal plots according to a preferred embodiment of the present invention;
[0067] Figure 7 This is a partial extraction result diagram of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0068] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0069] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0070] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.
[0071] like Figure 1 As shown, this embodiment provides a method for extracting cultivated land plots from high-resolution remote sensing images combining edge detection and multi-task learning, including the following steps:
[0072] Step S1: Select high-resolution remote sensing images covering the study area and preprocess the images
[0073] Step S2: Construct a multi-task convolutional neural network model that integrates the multi-level spatial information of bottom-level vertices, middle-level line segments, and high-level masks to extract image features at different levels.
[0074] Step S3: Based on step S2, the line segment attraction direction map representation method is used to constrain the middle-level line segment features to improve the problems of irregular geometric shapes, discontinuous boundaries, and incomplete closure in land parcel extraction. Specifically, this is achieved by calculating the projection distance from each pixel on the map to the nearest line segment.
[0075] Step S4: Establish a joint loss function, select the optimizer of the model and perform parameter fine-tuning.
[0076] Step S5: Train the multi-task network model of step S2, and use the trained model to perform regularized plot extraction on the remote sensing image obtained in step S1.
[0077] Step S6: Regularization of model output results, including polygon generation and vertex optimization.
[0078] In this embodiment, step S1 specifically includes the following steps:
[0079] Step S11: Obtain high-resolution remote sensing images of the study area and perform preprocessing operations on the images, including radiation correction, geometric correction, image cropping, and multispectral-panchromatic fusion operations.
[0080] In this embodiment, step S2 specifically includes the following steps:
[0081] Step S21: construct a convolutional neural network encoder, wherein the encoder consists of four parts, each of which includes a convolution layer, a pooling layer and an activation function;
[0082] Step S22: construct a multi-task convolutional neural network decoder, wherein the decoder consists of two parts, the first half includes four upsampling, feature channel splicing operations and a spatial self-attention module, and the second half includes three branches, one of which is used to output the middle-layer line segment enhancement mask, and the other two branches are used to output the vertex heat map and the high-level mask. Each branch includes a convolution layer, a pooling layer, an activation function, and a task output module. The line segment and vertex and mask branches all contain a spatial information fusion module, and the middle-layer line segment branch also has a channel enhancement module.
[0083] Step S23: Use the multi-task convolutional neural network model of different levels of feature detection constructed in step S21 to automatically learn different tasks Task1, Task2, ....Task n The spatial information between the two tasks is extracted and features are shared to extract multi-task features.
[0084] In this embodiment, step S3 specifically includes the following steps:
[0085] Step S31: Calculate the projection distance from each pixel to all line segments, and assign it to the line segment with the shortest projection distance. Calculate the projection distance D from the pixel to the line segment i Use the following formula:
[0086]
[0087] In the formula, w i and h j Represents the horizontal and vertical coordinates of the pixel point, a i 、b i 、c i is the normal equation of the projection point of the pixel on the line segment (a i x+b i y+c i =0).
[0088] Each pixel x in the graph can be expressed as:
[0089]
[0090] That is, the pixel is assigned to the line segment i whose projection distance to the line segment is the shortest.
[0091] The line segment attraction direction graph can be expressed as
[0092] S = x' i -x i , x i ∈δ i
[0093] Where x′ i belongs to line segment l i Pixel x i On line segment l i Projection on.
[0094] In this embodiment, step S4 specifically includes the following steps:
[0095] Step S41: Based on the multi-task network model constructed in step S2, a multi-task loss function is designed to update the network weights to reduce the error between the actual and predicted results. Specifically, a negative log-likelihood function is used as the landmass prediction mask loss function, and the calculation formula is:
[0096] l Mask =w1l Mseg +w2l Mafm
[0097] Where, L Mask It represents the total loss value of the mask prediction part, w1 and w2 are the weight parameters of the loss of each part; the specific single loss calculation formula is:
[0098]
[0099]
[0100] In the formula, l Mseg , l Mafm They represent the classification errors calculated by comparing the masks generated by fusion of different levels of features and the masks generated by fusion of line segment gravity maps with the label pixels; n represents the number of pixels, p m· (x; f m. ) represents the predicted probability of the model after the Softmax function, f m· is the true label.
[0101] Step S42: The line segment attraction graph uses the L1 loss function:
[0102]
[0103] Where l Afm represents the error calculated using L1 loss, where x, f afm Represent the predicted value and the true label respectively.
[0104] Step S43: L1 loss function and cross entropy loss function are used as the point prediction offset and vertex heat map estimation task loss functions respectively:
[0105] l Ver =λ1l Mseg +λ2l Mafm
[0106] In the formula, l Ver Represents the total loss value of the vertex prediction part, λ1 and λ2 are the weight parameters of the loss of each part; the specific single loss calculation formula is:
[0107]
[0108]
[0109] Where f ver is the true label of the sample, the positive class is 1, and the negative class is 0; p is the probability that the sample is predicted to be a positive class. voff It represents the short-distance offset of positive pixels in the range of [-0.5, 0.5), and e is the predicted value of the offset.
[0110] Step S44: Setting the loss function l of the multi-task neural network Total for
[0111] l Total =θ1l Mask +θ2l Afm +θ3l Ver
[0112] In the formula, θ1, θ2, θ3 represent the weight ratio of each part, l Total It is the loss function of the feature extraction part.
[0113] Step S45: Adopt the Adam optimizer and set the initial learning rate Lr as the optimizer for model prediction, set the training batch of the model to B, and set the number of model iterations to E;
[0114] Step S45: Use the weights pre-trained on the large image processing dataset ImageNet to initialize the network model to speed up the model convergence, improve the model training efficiency, and obtain a better model weight file.
[0115] In this embodiment, step S5 specifically includes the following steps:
[0116] Step S51: using the remote sensing image obtained in step S1, making a sample binary mask and performing format conversion, training a multi-task neural network model with the converted labels, and saving the model weights;
[0117] Step S52: Use the model to predict the target image obtained in step S1 to obtain the vertex heat map and mask detection results of the cultivated land.
[0118] In this embodiment, step S6 specifically includes the following steps:
[0119] Step S61: Use the vertex heat map of cultivated land and the mask detection results obtained in S5. First, process the mask into n instances, each instance is represented by the coordinates converted from the mask boundary pixels. i ={x i1 ,x i2 ,...,x im}, where i∈{1,2,3...n}.
[0120] Step S62: Using the vertex heat map result of the cultivated land block obtained in step S5, perform a local non-maximum suppression (NMS) operation and retain the top n points with the highest prediction scores (pixel values in the image) to obtain a vertex sparse matrix, and further refine it using the position offset vector. The obtained sparse vertex matrix is represented as V = {v1, v2, ..., v h}.
[0121] Step S63: Using the sparse vertex matrix obtained in step S61, for the polygon boundary coordinates S generated in step S61 i ={x i1 ,x i2 ,...,x im}, calculate the Euclidean distance between the boundary coordinates of each instance and the sparse vertex matrix, and assign a set of indices containing the shortest vertex and distance information to each boundary coordinate. The process can be expressed as:
[0122]
[0123] Where represents the jth polygon, k represents the kth boundary coordinate point of the polygon, and i represents the ith vertex of the sparse matrix. The boundary points with non-minimum distance and minimum distance greater than a certain pixel with the same index information are deleted to obtain preliminary vertices for constructing polygons.
[0124] Step S64: Using the preliminary polygon vertex results obtained in step S63, the Douglas-Peucker algorithm is used to simplify the rules. First, a straight line is connected between the starting point and the end point of the curve or track to divide the curve or track into two parts. Then, the distance from each intermediate point to the straight line is calculated, and the point with the largest distance is selected as the key point. If the maximum distance is less than the set threshold, the starting point and the end point are retained to form a line segment. If the maximum distance is greater than the threshold, the key point is retained and the curve or track is divided into two parts. For the two divided parts, continue to use the above algorithm until all key points are determined. Finally, all key points are connected to form a simplified regular polygon representation.
[0125] In this example, the GF-2 remote sensing image of Korla, Xinjiang in July 2021 was used. After preprocessing in step 1, the spatial resolution was 1m, and the bands were red, green, and blue. This example used 4576 images with a pixel size of 256×256 and corresponding plot labels for model training.
[0126] like Figure 2 As shown, this is the structure diagram of the spatial self-attention CA module of this embodiment. As can be seen from the figure, the module first performs an average pooling operation on the feature matrix on two different direction axes, and then splices the pooled feature tensors. After two-dimensional convolution, normalization and nonlinear convolution layers, two two-dimensional convolutions are input respectively and activated with Sigmoid function. The activated feature matrix is used as the weight to multiply the initial input feature map, and finally a feature tensor with contextual interaction information is obtained to improve the feature expression of key areas. This solves the problem of fuzzy and discontinuous boundary prediction of target objects.
[0127] like Figure 3 As shown in FIG. 1 , the structure diagram of the spatial information fusion SCSE module of this embodiment is shown. Taking the fusion of middle-level line segments and high-level mask features as an example, the module first transforms the middle-level line segment feature vector F afm With Advanced Mask F seg The feature information of the sum is added, and then the added feature tensor is divided into two parts for operation. One part is subjected to two-dimensional convolution and Sigmoid activation function, and the other part is subjected to two two-dimensional convolutions after the global average pooling layer. Then the feature tensors after these two parts of operation are respectively compared with F seg The feature tensor is dot-multiplied with F seg The tensors are added to obtain enhanced spatial key information and weaken the feature expression of irrelevant background areas. This module is used to improve the prediction accuracy of polygon key points.
[0128] like Figure 4As shown, this is the structure diagram of the task output LAM module of this embodiment. As can be seen from the figure, the module first uses 3 consecutive convolution, normalization and Relu activation operations to expand the number of channels of the segment gravity map FAM, and then passes the feature tensor after the expansion of the number of channels through a global average pooling layer, a nonlinear convolution layer, a Relu activation function module and a nonlinear convolution layer, a Sigmoid activation function module, and then performs a dot multiplication operation on the expanded feature tensor, and finally adds it to the backbone network feature as the feature tensor of mask prediction. Through the above operations, the feature tensor has correct geometric information to solve the problem of geometric correctness of the output mask and enhance the regularization of the edge of the output mask.
[0129] like Figure 5 As shown, this is the polygon vertex detection model BsiNet of this embodiment. hisup As can be seen from the figure, BsiNet hisup The network model first undergoes a 3×3 convolution to expand the number of channels to 32, and then uses 4 3×3 convolutions, activation functions, and downsampling of the maximum pooling layer to obtain high-level semantic information. In order to reduce spatial information loss, the Concat operation is used to combine the semantic information of different levels of the encoding with the upsampling information of the four corresponding decoders. Each upsampling reduces the current number of channels by half and doubles the height and width. Finally, a 3×3 convolution operation with a step size of 1 is used to obtain a set of 32×256×256 feature tensors. Use Figure 2 The spatial self-attention module is used to enhance the spatial feature information of the model and then inputs into the three branch decoders. After each branch is convolved, normalized and activated, the high-level mask and the bottom-level vertex are respectively Figure 3 The spatial information fusion module is fused with the mid-level line segment features, and finally the output features are converted to 2×256×256 to obtain masks and vertices. The mid-level line segment gravity map is used as the intermediate output feature of the model and supervised by the loss function. Figure 4 In the task output module, it is fused with the previous backbone features to output a predicted mask with geometric correctness.
[0130] like Figure 6As shown, the regular polygonal plot extraction of this embodiment. As can be seen from the figure, the mask, vertex heat map and vertex position offset are generated by the model. First, the mask is processed into n instances, each instance is represented by the coordinates converted from the boundary pixels of the mask. The vertex heat map performs a local 3×3 non-maximum suppression operation (NMS) and retains the top 300 points with the highest prediction scores (pixel values in the figure) to obtain a vertex sparse matrix, which is further refined using the position offset vector, and the polygon vertices are selected by mapping the sparse vertices to the boundary pixels. Finally, the Douglas-Peucker algorithm is used to further simplify the selected vertices, and finally a regularized polygon of the target plot is generated.
[0131] like Figure 7 As shown in FIG. 1 , some experimental results of this embodiment on the GF-2 remote sensing image of Korla, Xinjiang are shown. It can be seen from the figure that the land polygons obtained using the proposed method are highly consistent with the ground truth labels; secondly, the extracted land boundaries are regular and square without any unclear phenomenon; the experimental results show the effectiveness of the proposed method in extracting cultivated land plots.
[0132] The present invention not only overcomes the problems of blurred edges and irregular plot extraction results in traditional convolutional neural networks in farmland plot identification, but also overcomes the problems of difficulty in using the extraction results of traditional prediction methods and complex post-processing. The method of the present invention further improves the geometric accuracy of farmland plot extraction by combining vertex detection and multi-task learning. The present invention considers the comprehensive utilization of vertex detection and mask recognition for farmland plot extraction.
[0133] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.
Claims
1. A regularized land parcel extraction method for high-resolution remote sensing images with multi-task line segment constraints, characterized in that: The multi-task neural network construction and land shape regularization include the following steps: Step S1: Select high-resolution remote sensing images covering the study area and preprocess the images; Step S2: construct a multi-task convolutional neural network model that integrates the spatial information of bottom-level vertices, middle-level line segments, and high-level masks to extract image features at different levels; Step S3: Based on step S2, the line segment attraction direction map representation method is used to constrain the middle-level line segment features to improve the problems of irregular geometric shapes, discontinuous boundaries, and non-closure in the land parcel extraction; specifically, this is achieved by calculating the projection distance from each pixel on the map to the nearest line segment; Step S4: Establish a joint loss function, select the optimizer of the model and perform parameter fine-tuning; Step S5: training the multi-task network model of step S2, and using the trained model to extract regularized plots from the remote sensing image obtained in step S1; Step S6: Regularization of model output results; specifically including polygon generation and vertex optimization; Step S3 specifically includes the following steps: Step S31: Calculate the projection distance from each pixel to all line segments on the graph, and assign it to the line segment with the shortest projection distance; Calculate the projection distance D from the pixel to the line segment i Use the following formula: In the formula, w i and h j Represents the horizontal and vertical coordinates of the pixel point, a i 、b i 、c i is the normal equation of the projection point of the pixel on the line segment (a i x+b i y+c i =0) three parameters; Each pixel x in the graph is represented by: That is, the pixel is assigned to the line segment i whose projection distance to the line segment is the shortest; The line segment attraction direction diagram is expressed as S=x′ i -x i ,x i ∈δ i Where x′ i belongs to line segment l i Pixel x i On line segment l i projection on; Step S4 specifically includes the following steps: Step S41: Based on the multi-task network model constructed in step S2, a multi-task loss function is designed to update the network weights to reduce the error between the actual and predicted results; specifically, a negative log-likelihood function is used as the land parcel prediction mask loss function, and the calculation formula is: l Mask =w1l Mseg +w2l Mafm In the formula, l Mask It represents the total loss value of the mask prediction part, w1 and w2 are the weight parameters of the loss of each part; the specific single loss calculation formula is: In the formula, l Mseg , l Mafm They represent the classification errors calculated by comparing the masks generated by fusion of different levels of features and the masks generated by fusion of line segment gravity maps with the label pixels; n represents the number of pixels, p m· (x; f m· ) represents the predicted probability of the model after the Softmax function, f m· is the true label; Step S42: The line segment attraction graph uses the L1 loss function: Where l Afm Represents the error calculated using L1 loss, where x,f afm Represent the predicted value and the true label respectively; Step S43: L1 loss function and cross entropy loss function are used as the loss functions of point prediction offset and vertex heat map estimation tasks respectively: l Ver =λ1l Mseg +λ2l Mafm In the formula, l Ver Represents the total loss value of the vertex prediction part, λ1 and λ2 are the weight parameters of the loss of each part; the specific single loss calculation formula is: Where f ver is the true label of the sample, the positive class is 1 and the negative class is 0; p is the probability that the sample is predicted to be positive; f voff It represents the short-distance offset of positive pixels in the range of [-0.5, 0.5), and e is the predicted value of the offset; Step S44: Setting the loss function l of the multi-task neural network Total for l Total =θ1l Mask +θ2l Afm +θ3l Ver In the formula, θ1, θ2, θ3 represent the weight ratio of each part, l Total It is the loss function of the feature extraction part; Step S45: Adopt the Adam optimizer and set the initial learning rate Lr as the optimizer for model prediction, set the training batch of the model to B, and set the number of model iterations to E; Step S45: Use the weights pre-trained on the large image processing dataset ImageNet to initialize the network model to speed up the model convergence, improve the model training efficiency, and obtain the model weight file.
2. The method for extracting regularized land parcels from high-resolution remote sensing images with multi-task line segment constraints according to claim 1, characterized in that: Step S1 specifically includes the following steps: Step S11: Obtain high-resolution remote sensing images of the study area and perform preprocessing operations on the images, including radiation correction, geometric correction, image cropping, and multispectral-panchromatic fusion operations.
3. The method for extracting regularized land parcels from high-resolution remote sensing images with multi-task line segment constraints according to claim 1, characterized in that: Step S2 specifically includes the following steps: Step S21: construct a convolutional neural network encoder, wherein the encoder consists of four parts, each of which includes a convolution layer, a pooling layer, and an activation function; Step S22: construct a multi-task convolutional neural network decoder, wherein the decoder consists of two parts, the first half includes four upsampling, feature channel concatenation operations and a spatial self-attention module, and the second half includes three branches, one of which is used to output the middle-layer line segment enhancement mask, and the other two branches are used to output the vertex heat map and the high-level mask. Each branch includes a convolution layer, a pooling layer, an activation function, and a task output module. The line segment and vertex, and the mask branches all include a spatial information fusion module, and the middle-layer line segment branch also has a channel enhancement module. Step S23: Use the multi-task convolutional neural network model of different levels of feature detection constructed in step S2 to automatically learn different tasks Task1, Task2, ....Task n The spatial information between the two tasks is extracted and features are shared to extract multi-task features.
4. The method for extracting regularized land parcels from high-resolution remote sensing images with multi-task line segment constraints according to claim 1, characterized in that: Step S5 specifically includes the following steps: Step S51: using the remote sensing image obtained in step S1, making a sample binary mask and performing format conversion, training a multi-task neural network model with the converted labels, and saving the model weights; Step S52: Use the model to predict the target image obtained in step S1 to obtain the vertex heat map and mask detection results of the cultivated land.
5. The multi-task line segment constrained high-resolution remote sensing image regularized plot extraction method according to claim 1, characterized in that: Step S6 specifically includes the following steps: Step S61: Use the vertex heat map of cultivated land and the mask detection results obtained in S5; firstly process the mask into n instances, each instance is represented by the coordinates converted from the mask boundary pixels S i ={x i 1,x i 2,...,x i m }, where i∈{1,2,3...n}; Step S62: Using the vertex heat map result of the cultivated land block obtained in step S5, perform the local non-maximum suppression operation NMS and retain the first n points with the highest prediction scores to obtain the vertex sparse matrix, and further refine it with the position offset vector. The obtained sparse vertex matrix is expressed as V = {v1,v2,...,v h }; Step S63: Using the sparse vertex matrix obtained in step S61, for the polygon boundary coordinates S generated in step S61 i ={x i1 ,x i2 ,...,x im }, calculate the Euclidean distance between the boundary coordinates of each instance and the sparse vertex matrix, and assign a set of indices containing the shortest vertex and distance information to each boundary coordinate. The process is expressed as: Where represents the jth polygon, k represents the kth boundary coordinate point of the polygon, and i represents the i-th vertex of the sparse matrix; the boundary points with non-minimum distances and minimum distances greater than a certain number of pixels having the same index information are deleted to obtain preliminary vertices for constructing polygons; Step S64: Using the preliminary polygon vertex results obtained in step S63, the Douglas-Peucker algorithm is used to simplify the rules. First, a straight line is connected between the starting point and the end point of the curve or trajectory to divide the curve or trajectory into two parts; then, the distance from each intermediate point to the straight line is calculated, and the point with the largest distance is selected as the key point; if the maximum distance is less than the set threshold, the starting point and the end point are retained to form a line segment; if the maximum distance is greater than the threshold, the key point is retained and the curve or trajectory is divided into two parts; for the two divided parts, the above algorithm is continued until all key points are determined; finally, all key points are connected to form a simplified regular polygon representation.
Citation Information
Patent Citations
Remote sensing image cultivated land parcel extraction method combining edge detection and multi-task learning
CN114821315A
Self-attention feature fused high-resolution remote sensing image semantic change detection method
CN116486255A