High-resolution remote sensing image building and road vector extraction method

CN120107609APending Publication Date: 2025-06-06CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510265417.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

Smart Images

  • Figure CN120107609A_ABST
    Figure CN120107609A_ABST
Patent Text Reader

Abstract

The invention relates to the field of remote sensing image processing, and discloses a high-resolution remote sensing image building and road vector extraction method, which comprises the following steps: acquiring a high-resolution remote sensing image data set, and performing data enhancement on the remote sensing image data set to obtain a high-resolution remote sensing image data set; obtaining a remote sensing image, a road segmentation label, road vector connection information, a building label and building vector connection information; constructing a road and building collaborative extraction model based on the encoder parameters, wherein the road and building collaborative extraction model comprises a multi-scale feature extraction module, a feature interaction module, a double-branch feature decoding module and a vector topology prediction module; inputting a batch of data into the road and building collaborative extraction model, and training to obtain a trained model; obtaining a building and road vector result according to the trained model; according to the invention, the precision of building and road collaborative extraction is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of remote sensing image processing, and in particular to a method for extracting building and road vectors from high-resolution remote sensing images. Background Art

[0002] Road and building vectors are of great significance in many practical applications such as urban planning, map construction, and autonomous driving.

[0003] Building and road vector extraction mainly includes two categories: traditional extraction algorithms and deep learning-based algorithms. Among them, traditional algorithms are highly dependent on artificially designed features, such as color, shape, and texture, and are limited by complex actual scenes and have poor generalization capabilities. With the rapid development of deep learning technology in the field of computer vision, deep learning methods have gradually been applied to remote sensing building and road vector extraction tasks, and have significant advantages in automated road and building vector extraction.

[0004] Existing deep learning-based building and road vector extraction methods mainly include segmentation-based methods and vertex-based two-stage methods. The segmentation-based method predicts the semantic segmentation masks of roads and buildings, and obtains vector results through complex post-processing. The vectors obtained by this post-processing algorithm often have problems with topological errors and geometric distortions; the vertex-based two-stage method first predicts the road and building masks through a deep learning model, samples all vertex positions, and then predicts the connection relationship between vertices through the model to determine the topological structure of the vector. This type of method can effectively solve the problems of topological errors and geometric distortions. However, there are still some problems with the existing vertex-based methods. First, the vertex-based two-stage method regards the vector extraction of the two types of objects as independent tasks, and only uses the separation structure to directly extract vector results from remote sensing images, ignoring the spatial correlation and feature complementarity of buildings and roads in remote sensing images; secondly, since the building morphology is mostly clustered, the area difference is large and the spatial position is close, the edge segmentation is unclear and it is easy to present a sticky state. Roads are generally straight or regular curves, and the shape does not change much, but it is easily blocked by trees and shadows on both sides, resulting in extraction holes and even road disconnection. The existing vertex-based two-stage method only uses simple convolution operations to predict building and road masks, which cannot solve the problems of building edge information loss and road extraction holes and disconnections. Summary of the invention

[0005] The purpose of the present invention is to propose a method for extracting building and road vectors from high-resolution remote sensing images, which solves the technical problems that the vertex-based two-stage method only uses simple convolution operations to predict building and road masks, and cannot solve the loss of building edge information and road extraction holes and disconnections.

[0006] Specifically, the present invention provides a method for extracting building and road vectors from high-resolution remote sensing images, comprising the following steps:

[0007] S1: Obtain a high-resolution remote sensing image dataset, perform data enhancement on the remote sensing image dataset, and obtain remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information;

[0008] S2: Construct a road and building collaborative extraction model based on encoder parameters. The road and building collaborative extraction model includes: a multi-scale feature extraction module, a feature interaction module, a dual-branch feature decoding module, and a vector topology prediction module;

[0009] S3: input a batch of remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information into the road and building collaborative extraction model for training to obtain a trained model;

[0010] S4. Obtain building and road vector results based on the trained model.

[0011] A storage medium stores instructions and data for realizing a method for extracting building and road vectors from high-resolution remote sensing images.

[0012] A high-resolution remote sensing image building and road vector extraction device comprises: a processor and the storage medium; the processor loads and executes instructions and data in the storage medium to implement a high-resolution remote sensing image building and road vector extraction method.

[0013] The beneficial effects provided by the present invention are as follows: a method for collaborative extraction of road and building vectors from high-resolution remote sensing images is proposed. A feature interaction module is designed to upsample and fuse building features and road features so that building features and road features fully interact with each other, solving the problem that a single extraction task cannot utilize the spatial correlation between buildings and roads, resulting in feature redundancy. In view of the morphological differences between buildings and roads, a dual-branch feature decoding module for heterogeneous targets is designed. For the unclear edge extraction and adhesion problems caused by the proximity of the spatial positions of buildings, a gated attention mechanism is designed and the boundary entropy loss of buildings is introduced, so that the model pays more attention to the building morphology. The clarity of the building edges is enhanced by minimizing the boundary entropy loss. For the special hole strip morphology of the road after occlusion, a holed strip convolution is designed to fit its features and enhance the perception of the road occlusion area. The quality of building and road mask extraction is enhanced, the problem of vertex position offset and missing in the vector results is effectively solved, and the topological accuracy of building and road vector extraction is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1It is a schematic diagram of the process of the method of the present invention;

[0015] Figure 2 It is a schematic diagram of the road and building collaborative extraction model of the present invention;

[0016] Figure 3 is a schematic diagram of a multi-scale feature extraction module in an embodiment of the present invention;

[0017] Figure 4 is a schematic diagram of a feature interaction module in an embodiment of the present invention;

[0018] Figure 5 Schematic diagram of road branches of a difference multi-scale feature extraction module in an embodiment of the present invention;

[0019] Figure 6 is a schematic diagram of a building branch in an embodiment of the present invention;

[0020] Figure 7 is a schematic diagram of a quantitative topology prediction module in an embodiment of the present invention;

[0021] Figure 8 It is a schematic diagram of the working of the hardware device of an embodiment of the present invention. DETAILED DESCRIPTION

[0022] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0023] Before formally describing the present invention, the scheme of the present invention is first generally described for easy understanding.

[0024] Please refer to Figure 1 The present invention provides a method for extracting building and road vectors from high-resolution remote sensing images, comprising:

[0025] S1: Obtain a high-resolution remote sensing image dataset, perform data enhancement on the remote sensing image dataset, and obtain remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information;

[0026] It should be noted that the data enhancement in step S1 includes: random rotation and random cropping.

[0027] As an embodiment, the remote sensing image data set used in the present invention is the Massachusetts building and road data set. The Massachusetts data set consists of aerial images of cities and suburbs in the Boston area, with an image size of 1500x1500 and a spatial resolution of 1 meter. According to the needs of the task, a total of 151 aerial images with both building labels and road labels were screened, and then each image was cropped to a size of 512*512 with a step size of 256. It was observed that some images in the data set were blocked by large areas. After removing these useless images, 3527 images were obtained, of which 3077, 200 and 250 were used for training, verification and testing, respectively.

[0028] S2: Construct a road and building collaborative extraction model based on encoder parameters. The road and building collaborative extraction model includes: a multi-scale feature extraction module, a feature interaction module, a dual-branch feature decoding module, and a vector topology prediction module;

[0029] It should be noted that the present invention constructs a road and building collaborative extraction model based on SAM encoder parameters. Figure 2 , Figure 2 It is a schematic diagram of the road and building vector collaborative extraction model of the present invention.

[0030] S3: input a batch of remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information into the road and building collaborative extraction model for training to obtain a trained model;

[0031] It should be noted that step S3 is specifically as follows:

[0032] S31, input a batch of remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information into a multi-scale feature extraction module to obtain intermediate features, road coding features and building coding features as input data for subsequent modules;

[0033] Each batch of remote sensing images is input into the multi-scale feature extraction module, which has four layers. The first three shallow layers extract local detail features of the image and retain the basic geometric forms of roads and buildings. The last deep layer extracts global semantic information and high-level representation features to generate the final coded features of roads and buildings.

[0034] The shallow layer features, road coding features and building coding features are saved separately as input data for subsequent modules;

[0035] As an exemplary embodiment, during the training process, the multi-scale feature extraction module uses the pre-trained SAM (Segment Everything Model) encoder parameters. It uses the VIT architecture suitable for high-resolution images. The present invention uses its smallest VIT-B variant to construct a multi-scale feature extraction module.

[0036] Please refer to Figure 3 Specifically, its multi-scale feature extraction module structure is divided into four layers, each of which is composed of multiple Transformer blocks stacked together: the first three levels of shallow layers extract local detail features of the image, focus on the texture, edge and low-level semantic information of the target, output multi-scale intermediate features, and retain the basic geometric forms of roads and buildings. The last layer extracts global semantic information and high-level representation features, generates coding features for distinguishing roads and buildings, and enhances the model's perception of large-scale and complex scenes. All parameters of the SAM encoder are adjusted throughout the training process, and the ability to extract roads and buildings in remote sensing images is further improved while retaining the original strong representation ability.

[0037] S32: Send the coded features of roads and buildings to the feature interaction module for preliminary fusion and decoding to obtain task shared features;

[0038] Please refer to Figure 4 , Figure 4 It is a schematic diagram of the characteristic interaction module of the present invention.

[0039] As an embodiment, the present invention sends the generated road coding features and building coding features to a feature interaction module, and interacts the coding features of different tasks through multi-scale feature alignment.

[0040] Subsequently, the aligned features are preliminarily decoded to obtain the task-shared features.

[0041] Multi-scale feature alignment is to eliminate the spatial differences or distribution inconsistencies between features of different scales through upsampling; preliminary decoding is to simply concatenate the road coding features and the building coding features in the channel dimension, and then send them to the convolution layer to form a feature representation shared by multiple tasks.

[0042] S33: After the road coding features, the intermediate features and the task shared features are sent to the dual-branch feature decoding module for decoding, the processed features are layered and fused with the shallow features, and added with the task shared features in the last layer to obtain the final road mask;

[0043] It should be noted that the dual-branch feature decoding module includes: a road branch and a building branch.

[0044] Please refer to Figure 5 , Figure 5Schematic diagram of the road branch of the dual-branch feature decoding module in an embodiment of the present invention. The dual-branch feature decoding module designs multi-scale and multi-directional strip convolution for the road branch to enhance the occlusion perception of the road, and then fuses the processed features with the shallow features in layers, and adds the task shared features in the last layer to obtain the final road mask.

[0045] Specifically, step S33 further includes the following steps:

[0046] S331, the road coding feature is replicated three times, and each feature is sent to three different scales of regular sequence convolution blocks based on dilated convolution, allowing the model to capture features at different scales and enhance the recognition ability of roads of different widths;

[0047] S332: In each convolution block, the input features are replicated four times. Then, four strip convolution kernels in different directions (vertical strips, horizontal strips, and two diagonal strips) are used for calculation, and the features in the four directions are fused to ensure that the model captures road features from different angles. The calculation formula is as follows:

[0048]

[0049] Where X is the input tensor, D h ,D w represents the direction tensor, which is used to determine the direction of the convolution kernel w, f is the dilation rate, which determines the scale of the convolution block. m is the convolution kernel index, u and v represent the index of the input tensor X, and n represents the number of convolution blocks.

[0050] S333: After the output feature maps of different scales are spliced, they are fused with shallow features as the output of the occlusion perception module, so that the module can use local features and global features at the same time. The calculation formula is:

[0051]

[0052] in is the feature fusion result of four directions, G i represents the channel function, and Z is the final output result of the occlusion perception module.

[0053] S34: Send the building encoding features, intermediate features and task shared features together to the dual-branch feature decoding module for decoding, and finally use the building boundary entropy loss to obtain a building mask with clear boundaries;

[0054] It should be noted that in step S34, decoding is performed through the building branch, and the building branch designs a gated attention convolution block to enhance the ability to perceive the spatial shape of the building.

[0055] As an example, please refer to Figure 6 , Figure 6 Schematic diagram of building branches in an embodiment of the present invention.

[0056] The boundary entropy loss calculation formula is as follows:

[0057] H(p i )=-p i log(p i )-(1-p i )log(1-p i )

[0058] where p i is the model's probability that pixel i belongs to the foreground (building).

[0059] Specifically, step S34 further includes the following steps:

[0060] S341: After passing the single-layer building coding features through the channel attention block and the spatial attention block, the coding features of the previous layer are sent to the door control attention block for calculation, and finally the two are added and sent to the next layer. After four layers of processing, the building decoding features are obtained.

[0061] S342: Obtain building labels and obtain the contour area mask of the building. Calculate the entropy loss for the edge area of ​​the building. Use the entropy loss to constrain the boundary, train the model until convergence, and obtain the final building mask.

[0062] S35: Processing the road mask and the building mask using a maximum suppression algorithm, and extracting vertices from the masks of the two respectively;

[0063] It should be noted that step S35 is specifically as follows:

[0064] S351: Based on experience, a threshold of 0.5 is set to filter out low-confidence pixels in the road mask and the building mask.

[0065] S352: traverse the pixels in descending order of pixel probability, and for the currently traversed pixel, remove all other pixels within a 16-pixel radius of the pixel.

[0066] S353: After the traversal is completed, the remaining pixels that have not been removed are retained as vertices, and the pixel position information will be used as the coordinates of the vertex.

[0067] S36: Sample the feature embedding of the corresponding vertex from the road and building coding features output in step S3, and input it into the vector topology prediction module together with the vertex. The vector topology prediction module predicts the connection relationship between the vertices to obtain a trained model.

[0068] S4. Obtain building and road vector results based on the trained model.

[0069] Please refer to Figure 7 , Figure 7 Schematic diagram of a quantitative topology prediction module in an embodiment of the present invention.

[0070] The specific process of obtaining the final building and road vector results is as follows:

[0071] S41: Perform a breadth-first search on all vertices, the currently selected node is set as the source vertex, and neighbor nodes are searched within a radius of 64 pixels from the source vertex.

[0072] S42: Use bilinear sampling from the road and building coded features output by S3 to obtain the corresponding vertex features, and input them into the vector topology prediction module together with the vertices as feature embedding.

[0073] S43: The topology prediction module predicts whether there is an edge between the source vertex and the neighboring node. The final classifier converts the prediction result into a probability value to obtain the final building and road vector result.

[0074] See also Figure 8 , Figure 8 4 is a schematic diagram of the working of the hardware device of an embodiment of the present invention, wherein the hardware device specifically comprises: a high-resolution remote sensing image building and road vector extraction device 401, a processor 402 and a storage medium 403.

[0075] A high-resolution remote sensing image building and road vector extraction device 401: The high-resolution remote sensing image building and road vector extraction device 401 implements the high-resolution remote sensing image building and road vector extraction method.

[0076] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the method for extracting building and road vectors from high-resolution remote sensing images.

[0077] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the method for extracting building and road vectors from high-resolution remote sensing images.

[0078] The beneficial effects of the present invention are as follows: a method for collaborative extraction of road and building vectors from high-resolution remote sensing images is proposed. A feature interaction module is designed to upsample and fuse building features and road features so that building features and road features fully interact with each other, solving the problem that a single extraction task cannot utilize the spatial correlation between buildings and roads, resulting in feature redundancy. In view of the morphological differences between buildings and roads, a dual-branch feature decoding module for heterogeneous targets is designed. For the unclear edge extraction and adhesion problems caused by the proximity of the spatial positions of buildings, a gated attention mechanism is designed and the boundary entropy loss of buildings is introduced, so that the model pays more attention to the building morphology. The clarity of the building edges is enhanced by minimizing the boundary entropy loss. For the special hole strip morphology of the road after occlusion, a strip convolution with holes is designed to fit its features to enhance the perception of the road occlusion area. The quality of mask extraction of buildings and roads is enhanced, the problem of vertex position offset and missing in the vector results is effectively solved, and the topological accuracy of building and road vector extraction is improved.

[0079] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for extracting building and road vectors from high-resolution remote sensing images, characterized in that: include: S1: Obtain a high-resolution remote sensing image dataset, perform data enhancement on the remote sensing image dataset, and obtain remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information; S2: Construct a road and building collaborative extraction model based on encoder parameters. The road and building collaborative extraction model includes: a multi-scale feature extraction module, a feature interaction module, a dual-branch feature decoding module, and a vector topology prediction module; S3: input a batch of remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information into the road and building collaborative extraction model for training to obtain a trained model; S4. Obtain building and road vector results based on the trained model.

2. A method for extracting building and road vectors from high-resolution remote sensing images as claimed in claim 1, characterized in that: The data enhancement in step S1 includes: random rotation and random cropping.

3. A method for extracting building and road vectors from high-resolution remote sensing images as claimed in claim 2, characterized in that: Step S3 is as follows: S31, input a batch of remote sensing images, road segmentation labels and road vector connection information, building labels and building vector connection information into a multi-scale feature extraction module to obtain intermediate features, road coding features and building coding features as input data for subsequent modules; S32: Send the coded features of roads and buildings to the feature interaction module for preliminary fusion and decoding to obtain task shared features; S33: After the road coding features, the intermediate features and the task shared features are sent to the dual-branch feature decoding module for decoding, the processed features are layered and fused with the shallow features, and added with the task shared features in the last layer to obtain the final road mask; S34: Send the building encoding features, intermediate features and task shared features together to the dual-branch feature decoding module for decoding, and finally use the building boundary entropy loss to obtain a building mask with clear boundaries; S35: Processing the road mask and the building mask using a maximum suppression algorithm, and extracting vertices from the masks of the two respectively; S36: Sample the feature embedding of the corresponding vertices from the road and building coding features output by S3, and input them together with the vertices into the vector topology prediction module. The vector topology prediction module predicts the connection relationship between the vertices to obtain a trained model.

4. A method for extracting building and road vectors from high-resolution remote sensing images as claimed in claim 3, characterized in that: The dual-branch feature decoding module includes: a road branch and a building branch.

5. A method for extracting building and road vectors from high-resolution remote sensing images as claimed in claim 4, characterized in that: In step S33, through road branch decoding, the road branch designs multi-scale and multi-directional strip convolution to enhance the occlusion perception ability of the road.

6. A method for extracting building and road vectors from high-resolution remote sensing images as claimed in claim 4, characterized in that: In step S34, decoding is performed through the building branch, and the building branch designs a gated attention convolution block to enhance the ability to perceive the spatial shape of the building.

7. A storage medium, characterized in that: The storage medium stores instructions and data for implementing a method for extracting building and road vectors from high-resolution remote sensing images as described in any one of claims 1 to 6.

8. A high-resolution remote sensing image building and road vector extraction device, characterized by: include: Processor and storage medium; the processor loads and executes instructions and data in the storage medium to implement a method for extracting high-resolution remote sensing image building and road vectors as described in any one of claims 1 to 6.