A Global-Local Decoding Method for Extracting Road Networks from Remote Sensing Images
Through the global-local decoding remote sensing image road network extraction method, a primary road network is built and iteratively searched using the Transformer encoder and decoder, noise samples and Hungarian algorithm, which solves the problems of low accuracy and long search time in remote sensing image road network extraction, and achieves efficient and complete road network extraction.
Patent Information
- Application Number
- CN202510388552.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the prior art, the remote sensing image road network extraction method has problems such as low accuracy, poor integrity and long search time, especially in multi-lane scenarios, the problems of road node connections and slow iteration search speed are prone to occur.
The global-local decoding method is adopted, and the road query is used to use the Transformer encoder and decoder, and the noise sample construction and Hungarian algorithm matching are combined to build a primary road network through topological connection networks, and iterative search and completion are performed locally to finally generate a complete road network.
It improves the accuracy and completeness of road network extraction, alleviates the unclear direction of road nodes in multi-lane scenarios, reduces the number of iterative vertices, and shortens the search time of road network.
Smart Images

Figure CN119884410B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image road network extraction, and in particular to a global-local decoding method for remote sensing image road network extraction. Background Art
[0002] Remote sensing image road network extraction refers to extracting the centerlines of roads in remote sensing images by means of manual drawing or machine learning model recognition. As an important basic geographic information, the road network reflects the structure and spatial layout of roads and has important applications in map updating, autonomous driving, and urban planning.
[0003] In the field of remote sensing image road network extraction research, since the method of manually drawing roads is time-consuming and costly, automatically extracting the road network in remote sensing images based on deep learning models is a current research hotspot. For example, there are mainly two categories. One is to use the global information of the image to output the road network in parallel. This method encodes the road network as a tensor representation, and conducts supervised training and prediction on whether each pixel is the center point of the road and the connection direction with other pixels, and then forms the road network result through a connection algorithm. However, there are connection confusion problems in the actual data production process. This method quantifies the road direction into 6 directions, so it is easy to misconnect to other lanes in the case of adjacent multi-lanes. The other is to use the local information of the image to iteratively output the road network. This method first uses a semantic segmentation network to extract the intersection points in the image as the initial points. Then, using the initial points as the center points of the local images, it iteratively searches for this road segment to form a road segment starting from the initial points until all the initial points are searched to form the final road network result. However, in the actual data production process, compared with the method of outputting the road network in parallel, there is a problem of slow retrieval speed. Each iterative retrieval of the next vertex in this method requires using the influence of the previous vertex area for model prediction, resulting in a significantly higher retrieval time than the method of outputting the road network in parallel. Summary of the Invention
[0004] The present invention provides a global-local decoding method for remote sensing image road network extraction to solve the problems of low accuracy and integrity of road network extraction and long road network retrieval time in the prior art when extracting road networks.
[0005] The present invention provides a global-local decoding method for remote sensing image road network extraction, including:
[0006] Obtain a remote sensing image slice of the ground where the road network to be extracted is located, and call a Transformer encoder to encode the remote sensing image slice to obtain an initial road query;
[0007] Construct a noise sample according to the ground truth coordinates of the ground where the road network to be extracted is located;
[0008] Input the initialized road query and noise samples into the Transformer decoder for decoding to obtain candidate road queries that have completed information interaction, and call the Hungarian algorithm to perform bipartite matching on the candidate road queries to obtain the target road query;
[0009] Call the global prediction head to map the target road query to the 2D coordinates and 36D direction descriptors of road nodes respectively, and input the combined features constructed using the 2D coordinates and the 36D direction descriptors into the topological connection network to obtain the primary road network of the ground to be extracted;
[0010] During the iterative retrieval of road endpoints in the primary road network, complete the road network for the road endpoints to obtain the final road network of the ground to be extracted.
[0011] In some embodiments, the calling the Transformer encoder to perform encoding processing on the remote sensing image slices to obtain the initialized road query includes:
[0012] Extract the multi-scale feature maps of each remote sensing image slice through the Transformer encoder;
[0013] Perform multi-scale dilated attention calculation on the multi-scale feature maps according to the ground sampling points to obtain the corresponding sampling point queries;
[0014] Select multiple sampling point queries with the largest confidence as the initialized road query.
[0015] In some embodiments, the constructing the noise samples according to the ground truth coordinates of the ground to be extracted includes:
[0016] For each ground truth coordinate in the ground to be extracted, add two random noise values respectively;
[0017] When the random noise value is less than or equal to the preset noise threshold, use the ground truth coordinate added with the random noise value as a positive sample;
[0018] When the random noise value is greater than the noise threshold and less than twice the noise threshold, use the ground truth coordinate added with the random noise value as a negative sample;
[0019] Use the positive and negative samples of each ground truth coordinate as the noise samples.
[0020] In some embodiments, the inputting the combined features constructed using the 2D coordinates and the 36D direction descriptors into the topological connection network to obtain the primary road network of the ground to be extracted includes:
[0021] Map the 2D coordinates to a coordinate feature representation through a coordinate multi-layer perceptron, and map the 36D direction descriptor to a direction feature representation through a direction multi-layer perceptron;
[0022] Concatenate the coordinate feature representation and the direction feature representation to obtain a combined feature;
[0023] In the topological connection network, for each target road node, determine the adjacent road nodes of the target road node, and merge the combined features of the target road node and the adjacent road nodes to obtain a merged feature;
[0024] Call an activation function to perform activation processing on the merged feature to obtain a corresponding feature tensor;
[0025] Call a fully connected layer to perform mapping processing on the feature tensor to obtain a corresponding tensor value;
[0026] When the tensor value is greater than the tensor threshold, connect the target road node and the adjacent road nodes to obtain a primary road network of the ground of the road network to be extracted.
[0027] In some embodiments, during the iterative retrieval of the road endpoints in the primary road network, network completion is performed on the road endpoints to obtain the final road network of the ground of the road network to be extracted, including:
[0028] Construct an initial retrieval point set according to the road endpoints in the primary road network, where the road endpoints are road nodes with less than 2 adjacent road nodes in the primary road network;
[0029] Select a center point from the initial retrieval point set according to a preset iterative strategy;
[0030] For each selected center point, extract a local remote sensing image from the remote sensing image slice and extract a local road network grid result from the primary road network;
[0031] Determine the local road query of the center point according to the local remote sensing image and the local road network grid result, and perform network completion on the primary road network according to the local road query to obtain the final road network of the ground of the road network to be extracted.
[0032] In some embodiments, the preset iterative strategy includes:
[0033] In the first round of iteration, arbitrarily select a road endpoint from the initial retrieval point set as the starting center point;
[0034] If the number of road nodes in the local road query of the starting center point is 0, randomly select a road endpoint from the candidate point set as the center point for the next iteration, where the candidate point set is the set of the remaining road endpoints in the initial retrieval point set except the starting center point;
[0035] If the number of road nodes in the local road query is 1, use the road node in the local road query as the center point for the next iteration;
[0036] If the number of road nodes in the local road query is 2, randomly select a road node in the local road query as the center point for the next iteration, and add the remaining road nodes in the local road query to the candidate point set;
[0037] When the candidate point set is an empty set, end the iteration.
[0038] In some embodiments, the method further includes:
[0039] Integrate the Transformer encoder, the Transformer decoder, and the topological connection network into a road network extraction model, and the training process of the road network extraction model includes:
[0040] Obtain a road network remote sensing image sample, where the road network remote sensing image sample includes the ground truth of road network connections;
[0041] Input the road network remote sensing image sample into the road network extraction model for forward propagation;
[0042] In the parallel stage of forward propagation, construct a parallel stage loss function according to the ground truth of road network connections and the predicted coordinate values, direction predicted values, and node connection predicted values of the road nodes output by the road network extraction model;
[0043] In the iterative stage of forward propagation, construct an iterative stage loss function according to the ground truth of road network connections and the predicted probability values of the road endpoints output by the road network extraction model;
[0044] Take the sum of the parallel stage loss function and the iterative stage loss function as the final loss function, and perform backpropagation in the road network extraction model through the final loss function to update the parameters of the road network extraction model.
[0045] In some embodiments, the construction process of the ground truth of road network connections includes:
[0046] In the ground truth road map of the road network remote sensing image sample, perform rasterization processing according to line segments of a preset pixel length;
[0047] When a road node in the road network remote sensing image sample is within the range of the line segment, determine the road node as a valid node;
[0048] For each valid node, determine the corresponding projection point in the true road map, where the projection point is the center line point in the true road map with the minimum Euclidean distance from the valid node;
[0049] According to the connection relationship of the projection points, determine the connectivity relationship of the road nodes in the road network remote sensing image sample, and connect the corresponding road nodes in the road network remote sensing image sample based on the connectivity relationship to obtain the road network connection truth value.
[0050] In some embodiments, the construction process of the parallel stage loss function includes:
[0051] Determine the true coordinate value, true direction value, and true node connection value of the road node from the road network connection truth value;
[0052] Construct a position loss according to the true coordinate value and the predicted coordinate value;
[0053] Construct a direction loss according to the true direction value and the predicted direction value;
[0054] Construct a connection loss according to the true node connection value and the predicted node connection value;
[0055] Weight the position loss, the direction loss, and the connection loss respectively according to the preset first position weight, direction weight, and connection weight, and sum the obtained weighted results to obtain the parallel stage loss function, where the values of the direction weight and the connection weight are positively correlated with the number of training rounds of the road network extraction model;
[0056] The construction process of the iterative stage loss function includes:
[0057] Determine the true probability value of the road end point from the road network connection truth value;
[0058] Construct a probability loss of the road end point according to the true probability value and the predicted probability value;
[0059] Weight the position loss and the probability loss respectively through the preset second position weight and probability weight, and sum the obtained weighted results to obtain the iterative stage loss function.
[0060] The present invention also provides a global-local decoding remote sensing image road network extraction device, including:
[0061] An encoding module, configured to obtain remote sensing image slices of the ground of the road network to be extracted, and call a Transformer encoder to perform encoding processing on the remote sensing image slices to obtain an initial road query;
[0062] A construction module, configured to construct noise samples according to the ground truth coordinates of the ground of the road network to be extracted;
[0063] A decoding module, configured to input the initial road query and the noise samples into a Transformer decoder for decoding processing to obtain a candidate road query that has completed information interaction, and call the Hungarian algorithm to perform bipartite matching on the candidate road query to obtain a target road query;
[0064] A global query module, configured to call a global prediction head to map the target road query to a 2D coordinate and a 36D direction descriptor of a road node respectively, and input a combined feature constructed by using the 2D coordinate and the 36D direction descriptor into a topological connection network to obtain a primary road network of the ground of the road network to be extracted;
[0065] A local query module, configured to perform road network completion on the road endpoints during the iterative retrieval of the road endpoints in the primary road network to obtain a final road network of the ground of the road network to be extracted.
[0066] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the remote sensing image road network extraction method of global-local decoding as described in any one of the above is implemented.
[0067] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the remote sensing image road network extraction method of global-local decoding as described in any one of the above is implemented.
[0068] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the remote sensing image road network extraction method of global-local decoding as described in any one of the above is implemented.
[0069] The global-local decoding remote sensing image road network extraction method provided by the present invention first realizes the query of road nodes in the remote sensing image through the Transformer model, maps the query of road nodes to the 2D coordinates and 36D direction descriptors of road nodes respectively, can refine the direction of road nodes, alleviate the problem of unclear direction indication of road nodes in multi-lane scenarios, and avoid the confusion of road node connections. In addition, the present invention also constructs noise samples for the prediction of road nodes, reduces semantic ambiguity at complex intersections, and improves the prediction accuracy of road points in complex road networks. Next, a road network construction process from global to local is constructed. Through global query, most of the road network can be extracted quickly. In local query, only the iterative search of road endpoints is performed to complete the road network, which not only ensures the integrity of road network extraction, but also can significantly reduce the number of iterative vertices and shorten the road network retrieval time compared with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art one by one. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0071] Figure 1 is a schematic flowchart of the global-local decoding remote sensing image road network extraction method provided by the present invention.
[0072] Figure 2 is a schematic principle diagram of the global-local decoding remote sensing image road network extraction method provided by the present invention.
[0073] Figure 3 is a schematic diagram of road node representation provided by the present invention.
[0074] Figure 4 is a schematic diagram of the effect of remote sensing image road network extraction provided by the present invention.
[0075] Figure 5 is a comparison diagram of the effects of various remote sensing image road network extractions provided by the present invention.
[0076] Figure 6 is a schematic structural diagram of the global-local decoding remote sensing image road network extraction device provided by the present invention.
[0077] Figure 7 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0079] The following describes the flowchart of the method for extracting road networks from remote sensing images with global-local decoding in the present invention in conjunction with the accompanying drawings. As Figure 1 shown, the method includes the following steps 101 to 105.
[0080] Step 101: Obtain a remote sensing image slice of the ground of the road network to be extracted, and call the Transformer encoder to perform encoding processing on the remote sensing image slice to obtain an initial road query.
[0081] First, obtain the remote sensing image of the ground of the road network to be extracted, and then crop this remote sensing image to obtain a remote sensing image slice. The cropping process can be to evenly crop the remote sensing image with a size of h×w×c into remote sensing image slices with a size of s×s×c through a non-overlapping sliding window with a size of s×s. Here, h is the image height, w is the image width, c is the number of image channels, and s is the size of the sliding window with the same width and height. If the height of the remote sensing image slice in the last row of the sliding window is less than s, then the part with a width less than s in the last column is cropped backward to obtain a remote sensing image slice with a size of s×s×c.
[0082] Then, for each remote sensing image slice, call the Transformer encoder to perform encoding processing on the remote sensing image slice to obtain an initial road query. Here, first extract the multi-scale feature map of the remote sensing image slice through the Transformer encoder, and then perform sampling point query on the multi-scale feature map to obtain an initial road query. In the embodiments of the present invention, the road query can be the pixel coordinate values representing the road part in the remote sensing image, which is used to determine the specific road nodes in the remote sensing image slice.
[0083] Specifically, as Figure 2 shown, first extract the multi-scale feature map of each remote sensing image slice through the Transformer encoder. Here, use the backbone network in the Transformer encoder to extract feature maps with feature scales of H / 4×W / 4×256, H / 8×W / 8×256, H / 16×W / 16×256, and H / 32×W / 32×256. During the process of extracting the feature maps, corresponding position encodings will also be added.
[0084] Next is the Transformer encoding process. The Transformer encoder consists of a multi-scale deformable attention module and a multi-layer perceptron, aiming to fuse remote information into multi-scale features. First, multiple ground sampling points are determined on the multi-scale feature map, and then multi-scale dilated attention calculation is performed on the multi-scale feature map according to the ground sampling points to obtain corresponding sampling point queries. The multi-scale dilated attention (MSDA) calculation is implemented through the multi-scale deformable attention module, and the formula is as follows:
[0085] (1)
[0086] In the above formula (1), , q, and represent the multi-scale feature map, the road query, and the reference point coordinates corresponding to the road query respectively. , , and represent the m-th attention head, the l-th feature level, and the k-th sampling point in the model respectively. and are both learnable weight matrices. represents the attention weight of the road query q, the -th attention head, the -th level feature, and the -th sampling point. is to scale the normalized reference point to adapt to the -th feature layer. represents the learnable sampling offset parameter. represents the sampling position of the road query on the feature map .
[0087] As shown in Figure 2 , during the Transformer encoding process, through the above formula (1), the Transformer encoder will output multiple sampling point queries, and these sampling point queries may all represent road regions. Then, the top N sampling point queries are selected here. Specifically, multiple sampling point queries with the largest confidence are selected as the initial road queries. That is, N is taken as 500. By calculating the confidence of the sampling point queries and then sorting them from largest to smallest, finally, the 500 sampling point queries with the largest confidence are selected from the head of the sequence. These sampling point queries have the greatest possibility of representing road regions and can therefore be used as the initial road queries.
[0088] Step 102: Construct noise samples according to the ground truth coordinates of the ground to be extracted road network.
[0089] To enhance the adaptability to noise, the embodiment of the present invention also constructs a denoising strategy, by adding noise to the ground truth coordinates (Ground Truth, GT) in the initialization road query and participating in the prediction process of the road query. By predicting the road area under the influence of noise, the possible noise influence can be effectively eliminated during the prediction of the road area.
[0090] In some embodiments, noise samples are constructed according to the ground truth coordinates of the ground of the road network to be extracted. Specifically, first, for each ground truth coordinate in the ground of the road network to be extracted, two random noise values are respectively added, which are respectively expressed as ( ), and ( ). Then, the embodiment of the present invention performs the distinction between positive and negative samples on the ground truth coordinates with added random noise.
[0091] When the random noise value is less than or equal to the preset noise threshold, the ground truth coordinate with the added random noise value is used as a positive sample. The noise threshold value is , if and are satisfied, it means that the ground truth coordinate with the added random noise value can be used as a positive sample.
[0092] When the random noise value is greater than the noise threshold and less than twice the noise threshold, the ground truth coordinate with the added random noise value is used as a negative sample. Here, if and are satisfied, it means that the ground truth coordinate with the added random noise value can be used as a negative sample. Among them, means the preset noise magnitude hyperparameter, and the value here is 5.
[0093] In this way, each ground truth coordinate can construct a positive sample and a negative sample, and then the positive and negative samples of each ground truth coordinate are used as noise samples.
[0094] In the embodiment of the present invention, when predicting the road query by extracting multi-scale feature maps, by constructing corresponding noise samples to participate in the prediction, it can also perform prediction when there may be noise affecting the road prediction, which can ensure the adaptability of the road network extraction process to noise, reduce the semantic ambiguity of complex intersections, and improve the prediction accuracy of road points in complex road networks.
[0095] Step 103: Input the initialization road query and the noise samples into the Transformer decoder for decoding processing to obtain candidate road queries that have completed information interaction, and call the Hungarian algorithm to perform bipartite matching on the candidate road queries to obtain the target road query.
[0096] As Figure 2As shown, after constructing the noise samples, the next step is the decoding process of the Transformer decoder, that is, Transformer decoding. Here, the initial road query and the noise samples are input into the Transformer decoder for decoding processing to obtain candidate road queries that have completed information interaction. The structure of the Transformer decoder is the same as that of the encoder, and it is also composed of 6 layers of multi-head attention modules, multi-scale deformable attention modules, and multi-layer perceptrons. The candidate road queries can not only represent possible road regions but also represent the information interaction between roads, thus facilitating the accurate determination of road nodes in the road regions. In this way, the process of query extraction ends.
[0097] Next, in order to determine the accurate road node information based on the road queries, the embodiment of the present invention performs bipartite matching on the candidate road queries by calling the Hungarian algorithm to obtain the target road queries. Here, the ground truth of each road information can be determined from the remote sensing image of the road network to be extracted, and then the candidate road queries are bipartitely matched with the ground truth of the road information through the Hungarian algorithm, and the optimal match is selected. Then, the target road queries are screened out from the candidate road queries according to the optimal match. The target road queries can accurately reflect the road nodes in the road regions and facilitate road network construction.
[0098] Step 104: Call the global prediction head to map the target road queries to the 2D coordinates and 36D direction descriptors of the road nodes respectively, and input the combined features constructed by the 2D coordinates and 36D direction descriptors into the topological connection network to obtain the primary road network of the road network ground to be extracted.
[0099] After determining the target road queries of the road network ground to be extracted through step 103, the first-stage global query of road network extraction is then performed. During the global query process, first, the coordinates and directions of the road nodes are determined according to the target road queries.
[0100] As Figure 2 shown, during the global query process, for the target road queries, here the global prediction head is called to map the target road queries to the 2D coordinates and 36D direction descriptors of the road nodes respectively. The global prediction head is composed of two different multi-layer perceptrons (MLPs), specifically, the coordinate multi-layer perceptron and the direction multi-layer perceptron, which are respectively used for coordinate prediction and direction prediction of road nodes.
[0101] See Figure 3 , Figure 3 illustrates a schematic diagram of the 2D coordinates and 36D direction descriptors of a road node. As Figure 3As shown on the left, the center of the circle represents the coordinates of the road center point. The horizontal right direction is 0 degrees, and each increment is 10 degrees counterclockwise, with a total of 36 increments. When the direction reaches 360 degrees, it coincides with the 0-degree direction, and the direction value is 0. There are adjacent road nodes in the 3, 9, 21, and 27 directions of this road node, and its direction representation is as follows Figure 3 As shown on the right, the road node coordinates are the coordinates of the road center point, belonging to a 2D vector (X, Y). Among the vectors of the 36D direction descriptor, the 3rd, 9th, 21st, and 27th positions are 1, and the rest are 0.
[0102] Next, for each road node, the combined features constructed from the 2D coordinates and the 36D direction descriptor are input into the topological connection network to obtain the primary road network of the ground to be extracted. Here, the coordinates and directions of the road nodes are used as the features of the road nodes and input into the topological connection network. In the topological connection network, by calculating the tensor values of adjacent road nodes and selecting the corresponding adjacent road nodes for connection according to the preset tensor threshold, the primary road network of the ground to be extracted is obtained.
[0103] The following specifically introduces the process of constructing the primary road network. As follows Figure 2 As shown, first, the features of the road nodes are processed. Here, the 2D coordinates are mapped to a coordinate feature representation (Point Embedding) through a coordinate multi-layer perceptron, and the 36D direction descriptor is mapped to a direction feature representation Direct Embedding through a direction multi-layer perceptron. Then, the fusion of the coordinate feature and the direction feature is realized. The coordinate feature representation and the direction feature representation are concatenated to obtain the combined feature, and finally, it is sent into the global topological connection network to perform the construction of the primary road network.
[0104] The input received by the topological connection network is the combined features of the road node and its adjacent road nodes. The topological connection network can be regarded as a fully connected neural network model. Therefore, in the topological connection network, for each target road node, first determine the adjacent road nodes of the target road node, and merge the combined features of the target road node and the adjacent road nodes to obtain the merged feature. Each road node has adjacent road nodes. Therefore, for each road node here, first determine the adjacent road nodes, and then the combined features of the road node and the combined features of the adjacent road nodes are merged to obtain the merged feature, denoted as , where is the maximum number of adjacent nodes.
[0105] Next, call the activation function to perform activation processing on the merged feature to obtain the corresponding feature tensor. During the activation processing, first project the merged feature into The eigenvector is then activated through the Relu activation function of the fully connected layer to obtain The eigenvector, and then the eigenvector is activated through the sigmoid function of the fully connected layer to obtain The eigen-tensor.
[0106] Finally, the fully connected layer is called to perform a mapping process on the eigen-tensor to obtain the corresponding tensor value. Here, The eigen-tensor is input into a fully connected layer, and a mapping process is performed through the sigmoid function of the fully connected layer to obtain the corresponding tensor value, denoted as .
[0107] When constructing the primary road network, an embodiment of the present invention presets a tensor threshold of 0.5. When the tensor value is greater than the tensor threshold, the target road node is connected to the adjacent road node to obtain the primary road network on the ground of the road network to be extracted. Here, by performing tensor calculations on the combined features of each target road node, when the calculated tensor number is greater than 0.5, the target road node can be connected to the adjacent road node. In this way, each target road node can select the corresponding adjacent road node for connection, and finally a primary road network is formed.
[0108] In the embodiment of the present invention, in the process of constructing the road network from global to local, through the global query stage, a preliminary road network structure is first constructed according to the predicted road nodes, which can quickly extract most of the road network, with high efficiency and reduced retrieval time and calculation amount for subsequent road network completion.
[0109] Step 105: During the iterative retrieval of the road endpoints in the primary road network, the road network is completed for the road endpoints to obtain the final road network on the ground of the road network to be extracted.
[0110] Through step 104, the first-stage global query process of road network extraction is completed. Next, enter the second stage of road network extraction, the local query process. This part is mainly for completing the dead-end roads in the road network, that is, for completing the road endpoints. The road endpoints are the road nodes with the number of adjacent road nodes less than 2 in the primary road network. In this process, first determine the road endpoints to be completed, and according to the local remote sensing image and the local road network raster result corresponding to the road endpoints, perform local road queries, and then perform road network completion. In this process, an iterative retrieval process of the road endpoints needs to be performed, and road network completion is performed once in each retrieval process, so as to obtain the final road network on the ground of the road network to be extracted. The schematic diagram of the effect of the final road network can be as shown in Figure 4 shown, Figure 4 The yellow road node line segments in it are the extracted final road network.
[0111] Specifically, first, an initial retrieval point set is constructed based on the road endpoints in the primary road network. The road endpoints are road nodes in the primary road network where the number of adjacent road nodes is less than 2. Here, all road endpoints are added to the initial retrieval point set, denoted as , where , represents the number of adjacent nodes for subsequent iterative retrieval.
[0112] Then, according to a preset iterative strategy, a center point is selected from the initial retrieval point set. For each selected center point, a local remote sensing image is extracted from the remote sensing image slice, and a local road network raster result is extracted from the primary road network. Specifically, from the remote sensing image slice, a local remote sensing image with a width and height of 128×128 is intercepted with the center point as the center, and a local road network raster result with a width and height of 128×128 is also extracted from the corresponding primary road network.
[0113] Next, a local road query for the center point is determined based on the local remote sensing image and the local road network raster result. As Figure 2 shown, during the local query process, first, both the local remote sensing image and the local road network raster result are input into the backbone network for feature extraction. The local road network raster result undergoes feature extraction through the backbone network to obtain the corresponding mask feature, while the local remote sensing image undergoes feature extraction through the backbone network to obtain the corresponding RGB feature. The backbone network can be a Transformer model. Then, the mask feature and the RGB feature are fused, and finally, the fused feature is input into the query extraction module for query extraction to output the corresponding local road query.
[0114] Next, based on the local road query, the road network in the primary road network is completed to obtain the final road network of the ground to be extracted. Here, still, the local road query is feature-mapped through the prediction head to map the local road query to the coordinates of the corresponding road endpoints, and then input into the topological connection network for tensor calculation. Then, the adjacent road nodes with tensor values less than the tensor threshold of 0.5 are determined, and the road endpoints and the adjacent road nodes are connected to complete the road network.
[0115] Of course, in some embodiments, to improve the accuracy of the local road query, as Figure 2 shown, the embodiment of the present invention can also intercept multiple ( with the center point Figure 2Two local remote sensing images with a size of 128×128 are exemplified (there are two in total), and multiple local road network raster results of 128×128 are also extracted from the corresponding primary road network. The first local remote sensing image and the local road network raster result extract features through the backbone network, and then output the corresponding local road query through the query extraction module. After that, this local road query, together with the second local remote sensing image and the local road network raster result, extracts features through the backbone network again, and then outputs the corresponding local road query through the query extraction module. The local road query finally output in this process is used as the final local road query of the road endpoints.
[0116] In the embodiment of the present invention, in the local query process of the second stage of road network extraction, the central point is selected from the initial retrieval point set through iteration, and the local road query of each road endpoint is calculated to complete the primary road network. In the local query, only the iterative search of the road endpoints is performed to complete the road network, which not only ensures the integrity of the road network extraction, but also can significantly reduce the number of iterative vertices and shorten the road network retrieval time compared with the prior art.
[0117] In some embodiments, in order to ensure the integrity of the road network extraction, the embodiment of the present invention selects the central point from the initial retrieval point set according to the preset iteration strategy to complete the road network complement of the road endpoints. The specific iteration process of the preset iteration strategy is introduced below.
[0118] The main purpose of the iteration strategy is to determine the central point selected in each round of iteration process. Here, in the first round of iteration, an arbitrary road endpoint is selected from the initial retrieval point set as the starting central point. That is, an arbitrary road endpoint is selected from the initial retrieval point set as the starting central point. Next, the neighbor road nodes of the road endpoint are judged.
[0119] If the number of road nodes in the local road query of the starting central point is 0, a road endpoint is randomly selected from the candidate point set as the central point of the next round of iteration, where the candidate point set is the point set composed of the remaining road endpoints in the initial retrieval point set except the starting central point.
[0120] Here, if it is judged that the number of road nodes in the local road query of the starting central point is 0, it means that in this local road query, the neighbor road nodes of the starting central point cannot be found, so the starting central point cannot be used for road network construction and is directly skipped, and this round of iteration ends. Then, another road endpoint is selected from the remaining road endpoints in the initial retrieval point set except the starting central point as the central point of the next round of iteration for local road query.
[0121] If the number of road nodes in the local road query is 1, then the road node in the local road query is used as the center point for the next iteration.
[0122] Here, if it is determined that the number of road nodes in the local road query of the starting center point is 1, it means that in this local road query, a neighbor road node of the starting center point can be found. Then, this road node in the local road query is used as the center point for the next iteration. Furthermore, according to the local road query, the road network construction between the starting center point and this road node is performed to complete the road network.
[0123] If the number of road nodes in the local road query is 2, then a road node randomly selected from the local road query is used as the center point for the next iteration, and the remaining road nodes in the local road query are added to the candidate point set.
[0124] Here, if it is determined that the number of road nodes in the local road query of the starting center point is 2, it means that in this local road query, two neighbor road nodes of the starting center point can be found. Then, one of the two neighbor road nodes is randomly selected as the center point for the next iteration. Furthermore, according to the local road query, the road network construction between the starting center point and the randomly selected road node in the local road query is performed to complete the road network. And the road node not selected from the two neighbor road nodes is added to the candidate point set.
[0125] Among them, in each iteration process, after the road network is completed for the selected center point, the next road endpoint is continuously selected from the candidate point set to perform the iteration. When the candidate point set is an empty set, it means that the road network completion of all road endpoints is completed, and the iteration ends.
[0126] In the embodiment of the present invention, by performing iterative search for road endpoints, the local road query is calculated for each road endpoint to judge the neighbor road nodes of the road endpoints, thereby realizing the completion of the road network. That is, it ensures the integrity of the road network extraction. Compared with the prior art, it can greatly reduce the number of iterative vertices and shorten the road network retrieval time.
[0127] In some embodiments, considering the complex road network structure and numerous road nodes of some road network ground to be extracted, there are still defects in the accuracy of road network extraction, and some requirements for road network structure extraction with higher requirements cannot be met. Therefore, the embodiments of the present invention also propose a method for performing road network extraction using a road network extraction model. By integrating a Transformer encoder, a Transformer decoder, and a topological connection network into a road network extraction model, that is, integrating the Transformer encoder used in step 101, the Transformer decoder used in step 103, and the topological connection network used in step 104 into a road network extraction model. The training process of the road network extraction model is introduced below.
[0128] First, obtain road network remote sensing image samples, where the road network remote sensing image samples include road network connection ground truths. Here, the road network remote sensing image samples are remote sensing images of some road network ground to be extracted, and the road network connection ground truths are ground truth data extracted from the ground truth road map with the road network structure already annotated, serving as training labels to construct a loss function. The ground truth data specifically includes: the true coordinates of road nodes, the true directions, the true node connections, and the true probabilities of road endpoints.
[0129] At the start of training, input the road network remote sensing image samples into the road network extraction model for forward propagation, and the forward propagation is completed in two stages.
[0130] In the parallel stage of forward propagation, construct a parallel stage loss function based on the road network connection ground truths and the predicted coordinate values, direction predicted values, and node connection predicted values of road nodes output by the road network extraction model.
[0131] Specifically, when the road network extraction model performs road network extraction, it will calculate road queries to predict the predicted coordinate values and direction predicted values of each road node. When constructing the primary road network, determine the connection relationship of road nodes through the predicted coordinate values and direction predicted values of road nodes, and thus obtain the predicted node connection values of the corresponding road nodes. Therefore, during the forward propagation process, a parallel stage loss function can be constructed based on the predicted coordinate values, direction predicted values, node connection predicted values, and the ground truth data included in the road network connection ground truths to train the accuracy of the road network extraction model for road network extraction.
[0132] In the iterative stage of forward propagation, construct an iterative stage loss function based on the road network connection ground truths and the predicted probability values of road endpoints output by the road network extraction model.
[0133] During the process of road network completion, first, the road network extraction model determines each road endpoint in the primary road network. The position information of the road endpoints also includes the predicted probability of the road endpoints. In the embodiment of the present invention, the predicted probability value of the road endpoints and the true value data included in the true value of road network connection are selected here to construct the loss function in the iterative stage, so as to train the road network extraction model for the accuracy of road network completion of the road endpoints.
[0134] Finally, the sum of the loss function in the parallel stage and the loss function in the iterative stage is used as the final loss function, and the parameters of the road network extraction model are updated through the backpropagation of the final loss function in the road network extraction model.
[0135] Through the preset maximum number of iterations, when the final loss function starts to converge or reaches the maximum number of iterations, the iteration is stopped and the training process is ended. The trained road network extraction model can then perform road network extraction on the remote sensing images of various road networks to be extracted on the ground and obtain the corresponding final road network.
[0136] In the embodiment of the present invention, an end-to-end road network extraction method is designed. By directly executing road network extraction through the designed road network extraction model, when training the road network extraction model, by constructing the loss function in the parallel stage and the loss function in the iterative stage, not only the accuracy of the road network extraction model for road network extraction is ensured, but also the accuracy of the road network extraction model for road network completion of the road endpoints is ensured. It can meet some requirements with relatively high requirements for the extraction of road network structures.
[0137] Further, on the basis of the above embodiment, the construction process of the true value of road network connection is introduced below. As Figure 5 shown, Figure 5 in (a) is the true value road map of the road network structure already marked in a certain road network remote sensing image sample. In the true value road map of the road network remote sensing image sample, first, rasterization processing is performed according to line segments with a preset pixel length, and this pixel length can be 5 pixels.
[0138] When the road nodes in the road network remote sensing image sample are within the range of the line segment, the road nodes are determined as valid nodes. Through these methods, the road nodes that can be displayed in the true value road map can be determined as valid nodes.
[0139] For each valid node, the corresponding projection point is determined in the true value road map. In this way, the mapping relationship between the road nodes in the road network remote sensing image sample and the projection points in the true value road map can be established. The projection point is the center line point with the smallest Euclidean distance from the valid node in the true value road map. According to this standard, the projection points and the road nodes can be mapped one by one.
[0140] In the true-value road map, the projection points are the road nodes of the already-standardized road network structure. Therefore, there is a connection relationship between each pair of projection points. Since the projection points have been mapped one-to-one with the road nodes in the road network remote sensing image sample, the connectivity relationship of the road nodes in the road network remote sensing image sample can be determined according to the connection relationship of the projection points.
[0141] Finally, based on the connectivity relationship, the corresponding road nodes in the road network remote sensing image sample are connected to obtain the true value of the road network connection. By connecting the road nodes, the road network structure of the road network remote sensing image sample can be determined, and then it is easy to determine the true value data such as the true coordinates of the road nodes, the true direction values, the true node connection values, and the true probability values of the road endpoints one by one, as the true value of the road network connection.
[0142] In the embodiment of the present invention, by establishing the mapping relationship between the road network remote sensing image sample and the true-value road map, the true road network structure of the road network remote sensing image sample is determined, and further the true value of the road network connection. This can be used as the training label of the road network extraction model to construct the final loss function to train the road network extraction model, ensuring the training effectiveness of the model and the accuracy of the road network extraction.
[0143] In some embodiments, the road network extraction model depends on the constructed final loss function, and the final loss function is composed of the parallel-phase loss function and the iterative-phase loss function These two parts will be described one by one below, including the construction processes of the parallel-phase loss function and the iterative-phase loss function.
[0144] Regarding the parallel-phase loss function , in the embodiment of the present invention, the true coordinate values , true direction values , and true node connection values of the road nodes are first determined from the true value of the road network connection.
[0145] Next, according to the true coordinate values and the predicted coordinate values , a position loss is constructed. The position loss uses the L1 loss function and is expressed as the following formula:
[0146] (2)
[0147] Moreover, according to the true direction values and the predicted direction values , a direction loss is constructed. Here, considering that most of the road nodes in the road network have two directions, and only at the intersection part will there be three or more directions. To handle the problem of class imbalance, the direction loss Constructed by the Focal Loss function, expressed as:
[0148] (3)
[0149] In addition, according to the true value of node connection and the predicted value of node connection , construct the connection loss . In the road network connection part, the connection between road nodes is a binary classification problem, that is, whether the connection with neighboring road nodes is correct or failed. The loss function of the binary classification problem is constructed by the BCEWithLogitsLoss function, expressed as:
[0150] (4)
[0151] When constructing the loss function in the parallel stage , there should be a focus on the prediction of coordinates, directions, and connection relationships. Therefore, in the embodiments of the present invention, the first position weight , direction weight , and connection weight are respectively preset for the loss functions of the three parts to balance the loss values of the three parts.
[0152] According to the preset first position weight , direction weight , and connection weight , weight the position loss , direction loss , and connection loss respectively, and sum the obtained weighted results to obtain the loss function in the parallel stage , expressed as:
[0153] (5)
[0154] Among them, , , are the weights of the position loss , direction loss , and connection loss respectively. However, in the actual training process of the road network extraction model, the position prediction of road nodes is unstable in the early stage. Therefore, in the early training of the road network extraction model, the position loss of road nodes is the main one and will be assigned a greater weight, while the values of the direction weight and the connection weight are positively correlated with the training rounds epoch of the road network extraction model. That is, as the number of iterations increases, in the middle and late stages of training, the direction weight and connection weights The corresponding assignment will also increase.
[0155] Specifically, as the number of iterations increases, the convergence direction of the road network extraction model is obvious, and the direction loss and connection loss weights increase exponentially. Among them, the direction weight is specifically set to , and the connection loss is specifically set to , where epoch is the current training round of the road network connection model.
[0156] For the loss function in the iteration stage, the embodiment of the present invention still determines the true probability value of the road endpoint from the true value of the road network connection. Here, the true probability value is the true prediction probability of the road endpoint. In the true value road map, the connection relationship between the road endpoint and the neighbor nodes is determined one by one, and the probability is 1. Therefore is 1.
[0157] Next, according to the true probability value and the predicted probability value , the probability loss of the road endpoint is constructed, and it is constructed through the BCEWithLogitsLoss loss function, which is expressed as:
[0158] (6)
[0159] In addition, the loss function in the iteration stage also includes the coordinate loss part of the road node, that is . Here, through the preset second position weight and probability weight , the position loss and probability loss are weighted respectively, and the weighted results are summed to obtain the loss function in the iteration stage, which is expressed as:
[0160] (7)
[0161] In the process of constructing the loss function of the road network extraction module in the embodiments of the present invention, a parallel-phase loss function is constructed by predicting the coordinates, directions, and connections of road nodes, enabling the model to accurately predict the positions, directions, and connection relationships of road nodes in the primary road network. The iterative-phase loss function is constructed by the predicted probability values of road endpoints, enabling the model to complete the road network complementation of road endpoints in the primary road network and ensuring the accuracy of road network extraction.
[0162] To verify the effectiveness of the road network extraction model of the present invention, the embodiments of the present invention also set up a comparative experiment, using models for road network extraction commonly used in the prior art, such as Sat2Graph and RNGDet, for comparative experiments and comparing them with the ground truth road map. The comparative results for the extraction of a certain ground road network are as Figure 5 shown, where (b) is the road network extraction effect of the Sat2Graph model, (c) is the road network extraction effect of the RNGDet model, (d) is the road network extraction effect of the road network extraction model proposed by the present invention, and (a) is the ground truth road map of this area. It can be Figure 5 proven that the road network extraction model proposed by the present invention is superior to models for road network extraction such as Sat2Graph and RNGDet in terms of road network extraction integrity and road node accuracy.
[0163] The global-local decoding remote sensing image road network extraction device provided by the present invention will be described below. The global-local decoding remote sensing image road network extraction device described below can be mutually referred to the global-local decoding remote sensing image road network extraction method described above.
[0164] As Figure 6As shown in the figure, the global-local decoding remote sensing image road network extraction device specifically includes: an encoding module 601, configured to obtain a remote sensing image slice of the road network ground to be extracted, and call a Transformer encoder to perform encoding processing on the remote sensing image slice to obtain an initialized road query; a construction module 602, configured to construct a noise sample according to the ground truth coordinates of the road network ground to be extracted; a decoding module 603, configured to input the initialized road query and the noise sample into a Transformer decoder for decoding processing to obtain a candidate road query that has completed information interaction, and call the Hungarian algorithm to perform bipartite matching on the candidate road query to obtain a target road query; a global query module 604, configured to call a global prediction head to map the target road query into a 2D coordinate and a 36D direction descriptor of a road node respectively, and input a combined feature constructed by using the 2D coordinate and the 36D direction descriptor into a topological connection network to obtain a primary road network of the road network ground to be extracted; a local query module 605, configured to complete the road network for the road endpoints during the iterative retrieval of the road endpoints in the primary road network to obtain a final road network of the road network ground to be extracted.
[0165] It should be noted that the beneficial effects of the global-local decoding remote sensing image road network extraction device here correspond to those of the global-local decoding remote sensing image road network extraction method in the above text. Therefore, the beneficial effects of the global-local decoding remote sensing image road network extraction device are not elaborated here.
[0166] Figure 7 An entity structure diagram of an electronic device is exemplified, such as Figure 7As shown in the figure, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete communication with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute the method for extracting road networks from remote sensing images with global-local decoding. The method includes: obtaining a remote sensing image slice of the ground of the road network to be extracted, and calling a Transformer encoder to encode the remote sensing image slice to obtain an initialized road query; constructing a noise sample according to the ground truth coordinates of the ground of the road network to be extracted; inputting the initialized road query and the noise sample into a Transformer decoder for decoding to obtain a candidate road query that has completed information interaction, and calling the Hungarian algorithm to perform bipartite matching on the candidate road query to obtain a target road query; calling a global prediction head to map the target road query to the 2D coordinates and 36D direction descriptors of road nodes respectively, and inputting the combined features constructed by using the 2D coordinates and the 36D direction descriptors into a topological connection network to obtain a primary road network of the ground of the road network to be extracted; during the iterative retrieval of road endpoints in the primary road network, completing road network supplementation for the road endpoints to obtain the final road network of the ground of the road network to be extracted.
[0167] In addition, when the logical instructions in the above-mentioned memory 730 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0168] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the global-local decoding remote sensing image road network extraction method provided by the above-mentioned various methods. The method includes: obtaining a remote sensing image slice of the ground of the road network to be extracted, and calling a Transformer encoder to perform encoding processing on the remote sensing image slice to obtain an initial road query; constructing a noise sample according to the ground truth coordinates of the ground of the road network to be extracted; inputting the initial road query and the noise sample into a Transformer decoder for decoding processing to obtain a candidate road query that has completed information interaction, and calling the Hungarian algorithm to perform bipartite matching on the candidate road query to obtain a target road query; calling a global prediction head to map the target road query to the 2D coordinates and 36D direction descriptors of road nodes respectively, and inputting the combined features constructed by using the 2D coordinates and the 36D direction descriptors into a topological connection network to obtain a primary road network of the ground of the road network to be extracted; during the iterative retrieval of the road endpoints in the primary road network, completing the road network for the road endpoints to obtain the final road network of the ground of the road network to be extracted.
[0169] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the global-local decoding remote sensing image road network extraction method provided by the above-mentioned various methods. The method includes: obtaining a remote sensing image slice of the ground of the road network to be extracted, and calling a Transformer encoder to perform encoding processing on the remote sensing image slice to obtain an initial road query; constructing a noise sample according to the ground truth coordinates of the ground of the road network to be extracted; inputting the initial road query and the noise sample into a Transformer decoder for decoding processing to obtain a candidate road query that has completed information interaction, and calling the Hungarian algorithm to perform bipartite matching on the candidate road query to obtain a target road query; calling a global prediction head to map the target road query to the 2D coordinates and 36D direction descriptors of road nodes respectively, and inputting the combined features constructed by using the 2D coordinates and the 36D direction descriptors into a topological connection network to obtain a primary road network of the ground of the road network to be extracted; during the iterative retrieval of the road endpoints in the primary road network, completing the road network for the road endpoints to obtain the final road network of the ground of the road network to be extracted.
[0170] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for extracting road networks from remote sensing images with global-local decoding, characterized in that Including: Obtain remote sensing image slices of the ground of the road network to be extracted, and call a Transformer encoder to encode the remote sensing image slices to obtain an initial road query; Construct noise samples according to the ground truth coordinates of the ground of the road network to be extracted; Input the initial road query and the noise samples into a Transformer decoder for decoding processing to obtain candidate road queries that have completed information interaction, and call the Hungarian algorithm to perform bipartite matching on the candidate road queries to obtain target road queries; Call a global prediction head to map the target road queries to the 2D coordinates and 36D direction descriptors of road nodes respectively, and input the combined features constructed using the 2D coordinates and the 36D direction descriptors into a topological connection network to obtain a primary road network of the ground of the road network to be extracted; During the iterative retrieval of road endpoints in the primary road network, complete the road network for the road endpoints to obtain the final road network of the ground of the road network to be extracted; During the iterative retrieval of road endpoints in the primary road network, completing the road network for the road endpoints to obtain the final road network of the ground of the road network to be extracted includes: Construct an initial retrieval point set according to the road endpoints in the primary road network, where the road endpoints are road nodes with the number of adjacent road nodes less than 2 in the primary road network; Select a center point from the initial retrieval point set according to a preset iterative strategy; For each selected center point, extract a local remote sensing image from the remote sensing image slices and extract a local road network grid result from the primary road network; Input both the local remote sensing image and the local road network grid result into a backbone network for feature extraction, perform feature fusion on the obtained mask feature and RGB feature, input the obtained fusion feature into a query extraction module for query extraction to obtain corresponding local road queries, and complete the road network in the primary road network according to the local road queries to obtain the final road network of the ground of the road network to be extracted.
2. The global-local decoding-based remote sensing image road network extraction method according to claim 1, wherein The calling a Transformer encoder to encode the remote sensing image slices to obtain an initial road query includes: Extract multi-scale feature maps of each remote sensing image slice through a Transformer encoder; Perform multi-scale dilated attention calculation on the multi-scale feature maps according to ground sampling points to obtain corresponding sampling point queries; Select multiple sampling point queries with the largest confidence as the initial road query.
3. The global-local decoding-based method for extracting road networks from remote sensing images according to claim 1, wherein, The constructing noise samples according to the ground truth coordinates of the ground of the road network to be extracted includes: For each ground truth coordinate in the ground of the road network to be extracted, add two random noise values respectively; When the random noise value is less than or equal to a preset noise threshold, use the ground truth coordinate added with the random noise value as a positive sample; When the random noise value is greater than the noise threshold and less than twice the noise threshold, use the ground truth coordinate added with the random noise value as a negative sample; Use the positive samples and negative samples of each ground truth coordinate as noise samples.
4. The global-local decoding method for extracting road networks from remote sensing images according to claim 1, wherein, Inputting the combined feature constructed using the 2D coordinates and the 36D direction descriptor into a topological connection network to obtain a primary road network of the ground of the road network to be extracted, including: Mapping the 2D coordinates into a coordinate feature representation through a coordinate multi-layer perceptron, and mapping the 36D direction descriptor into a direction feature representation through a direction multi-layer perceptron; Performing a splicing process on the coordinate feature representation and the direction feature representation to obtain a combined feature; In the topological connection network, for each target road node, determining adjacent road nodes of the target road node, and merging the combined features of the target road node and the adjacent road nodes to obtain a merged feature; Invoking an activation function to perform activation processing on the merged feature to obtain a corresponding feature tensor; Invoking a fully connected layer to perform mapping processing on the feature tensor to obtain a corresponding tensor value; When the tensor value is greater than a tensor threshold, connecting the target road node and the adjacent road nodes to obtain a primary road network of the ground of the road network to be extracted.
5. The global-local decoding-based method for extracting road networks from remote sensing images according to claim 1, wherein The preset iterative strategy includes: In the first round of iteration, arbitrarily select a road endpoint from the initial retrieval point set as the starting center point; If the number of road nodes in the local road query of the starting center point is 0, randomly select a road endpoint from the candidate point set as the center point for the next round of iteration, where the candidate point set is a point set composed of the remaining road endpoints in the initial retrieval point set except the starting center point; If the number of road nodes in the local road query is 1, use the road node in the local road query as the center point for the next round of iteration; If the number of road nodes in the local road query is 2, randomly select a road node in the local road query as the center point for the next round of iteration, and add the remaining road nodes in the local road query to the candidate point set; When the candidate point set is an empty set, end the iteration.
6. The global-local decoding remote sensing image road network extraction method according to claim 1, wherein The method further includes: Integrating the Transformer encoder, the Transformer decoder, and the topological connection network into a road network extraction model, and the training process of the road network extraction model includes: Obtaining a road network remote sensing image sample, where the road network remote sensing image sample includes a road network connection ground truth; Inputting the road network remote sensing image sample into the road network extraction model for forward propagation; In the parallel stage of forward propagation, constructing a parallel stage loss function according to the road network connection ground truth and the coordinate prediction value, direction prediction value, and node connection prediction value of the road nodes output by the road network extraction model; In the iterative stage of forward propagation, constructing an iterative stage loss function according to the road network connection ground truth and the predicted probability value of the road endpoints output by the road network extraction model; Taking the sum of the parallel stage loss function and the iterative stage loss function as the final loss function, and performing backpropagation in the road network extraction model through the final loss function to update the parameters of the road network extraction model.
7. The global-local decoding remote sensing image road network extraction method according to claim 6, characterized in that The construction process of the road network connection ground truth includes: In the ground truth road map of the road network remote sensing image sample, rasterize it according to line segments of a preset pixel length; When a road node in the road network remote sensing image sample is within the range of the line segment, determine the road node as a valid node; For each valid node, determine the corresponding projection point in the ground truth road map, where the projection point is the center line point in the ground truth road map with the minimum Euclidean distance from the valid node; According to the connection relationship of the projection points, determine the connectivity relationship of the road nodes in the road network remote sensing image sample, and connect the corresponding road nodes in the road network remote sensing image sample based on the connectivity relationship to obtain the road network connection ground truth.
8. The global-local decoding-based remote sensing image road network extraction method according to claim 6, wherein, The construction process of the parallel stage loss function includes: Determine the true values of the coordinates, true values of the directions, and true values of the node connections of the road nodes from the road network connection ground truth; Construct a position loss according to the true value of the coordinates and the predicted value of the coordinates; Construct a direction loss according to the true value of the direction and the predicted value of the direction; Construct a connection loss according to the true value of the node connection and the predicted value of the node connection; Weight the position loss, the direction loss, and the connection loss respectively according to the preset first position weight, direction weight, and connection weight, and sum the obtained weighted results to obtain the parallel stage loss function, where the values of the direction weight and the connection weight are positively correlated with the number of training rounds of the road network extraction model; The construction process of the iterative stage loss function includes: Determine the true probability value of the road endpoints from the road network connection ground truth; Construct a probability loss of the road endpoints according to the true probability value and the predicted probability value; Weight the position loss and the probability loss respectively through the preset second position weight and probability weight, and sum the obtained weighted results to obtain the iterative stage loss function.
9. A global-local decoding remote sensing image road network extraction device, characterized in that, Include: An encoding module for obtaining a remote sensing image slice of the road network ground to be extracted and calling a Transformer encoder to perform encoding processing on the remote sensing image slice to obtain an initial road query; A construction module for constructing a noise sample according to the ground truth coordinates of the road network ground to be extracted; A decoding module for inputting the initial road query and the noise sample into a Transformer decoder for decoding processing to obtain a candidate road query that has completed information interaction, and calling the Hungarian algorithm to perform bipartite matching on the candidate road query to obtain a target road query; A global query module for calling a global prediction head to map the target road query into a 2D coordinate and a 36D direction descriptor of the road nodes respectively, and inputting the combined feature constructed by the 2D coordinate and the 36D direction descriptor into a topological connection network to obtain a primary road network of the road network ground to be extracted; A local query module for completing the road network of the road endpoints during the iterative retrieval of the road endpoints in the primary road network to obtain the final road network of the road network ground to be extracted; During the iterative retrieval of the road endpoints in the primary road network, the road network is complemented for the road endpoints to obtain the final road network of the ground to be extracted, including: Construct an initial retrieval point set based on the road endpoints in the primary road network, where the road endpoints are road nodes with the number of adjacent road nodes less than 2 in the primary road network; Select a center point from the initial retrieval point set according to a preset iterative strategy; For each selected center point, extract a local remote sensing image from the remote sensing image slice, and extract a local road network raster result from the primary road network; Input both the local remote sensing image and the local road network raster result into the backbone network for feature extraction, perform feature fusion on the obtained mask feature and RGB feature, input the obtained fusion feature into the query extraction module for query extraction to obtain the corresponding local road query, and perform road network complementation in the primary road network according to the local road query to obtain the final road network of the ground to be extracted.
Citation Information
Patent Citations
Remote sensing image road extraction method and system based on depth learning, storage medium, and electronic device
CN109493320A
Remote sensing image road extraction method and system combining semantic segmentation and angle prediction
CN115457379A