Remote sensing image road network intelligent extraction method combining dynamic reasoning and spatial topology
Through the combined dynamic inference and spatial topology, the deformation convolution and structural reparameterization technology are used to optimize the road extraction process, and the problem of incomplete road network extraction in remote sensing images is solved, and high-precision and coherent road network extraction is achieved.
Patent Information
- Application Number
- CN202510575174.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art is difficult to effectively extract road networks in remote sensing images, especially in the problem of discontinuity of complex morphology and topological structures, resulting in incomplete road extraction and insufficient multi-task synergy.
Using the combined dynamic inference and spatial topology method, the convolution kernel position is dynamically adjusted through deformation convolution and structural reparameterization technology, and feature extraction and recovery are performed in combination with encoder and decoder, road mask diagram and centerline are generated, and the road extraction process is optimized.
It improves the accuracy and coherence of road extraction, enhances the model's perception of complex road patterns, and realizes efficient road network extraction and update.
Smart Images

Figure CN120471935A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent interpretation of remote sensing images, and specifically relates to an intelligent extraction method of road networks from remote sensing images by combining dynamic reasoning and spatial topology. Background Art
[0002] The road network is one of the most important artificial linear objects on the Earth's surface and serves as infrastructure in multiple fields, including transportation, urban planning, and regional development. High-precision, timely road network information is of great significance for urban management, emergency response, and the updating of geographic information systems. Traditional road network updates rely mainly on manual field surveys or remote sensing image interpretation based on manual annotation. Such methods are not only time-consuming and labor-intensive, but also fail to meet the actual response needs when faced with road changes caused by rapid urban expansion and natural disasters. The wide-area coverage and multi-temporal characteristics of remote sensing images provide new technical means for the rapid acquisition and updating of road networks. However, road extraction and updating still face many challenges.
[0003] Currently, deep learning-based remote sensing road extraction methods demonstrate significant advantages in end-to-end modeling and feature extraction, effectively extracting road structure in high-contrast areas such as main roads. However, because roads in remote sensing imagery often exhibit complex patterns such as breaks, occlusions, and blurring, deep learning models are prone to missed detections, false detections, or topological discontinuities, particularly in urban-rural fringe areas and secondary road areas. Some studies have introduced graph neural networks or attention mechanisms to enhance the model's understanding of spatial structure, but these approaches still face bottlenecks such as poor robustness and inaccurate boundary localization.
[0004] On the other hand, road networks possess a natural spatial topological structure, encompassing important properties such as connectivity and directionality. Effectively incorporating this prior structural knowledge into the automated interpretation of remote sensing imagery is key to improving the accuracy and consistency of road network updates. The current lack of intelligent methods that can dynamically perceive road morphology while simultaneously modeling spatial topological constraints has hindered the efficient extraction of road networks from remote sensing imagery. Summary of the Invention
[0005] In order to solve the technical problems existing in the background technology, the present invention aims to provide an intelligent extraction method of road networks from remote sensing images that combines dynamic reasoning and spatial topology. Through a single network architecture, two types of road products, road surface and road centerline, are simultaneously obtained, solving the problems of incomplete road morphology information extraction and insufficient multi-task collaboration in the existing technology.
[0006] In order to solve the technical problem, the technical solution of the present invention is:
[0007] A method for intelligently extracting road networks from remote sensing images by combining dynamic reasoning and spatial topology, the method comprising:
[0008] S1: The original image is processed through cropping, deformation convolution, structural reparameterization, region and line stages, and the binary road mask M and the optimized road centerline are output.
[0009] S2: Based on the binary road mask obtained in step S1 and the optimized road centerline, the cropped original image is passed to the encoder for feature extraction, and the image features are gradually restored through the decoder to finally generate a binary probability map of the road area.
[0010] The role of the optimized road centerline:
[0011] Supporting information:
[0012] Although not directly used as input, the optimized road centerline can be used to guide the processing of feature maps in the subsequent feature extraction or decoding process, helping the model focus on areas that may contain road information.
[0013] The information of linear features can provide the system with additional context about the road structure, especially when performing subsequent path prediction or segmentation of complex road structures.
[0014] Post-processing steps:
[0015] At the decoder stage or output stage, the optimized road centerlines can be used for additional correction and optimization to ensure that the extracted roads have stronger connectivity with the centerlines.
[0016] The role of the binary road mask (M):
[0017] Training data:
[0018] The binary road mask image can be used for network training and verification, and as annotated true labels, it can provide an important reference, especially when monitoring the accuracy and effect of road extraction.
[0019] During the iterative training process, the feature maps output by the encoder are compared to help the network learn more effective feature representations.
[0020] Feature Guide:
[0021] The mask M indicates the precise location of roads in the image, helping the network focus on these specific areas during layer-by-layer restoration, thereby improving the overall extraction accuracy.
[0022] In the decoder stage, the position information of the mask image can be used to enhance the accuracy of decoding and reconstruction, allowing the network to better restore structurally complete and coherent road areas.
[0023] As you can see, the optimized road centerlines and binarized road masks not only provide additional context and guidance for feature extraction but also optimize the effectiveness of road extraction during training and evaluation. This information can effectively improve model performance, resulting in more accurate and consistent road extraction.
[0024] Furthermore, the step S1 includes:
[0025] S101: Cropping the original image to generate image blocks suitable for network input;
[0026] S102: Input the cropped image block, apply deformable convolution, dynamically adjust the convolution kernel position to enhance the perception of road shape, edge, and curvature features, and output the feature map generated by the deformable convolution;
[0027] S103: Based on the feature map output by the deformable convolution, the inference process is optimized through structural reparameterization, multiple convolution operations are merged into a single efficient operation, the inference process is accelerated, and an efficient feature map after structural reparameterization is obtained;
[0028] S104: In the regional stage, the feature map after the structural reparameterization is processed to generate a binary road mask map M. In the online stage, the road mask map M is processed to generate a preliminary road centerline CL. init The preliminary road centerline is input into the DS-Net network for optimization to obtain the final optimized road centerline.
[0029] Furthermore, the step S2 includes:
[0030] S201: Based on the cropped original image, the encoder performs step-by-step feature extraction on the image information. The encoder consists of five progressive convolutional modules. The first layer uses deformable convolution to dynamically adapt to the complex morphology in the input feature map. The middle three convolutional modules use standard convolution blocks to gradually extract deeper features and semantic information. The last convolutional module introduces structural reparameterization, which can combine the multiple convolution operations used in the training phase into a single efficient operation during the inference phase. Through convolution and pooling operations, high-level features are gradually extracted to obtain feature maps (I1, I2, ..., I5) extracted layer by layer.
[0031] S202: Based on the upsampled features and the feature map output by the encoder, the upsampled features are fused with the feature map of the corresponding layer in the encoder in the decoder to finally generate the fused feature map Z i , after the decoding process, the final road area probability map is obtained That is, the feature map is restored step by step through the decoder to complete the end-to-end recognition process.
[0032] Furthermore, the step S102 of dynamically capturing road morphology information through deformation convolution specifically includes:
[0033] This step uses deformable convolution at the first level of the encoder to adjust the convolution kernel sampling position through dynamic offset to enhance the ability to capture road edge and curvature morphological features;
[0034] The entire process of ordinary two-dimensional convolution is divided into two steps:
[0035] First, the input feature map I is sampled through a regular grid R, and then the sampled weights w are accumulated, where the size and expansion of the convolution kernel are defined by the grid R. Assuming that the grid R is defined as a convolution kernel of size 3×3 and a dilation degree of 1, the grid R is expressed as R = {(-1,-1), (-1,0), ..., (0,1), (1,1)}. For each position P on the output feature map O, i :
[0036]
[0037] Among them, P i Represents the current position coordinates on the output feature map, P n represents the nth offset position in R, w(·) represents the weight of a specific position, I(·) represents the value of the input feature map at a specific position, and O(·) represents the value of the output feature map at a specific position;
[0038] In deformable convolution, the regular grid R will introduce a position offset {Δp n |n=1,...,N}, where N is |R|;
[0039] p' n =p n +Δp n (2)
[0040]
[0041] After that, the irregular position p' after adding the bias n Sampling is performed on the n It is usually a fraction, so Formula 3 needs to be improved by bilinear interpolation:
[0042]
[0043] Among them, q represents the enumeration of all integer spatial positions in the input feature map, that is, the actual coordinates of the pixels; p' n Represents the coordinates of any position, that is, (p n +Δp n), G(.,.) represents a two-dimensional bilinear interpolation kernel, which can be divided into two one-dimensional kernels to calculate the fractional coordinate p' according to the pixel values of the surrounding integer coordinates q. n The value at:
[0044] G(q,p' n )=g(q x ,p' nx )·g(q y ,p' ny ) (5)
[0045] Among them, g(a,b)=max(0,1-|ab|) represents the linear interpolation weight, p' nx 、p' ny Represents p' n The horizontal and vertical coordinates, q x ,q y They represent the horizontal and vertical coordinates of q respectively. G(q,p) is non-zero only for a few q, so the calculation speed is very fast.
[0046] Assuming that the input feature map is I, the deformable convolution kernel is W, and the output feature map is O, the mathematical expression of deformable convolution can be expressed as:
[0047]
[0048] Among them, p represents any position in the output feature map O, K represents the size of the deformed convolution kernel, and W k is the weight parameter of the i-th sampling position, Δp k is the k-th position dynamic offset, I(p+p k +Δp k ) represents the pixel value at the fractional position calculated by bilinear interpolation.
[0049] Furthermore, the step S103 of structural reparameterization to accelerate the model inference process specifically includes:
[0050] This step optimizes the reasoning process of the network model by structural reparameterization. In the convolutional neural network, the convolution operation convolves the input image I with the convolution kernel K and outputs the value O of the zth channel at (x, y). Z (x,y); the convolution operation is expressed as:
[0051]
[0052] Among them, I(x,y) is the pixel value of the input image at (x,y), K Z (i, j) is the weight value of the convolution kernel in the zth channel, b Z is the bias term of the convolution kernel in the z channel, H and W represent the height and width of the convolution kernel respectively;
[0053] Based on the principle of convolution calculation, the convolution operation is a linear operation with distribution and additive properties. For convolutions of the same structure, the distribution is reflected in the decomposition of the convolution of two input images into the sum of separate convolutions:
[0054] Conv(I1+I2,K)=Conv(I1,K)+Conv(I2,K) (8)
[0055]
[0056] If the convolution kernel is decomposed into two parts, one is the 1×1 convolution kernel K1 and the other is the H×W convolution kernel K2, then the re-parameterized convolution is expressed as:
[0057]
[0058] Among them, K1(k,z) represents the weight of the 1×1 convolution kernel, K2(i,j) represents the weight of the H×W convolution kernel, and C represents the number of input channels.
[0059] Furthermore, the step S104 is a surface-to-line intelligent recognition mechanism, specifically including:
[0060] The recognition process is divided into a regional phase and a line phase. The regional phase extracts road surfaces, while the line phase learns strategies for improving the connectivity of road centerlines. In the online phase, the DS-Net road region extraction results are further processed to generate corresponding vectorized road centerlines. First, a post-processing algorithm converts the raster segmentation results into vector centerlines, and the results are converted into a road centerline raster map. Subsequently, the DS-Net network learns a strategy for stitching road centerlines to improve their connectivity.
[0061] In the regional stage, the input remote sensing image I is first segmented at the pixel level by DS-Net to obtain a binary road mask M:
[0062] M=DS-Net(I), M∈{0,1} H×W (11)
[0063] Among them, M{x,y}=1 means that the pixel point {x,y} is determined to be a road area; DS-Net(·) means it is processed by the DS-Net network; then, based on the outer contour boundary of the M extractor
[0064]
[0065] In the online stage, the skeleton extraction algorithm is used to extract the centerline of the road outer contour boundary to obtain the centerline candidate map CLinit :
[0066]
[0067] Input the initial centerline into DS-Net to obtain the final optimized road centerline
[0068] Furthermore, the step S201 specifically includes:
[0069] The DS-Net encoder mechanism consists of five progressive convolutional modules (Conv1-Conv5). Each layer combines the MaxPooling operation to achieve five-fold feature downsampling. The first convolutional module (Conv1: Conv Block-Conv Block3-Conv Block) introduces Deformable Convolution, which dynamically adjusts the convolution kernel shape by learning the convolution kernel offset, thereby enhancing the ability to capture complex morphological features.
[0070]
[0071] in, Indicates that deformable convolution is applied to the first layer of the encoder; BN(·) represents the batch normalization operation, which normalizes the distribution of feature maps and accelerates training convergence; ReLU(·) represents the linear activation function, and ReLU(x)=max(0,x);
[0072] The three-layer convolutional modules Conv2-Conv4 in the middle use the standard convolutional block Conv Block: (Conv-BN-Relu)*2, which successively obtain deeper semantic information in the remote sensing image and gradually improve the abstraction level of the features in the process of step-by-step downsampling;
[0073] Conv2~Conv4:I i =ReLU(BN(Conv(ReLU(BN(Conv(I i -1))))))→fori=2,3,4(16)
[0074] Conv i×i (·) indicates that a convolution operation is performed using a convolution kernel of size i×i; Conv(·) indicates a standard two-dimensional convolution operation;
[0075] The last convolution module Conv5 introduces structural re-parameterization to form a bridge layer. This layer can not only achieve efficient fusion of multi-branch features in the training phase and enhance the network's comprehensive ability to handle different scales and semantic information, but also reduce the computational burden of the inference phase by converting multiple convolution operations into a single efficient operation. If the input image is I0∈R H×W×3 , where H and W represent the height and width of the input image respectively, 3 represents the number of RGB channels, and the output of each layer uses I i If , the processing of the encoder stage is formulated as follows:
[0076] Conv5:I5=Conv reparam (I4) (17)
[0077] Among them, Conv reparam (·) The efficient single-branch convolution operation after structural reparameterization is equivalent to the merging of multi-branch structures in the training phase.
[0078] Furthermore, the step S202 specifically includes:
[0079] The decoder mechanism adopts a hierarchical structure symmetrical to the encoder, which consists of multiple upsampling blocks and feature map convolution layers. i First, the size is gradually restored through the upsampling block (Upsample-3×3Conv-BN-ReLU) At the same time, reducing the number of channels to obtain
[0080]
[0081] Then, the upsampling results are combined with the feature fusion mechanism Concat and the feature map I of the corresponding layer in the encoder L-i Perform fusion, keeping the feature map size unchanged but doubling the number of channels;
[0082]
[0083] The fused feature map Z i Then, the convolution module (3×3Conv—BN—ReLU) is used to restore the number of channels to the state before the fusion operation. i-1 , to avoid channel redundancy;
[0084] I i-1 =ReLU(BN(Conv 3×3 (Zi ))) (twenty one)
[0085] Among them, Upsample represents the upsampling operation; Concat represents the feature map concatenation operation, which often concatenates two feature maps in the channel dimension. The final predicted output is:
[0086]
[0087] Among them, δ(·) is the Sigmoid activation function, which is used to generate the binary road area probability map.
[0088] A remote sensing image road network intelligent extraction system combining dynamic reasoning and spatial topology, the system comprising:
[0089] one or more processors;
[0090] a memory for storing one or more programs;
[0091] When the one or more programs are executed by the one or more processors, the one or more processors execute any of the above-mentioned methods for intelligent extraction of road networks from remote sensing images by combining dynamic reasoning and spatial topology.
[0092] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements any one of the above-mentioned methods for intelligently extracting road networks from remote sensing images by combining dynamic reasoning and spatial topology.
[0093] Compared with the prior art, the advantages of the present invention are:
[0094] First, through the innovative dynamic training module (DTM-DC, Dynamic Training Module with Deformable Convolution) and dynamic reasoning module (DIM-SR, Dynamic Inference Module with Structural Re-parameterization), the joint optimization of the morphological connectivity of road boundaries and centerlines is achieved; secondly, a two-stage collaborative training framework is constructed for the first time to optimize the extraction performance of road surfaces and centerlines; finally, the model demonstrates strong generalization ability in cross-regional tests, revealing its practical value in different geographical environments, and providing an efficient dual-output solution for intelligent transportation systems and urban planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 , the main flow chart of the intelligent extraction method of road network from remote sensing images combining dynamic reasoning and spatial topology of the present invention;
[0096] Figure 2 , network structure diagram;
[0097] Figure 3 , model generalization diagram. DETAILED DESCRIPTION
[0098] The specific implementation of the present invention is described below in conjunction with embodiments:
[0099] It should be noted that the structures, proportions, sizes, etc. shown in this specification are only used to match the contents disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Any structural modification, change in proportional relationship or adjustment of size should still fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention.
[0100] At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" quoted in this specification are only for the convenience of description and are not used to limit the scope of implementation of the present invention. Changes or adjustments to their relative relationships should be regarded as the scope of implementation of the present invention without substantially changing the technical content.
[0101] Example 1:
[0102] like Figure 1 and Figure 2 As shown, the present invention proposes a semantic segmentation network DS-Net based on dynamic strategy optimization, which specifically includes the following steps:
[0103] (1) Data preparation and module construction
[0104] Step S1: Road data construction
[0105] Crop the original remote sensing or drone imagery into multiple image patches to fit the network input.
[0106] Step S2: Deformable convolution dynamically captures road morphology information.
[0107] This step uses deformable convolution in the first layer of the encoder, and adjusts the convolution kernel sampling position through dynamic offset to enhance the capture of morphological features such as road edges and curvature.
[0108] The principle of conventional two-dimensional convolution is simple, and the whole process is divided into two steps. First, the input feature map I is sampled through a regular grid R, and then the sampled weights w are accumulated. Among them, the size and expansion of the convolution kernel can be defined by the grid R. Assuming that the grid R is defined as a convolution kernel of size 3×3 and a dilation of 1, the grid R can be expressed as R = {(-1,-1), (-1,0), ..., (0,1), (1,1)}. For each position P on the output feature map Oi :
[0109]
[0110] Among them, P i Represents the current position coordinates on the output feature map, P n represents the nth offset position in R, w(·) represents the weight of a specific position, I(·) represents the value of the input feature map at a specific position, and O(·) represents the value of the output feature map at a specific position.
[0111] In deformable convolution, the regular grid R will introduce a position offset {Δp n |n=1,...,N}, where N is |R|.
[0112] p' n =p n +Δp n (2)
[0113]
[0114] After that, the irregular position p' after adding the bias n Sampling is performed on the n It is usually a fraction, so Formula 3 needs to be improved by bilinear interpolation:
[0115]
[0116] Where q represents the enumeration of all integer spatial positions in the input feature map (i.e., the actual coordinates of the pixels); p' n Represents the coordinates of any position, that is, (p n +Δp n ). G(.,.) represents a two-dimensional bilinear interpolation kernel, which can be divided into two one-dimensional kernels to calculate the fractional coordinate p' according to the pixel values of the surrounding integer coordinates q n The value at:
[0117] G(q,p' n )=g(q x ,p' nx )·g(q y ,p' ny ) (5)
[0118] Among them, g(a,b)=max(0,1-|ab|) represents the linear interpolation weight, p' nx 、p' ny Represents p' n The horizontal and vertical coordinates, q x ,q ywhere q is the horizontal and vertical coordinates of q respectively. G(q,p) is non-zero only for a small number of q, so the calculation speed is very fast.
[0119] Assuming that the input feature map is I, the deformable convolution kernel is W, and the output feature map is O, the mathematical expression of deformable convolution can be expressed as:
[0120]
[0121] Among them, p represents any position in the output feature map O, K represents the size of the deformed convolution kernel, and W k is the weight parameter of the i-th sampling position, Δp k is the k-th position dynamic offset, I(p+p k +Δp k ) represents the pixel value at the fractional position calculated by bilinear interpolation.
[0122] Step S3: Structural reparameterization accelerates the model inference process.
[0123] This step optimizes the reasoning process of the network model by means of structural reparameterization. In a convolutional neural network, the convolution operation convolves the input image I with the convolution kernel K and outputs O Z The value of the zth channel of (x,y) at (x,y). The convolution operation can be expressed as:
[0124]
[0125] Among them, I(x,y) is the pixel value of the input image at (x,y), K Z (i, j) is the weight value of the convolution kernel in the zth channel, b Z is the bias term of the convolution kernel in the z channel. H and W represent the height and width of the convolution kernel respectively.
[0126] Based on the principle of convolution calculation, the convolution operation is a linear operation with distribution and additive properties. For convolutions of the same structure, the distribution is reflected in the fact that the convolution of two input images can be decomposed into the sum of individual convolutions:
[0127] Conv(I1+I2,K)=Conv(I1,K)+Conv(I2,K) (8)
[0128]
[0129] If the convolution kernel is decomposed into two parts, one is the 1×1 convolution kernel K1 and the other is the H×W convolution kernel K2. The reparameterized convolution can be expressed as:
[0130]
[0131] Among them, K1(k,z) represents the weight of the 1×1 convolution kernel, K2(i,j) represents the weight of the H×W convolution kernel, and C represents the number of input channels.
[0132] Step S4: Intelligent recognition mechanism from surface to line.
[0133] The recognition process is divided into two phases: the regional phase extracts road surfaces, while the line phase learns strategies for improving road centerline connectivity. In the online phase, the DS-Net road region extraction results are further processed to generate corresponding vectorized road centerlines. First, a post-processing algorithm converts the raster segmentation results into vector centerlines, which are then converted into a road centerline raster map. Subsequently, the DS-Net network learns a strategy for stitching road centerlines to improve their connectivity.
[0134] In the regional stage, the neural network model DS-Net is first used to perform pixel-level segmentation on the input remote sensing image I to obtain a binary road mask M:
[0135] M=DS-Net(I), M∈{0,1} H×W (11)
[0136] Among them, M{x,y}=1 means that the pixel {x,y} is determined to be a road area; DS-Net(·) means it is processed by the DS-Net network. Then, based on the outer contour boundary of M extractor
[0137]
[0138] In the online stage, the skeleton extraction algorithm is used to extract the centerline of the road outer contour boundary to obtain the centerline candidate map CL init :
[0139]
[0140] Input the initial centerline into the DS-Net network to obtain the final optimized road centerline
[0141]
[0142] (2) Road extraction network framework based on dynamic reasoning and spatial topology
[0143] Step S1: extract features of the image information step by step through the encoder.
[0144] The DS-Net encoder mechanism consists of five progressive convolutional modules (Conv1-Conv5), each layer combining a max-pooling operation to achieve five-fold feature downsampling. In particular, the first convolutional module (Conv1: ConvBlock-Conv Block 3-Conv Block) introduces deformable convolution, which dynamically adjusts the convolution kernel shape by learning its offset, thereby enhancing the ability to capture complex morphological features. Compared to ordinary convolution, deformable convolution has stronger local adaptability.
[0145]
[0146] in, Indicates that deformable convolution is applied to the first layer of the encoder; BN(·) represents the batch normalization operation, which normalizes the feature map distribution and accelerates training convergence; ReLU(·) represents the linear activation function, and ReLU(x)=max(0,x).
[0147] The three middle convolutional modules (Conv2-Conv4) use standard convolutional blocks (Conv Block: (Conv-BN-Relu)*2), which successively obtain deeper semantic information in the remote sensing image and gradually improve the level of feature abstraction in the process of gradual downsampling.
[0148] Conv2~Conv4:I i =ReLU(BN(Conv(ReLU(BN(Conv(I i -1))))))→fori=2,3,4(16)
[0149] Conv i×i (·) indicates that a convolution operation is performed using a convolution kernel of size i×i; Conv(·) indicates a standard two-dimensional convolution operation; BN(·) and ReLU(·) have the same functions as above.
[0150] The last convolution module (Conv5) introduces the structural re-parameterization technology to form a bridge layer. This layer can not only achieve efficient fusion of multi-branch features in the training phase and enhance the network's ability to integrate information of different scales and semantics, but also reduce the computational burden of the inference phase by converting multiple convolution operations into a single efficient operation. H×W×3 (where H and W represent the height and width of the input image, respectively, and 3 represents the number of RGB channels). The output of each layer uses I iIf , the processing of the encoder stage is formulated as follows:
[0151] Conv5:I5=Conv reparam (I4) (17)
[0152] Among them, Conv reparam (·) The efficient single-branch convolution operation after structural reparameterization is equivalent to the merging of multi-branch structures in the training phase.
[0153] Step S2: The decoder restores the feature map step by step to complete the end-to-end recognition process.
[0154] The decoder mechanism adopts a hierarchical structure symmetrical to the encoder, consisting of multiple upsampling blocks (UpsamplingBlock) and feature map convolution layers (Convolutional Layer). The feature map I after encoder processing i First, the size is gradually restored through the upsampling block (Upsample-3×3Conv-BN-ReLU) At the same time, reducing the number of channels to obtain
[0155]
[0156] Then, the upsampling results are combined with the feature fusion mechanism (Concat) and the feature map I of the corresponding layer in the encoder L-i Perform fusion, keeping the feature map size unchanged but doubling the number of channels.
[0157]
[0158] The fused feature map Z i Then, the convolution module (3×3Conv—BN—ReLU) is used to restore the number of channels to the state before the fusion operation. i-1 , to avoid channel redundancy.
[0159] I i-1 =ReLU(BN(Conv 3×3 (Z i ))) (twenty one)
[0160] Among them, Upsample represents the upsampling operation; Concat represents the feature map concatenation operation, which often concatenates two feature maps in the channel dimension. The final prediction output is:
[0161]
[0162] Among them, δ(·) is the Sigmoid activation function, which is used to generate the binary road area probability map.
[0163] This paper uses public data from the Massachusetts road dataset and the Ottawa dataset to test the generalization evaluation results of the remote sensing image road network intelligent extraction method that combines dynamic reasoning and spatial topology under actual conditions. Figure 3 .
[0164] Example 2:
[0165] This embodiment provides a terminal device, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of a remote sensing image road network intelligent extraction method that combines dynamic reasoning and spatial topology, including the following steps:
[0166] S1: The original image is processed through cropping, deformation convolution, structural reparameterization, region and line stages, and the binary road mask M and the optimized road centerline are output.
[0167] S2: Based on the binary road mask obtained in step S1 and the optimized road centerline, the cropped original image is passed to the encoder for feature extraction, and the image features are gradually restored through the decoder to finally generate a binary probability map of the road area.
[0168] Example 3:
[0169] This embodiment provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0170] The processor may load and execute one or more instructions stored in a computer-readable storage medium to implement the corresponding steps of the above-mentioned embodiment regarding a method for intelligently extracting a road network from a remote sensing image by combining dynamic reasoning and spatial topology. The processor may load and execute the following steps:
[0171] S1: The original image is processed through cropping, deformation convolution, structural reparameterization, region and line stages, and the binary road mask M and the optimized road centerline are output.
[0172] S2: Based on the binary road mask obtained in step S1 and the optimized road centerline, the cropped original image is passed to the encoder for feature extraction, and the image features are gradually restored through the decoder to finally generate a binary probability map of the road area.
[0173] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0174] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0175] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0177] The preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
[0178] Many other changes and modifications can be made without departing from the spirit and scope of the present invention. It should be understood that the present invention is not limited to the specific embodiments, and the scope of the present invention is defined by the appended claims.
Claims
1. A method for intelligent road network extraction from remote sensing images combining dynamic reasoning and spatial topology, characterized in that: The method comprises: S1: The original image is processed through cropping, deformation convolution, structural reparameterization, region and line stages, and the binary road mask M and the optimized road centerline are output. S2: Based on the binary road mask obtained in step S1 and the optimized road centerline, the cropped original image is passed to the encoder for feature extraction, and the image features are gradually restored through the decoder to finally generate a binary probability map of the road area.
2. The intelligent road network extraction method based on remote sensing images combining dynamic reasoning and spatial topology according to claim 1 is characterized in that: The step S1 comprises: S101: Cropping the original image to generate image blocks suitable for network input; S102: Input the cropped image block, apply deformable convolution, dynamically adjust the convolution kernel position to enhance the perception of road shape, edge, and curvature features, and output a feature map generated by the deformable convolution; S103: Based on the feature map output by the deformable convolution, the inference process is optimized through structural reparameterization, multiple convolution operations are merged into a single efficient operation, the inference process is accelerated, and an efficient feature map after structural reparameterization is obtained; S104: In the regional stage, the feature map after the structural reparameterization is processed to generate a binary road mask map M. In the online stage, the road mask map M is processed to generate a preliminary road centerline CL. init The preliminary road centerline is input into the DS-Net network for optimization to obtain the final optimized road centerline.
3. The intelligent road network extraction method based on remote sensing images combining dynamic reasoning and spatial topology according to claim 2 is characterized in that: The step S2 includes: S201: Based on the cropped original image, the encoder performs step-by-step feature extraction on the image information. The encoder consists of five progressive convolutional modules. The first layer uses deformable convolution to dynamically adapt to the complex morphology in the input feature map. The middle three convolutional modules use standard convolution blocks to gradually extract deeper features and semantic information. The last convolutional module introduces structural reparameterization, which can combine the multiple convolution operations used in the training phase into a single efficient operation during the inference phase. Through convolution and pooling operations, high-level features are gradually extracted to obtain feature maps (I1, I2, ..., I5) extracted layer by layer. S202: Based on the upsampled features and the feature map output by the encoder, the upsampled features are fused with the feature map of the corresponding layer in the encoder in the decoder to finally generate the fused feature map Z i , after the decoding process, the final road area probability map is obtained That is, the feature map is restored step by step through the decoder to complete the end-to-end recognition process.
4. The method for intelligently extracting road networks from remote sensing images by combining dynamic reasoning and spatial topology according to claim 2 is characterized in that: The step S102 of deformable convolution dynamically captures road morphology information, specifically including: This step uses deformable convolution at the first encoder layer, adjusting the convolution kernel sampling position through dynamic offset to enhance the ability to capture road edge and curvature morphological features; The whole process of two-dimensional convolution is divided into two steps: First, the input feature map I is sampled through a regular grid R, and then the sampled weights w are accumulated, where the size and expansion of the convolution kernel are defined by the grid R. Assuming that the grid R is defined as a convolution kernel of size 3×3 and a dilation degree of 1, the grid R is expressed as R = {(-1,-1), (-1,0), ..., (0,1), (1,1)}. For each position P on the output feature map O, i : Among them, P i Represents the current position coordinates on the output feature map, P n represents the nth offset position in R, w(·) represents the weight of a specific position, I(·) represents the value of the input feature map at a specific position, and O(·) represents the value of the output feature map at a specific position; In deformable convolution, the regular grid R will introduce a position offset {Δp n |n=1,...,N}, where N is |R|; p' n =p n +Δp n (2) After that, the irregular position p' after adding the bias n Sampling is performed on the n It is usually a fraction, so Formula 3 needs to be improved by bilinear interpolation: Among them, q represents the enumeration of all integer spatial positions in the input feature map, that is, the actual coordinates of the pixels; p' n Represents the coordinates of any position, that is, (p n +Δp n ), G(.,.) represents a two-dimensional bilinear interpolation kernel, which can be divided into two one-dimensional kernels to calculate the fractional coordinate p' according to the pixel values of the surrounding integer coordinates q. n The value at: G(q,p' n )=g(q x ,p' nx )·g(q y ,p' ny ) (5) Among them, g(a,b)=max(0,1-|ab|) represents the linear interpolation weight, p' nx 、p' ny Represents p' n The horizontal and vertical coordinates, q x ,q y They represent the horizontal and vertical coordinates of q respectively. G(q,p) is non-zero only for a few q, so the calculation speed is very fast. Assuming that the input feature map is I, the deformable convolution kernel is W, and the output feature map is O, the mathematical expression of deformable convolution can be expressed as: Among them, p represents any position in the output feature map O, K represents the size of the deformed convolution kernel, and W k is the weight parameter of the i-th sampling position, Δp k is the k-th position dynamic offset, I(p+p k +Δp k ) represents the pixel value at the fractional position calculated by bilinear interpolation.
5. The method for intelligently extracting road networks from remote sensing images by combining dynamic reasoning and spatial topology according to claim 2, characterized in that: The step S103 structure reparameterization accelerates the model reasoning process, specifically including: This step optimizes the reasoning process of the network model by structural reparameterization. In the convolutional neural network, the convolution operation convolves the input image I with the convolution kernel K and outputs the value O of the z-th channel at (x, y). Z (x,y); the convolution operation is expressed as: Among them, I(x,y) is the pixel value of the input image at (x,y), K Z (i, j) is the weight value of the convolution kernel in the zth channel, b Z is the bias term of the convolution kernel in the z channel, H and W represent the height and width of the convolution kernel respectively; Based on the principle of convolution calculation, the convolution operation is a linear operation with distribution and additive properties. For convolutions of the same structure, the distribution is reflected in the decomposition of the convolution of two input images into the sum of separate convolutions: Conv(I1+I2,K)=Conv(I1,K)+Conv(I2,K) (8) If the convolution kernel is decomposed into two parts, one is the 1×1 convolution kernel K1 and the other is the H×W convolution kernel K2, then the re-parameterized convolution is expressed as: Among them, K1(k,z) represents the weight of the 1×1 convolution kernel, K2(i,j) represents the weight of the H×W convolution kernel, and C represents the number of input channels.
6. The method for intelligently extracting road networks from remote sensing images by combining dynamic reasoning and spatial topology according to claim 2, characterized in that: The step S104 is a surface-to-line intelligent recognition mechanism, specifically including: The recognition process is divided into a regional phase and a line phase. The regional phase extracts road surfaces, while the line phase learns strategies for improving the connectivity of road centerlines. In the online phase, the DS-Net road region extraction results are further processed to generate corresponding vectorized road centerlines. First, a post-processing algorithm converts the raster segmentation results into vector centerlines, and the results are converted into a road centerline raster map. Subsequently, the DS-Net network learns a strategy for stitching road centerlines to improve their connectivity. In the regional stage, the input remote sensing image I is first segmented at the pixel level by DS-Net to obtain a binary road mask M: M=DS-Net(I),M∈{0,1} H×W (11) Among them, M{x,y}=1 means that the pixel point {x,y} is determined to be a road area; DS-Net(·) means it is processed by the DS-Net network; then, based on the outer contour boundary of the M extractor In the online stage, the skeleton extraction algorithm is used to extract the centerline of the road outer contour boundary to obtain the centerline candidate map CL init : Input the initial centerline into DS-Net to obtain the final optimized road centerline 7. The method for intelligent road network extraction from remote sensing images combining dynamic reasoning and spatial topology according to claim 3 is characterized in that: The step S201 specifically includes: The DS-Net encoder mechanism consists of five progressive convolutional modules (Conv1-Conv5). Each layer combines the MaxPooling operation to achieve five-fold feature downsampling. The first convolutional module (Conv1: Conv Block-Conv Block 3-Conv Block) introduces Deformable Convolution, which dynamically adjusts the convolution kernel shape by learning the convolution kernel offset, thereby enhancing the ability to capture complex morphological features. in, Indicates that deformable convolution is applied to the first layer of the encoder; BN(·) represents the batch normalization operation, which normalizes the distribution of feature maps and accelerates training convergence; ReLU(·) represents the linear activation function, and ReLU(x)=max(0,x); The three-layer convolutional modules Conv2-Conv4 in the middle use the standard convolutional block Conv Block: (Conv-BN-Relu)*2, which successively obtain deeper semantic information in the remote sensing image and gradually improve the abstraction level of the features in the process of step-by-step downsampling; Conv2~Conv4:I i =ReLU(BN(Conv(ReLU(BN(Conv(I i -1))))))→fora=2,3,4 (16) Conv i×i (·) indicates that a convolution operation is performed using a convolution kernel of size i×i; Conv(·) indicates a standard two-dimensional convolution operation; The last convolution module Conv5 introduces structural re-parameterization to form a bridge layer. This layer can not only achieve efficient fusion of multi-branch features in the training phase and enhance the network's comprehensive ability to handle different scales and semantic information, but also reduce the computational burden of the inference phase by converting multiple convolution operations into a single efficient operation. If the input image is I0∈R H×W×3 , where H and W represent the height and width of the input image respectively, 3 represents the number of RGB channels, and the output of each layer uses I i If , the processing of the encoder stage is formulated as follows: Conv5:I5=Conv reparam (I4) (17) Among them, Conv reparam (·) The efficient single-branch convolution operation after structural reparameterization is equivalent to the merging of multi-branch structures in the training phase.
8. The method for intelligently extracting road networks from remote sensing images by combining dynamic reasoning and spatial topology according to claim 3 is characterized in that: The step S202 specifically includes: The decoder mechanism adopts a hierarchical structure symmetrical to the encoder, which consists of multiple upsampling blocks and feature map convolution layers. i First, the size is gradually restored through the upsampling block (Upsample-3×3Conv-BN-ReLU) At the same time, reducing the number of channels to obtain Then, the upsampling results are combined with the feature fusion mechanism Concat and the feature map I of the corresponding layer in the encoder L-i Perform fusion, keeping the feature map size unchanged but doubling the number of channels; The fused feature map Z i Then, the convolution module (3×3Conv—BN—ReLU) is used to restore the number of channels to the state before the fusion operation. i-1 , to avoid channel redundancy; and i-1 =ReLU(BN(Conv 3×3 (Z i ))) (21) Among them, Upsample represents the upsampling operation; Concat represents the feature map concatenation operation, which often concatenates two feature maps in the channel dimension. The final predicted output is: Among them, δ(·) is the Sigmoid activation function, which is used to generate the binary road area probability map.
9. A remote sensing image road network intelligent extraction system combining dynamic reasoning and spatial topology, characterized by: The system comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors execute the remote sensing image road network intelligent extraction method combining dynamic reasoning and spatial topology as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements a remote sensing image road network intelligent extraction method combining dynamic reasoning and spatial topology according to any one of claims 1 to 8.