Road weak information extraction method and system based on ASU-Net

By constructing the ASU-Net model and using dilated convolution to expand the receptive field and strip convolution to capture features, the imbalance and misclassification problems in road extraction in remote sensing images are solved, and the accuracy and connectivity of road extraction are improved.

CN120808157AActive Publication Date: 2025-10-17CHINA AERO GEOPHYSICAL SURVEY & REMOTE SENSING CENT FOR LAND & RESOURCES

Patent Information

Application Number
CN202510938853.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Road extraction from remote sensing images faces problems such as imbalance, misclassification, fragmentation, and loss of small road information, which leads to insufficient extraction accuracy and connectivity, affecting geographic information analysis.

Method used

An ASU-Net model is constructed by fusing strip convolution and dilated convolution. Dilated convolution is used to expand the receptive field, strip convolution is used to capture road features, and skip connections are used to fuse multi-scale features to improve the accuracy of road extraction.

Benefits of technology

A balance between high precision and recall rate of road extraction is achieved in complex environments, which improves the completeness and accuracy of road extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808157A_ABST
    Figure CN120808157A_ABST
Patent Text Reader

Abstract

The invention discloses a road weak information extraction method and system based on ASU-Net, and relates to the technical field of image feature recognition, and the method comprises the steps: employing an encoder-decoder structure to construct a road extraction network model ASU-Net, employing ResNet-50 as an encoder, introducing a cavity convolution module into the encoder, and introducing a strip convolution module into the decoder; inputting a remote sensing satellite road image into an ASU-Net model, performing down-sampling operation on the image by an encoder, and expanding a receptive field through parallel cascade branches with preset voidage in a void convolution module; up-sampling is carried out on features learned by the encoder through transposition convolution, and meanwhile, road linear feature extraction is carried out by adopting banded convolution in four different directions through a banded convolution module. Through the technical scheme of the invention, the performance of road extraction in a complex environment is improved, better balance between the accuracy rate and the recall rate is realized, and the accuracy of road extraction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image feature recognition, and particularly relates to a road weak information extraction method based on an ASU-Net and a road weak information extraction system based on the ASU-Net. BACKGROUND

[0002] The road extraction problem can be regarded as a semantic segmentation task, which needs to classify each pixel to determine whether it belongs to the road class or the background class. However, due to too much interference information from the background class, the road extraction problem of remote sensing images has always faced some major challenges:

[0003] (1) In the remote sensing road image, the road usually only occupies a relatively small proportion, and the remaining part is mostly background, which leads to a serious imbalance between the foreground and the background, making it difficult to train the road extraction model. Because the model generally tends to learn the background class with a larger proportion, and the learning of the road class is relatively insufficient, which will lead to the decline of the performance of the model and the accuracy will also be affected.

[0004] (2) In terms of geometric shape, texture features, and other aspects, the road is very similar to the background class such as rivers, gullies, farmland, etc., so it is easy to misclassify. This will make it difficult for the road extraction algorithm to accurately distinguish the classes of these objects, which may lead to the error of marking rivers, gullies, farmland and other objects as road class. Moreover, this situation will not only affect the accuracy of the road extraction result, but also may have an impact on the subsequent geographic information analysis and application.

[0005] (3) Urban roads usually have complex traffic environments, and vehicles, vegetation, buildings and other objects will also block the road. Based on the above factors, the result of road extraction often has fragmented discontinuous roads

[34] , because the model may interrupt or incorrectly identify the road, so it is very difficult to extract continuous roads. In addition, the complex traffic environment, such as the variety of road shapes or directions, makes it difficult to accurately capture the road.

[0006] (4) Because the small roads are weakly represented in the remote sensing image, it is easy to lose the information of the small roads in the feature extraction process, which leads to the omission of small roads

[31] . Therefore, when extracting the road, special attention should be paid to the identification of small roads, and the corresponding algorithm should be designed to solve this problem, so as to improve the integrity of the road extraction result.

[0007] In summary, the road extraction task of remote sensing images is very challenging, and it is also very difficult to design a method that can effectively solve the above difficulties. Therefore, in the task of road extraction of remote sensing images, many outstanding methods are still needed to further improve the connectivity and accuracy of road extraction, and to provide more comprehensive support for the field of geographic information. SUMMARY

[0008] In view of the above problems, the present application provides a road weak information extraction method and system based on ASU-Net, which constructs a road extraction network model ASU-Net that fuses strip convolution and dilated convolution, uses a strip convolution module to capture image features, reduces irrelevant area interference feature learning, and uses dilated convolution of the dilated convolution module to expand the receptive field, while fusing multi-scale features through skip connection, better capturing global context information and retaining lower-level features, improving the performance of road extraction in complex environments, achieving a better balance between precision and recall, and improving the accuracy of road extraction.

[0009] To achieve the above object, the present application provides a road weak information extraction method based on ASU-Net, comprising:

[0010] An encoder-decoder structure is used to construct a road extraction network model ASU-Net, wherein ResNet-50 is used as the encoder, and a dilated convolution module is introduced in the encoder, and a strip convolution module is introduced in the decoder.

[0011] A remote sensing satellite road image is input into the ASU-Net model, the encoder performs down-sampling operation on the image, and the receptive field is expanded through the parallel cascade branch of the preset dilated rate in the dilated convolution module.

[0012] The features learned by the encoder are up-sampled through transposed convolution, and the strip convolution module is used to extract road linear features through four different direction strip convolutions.

[0013] In the above technical solution, preferably, the encoder and the decoder are connected through a skip connection for multi-scale feature information fusion.

[0014] In the above technical solution, preferably, the ResNet-50 includes a preset number of convolution layers, and uses a bottleneck module as a residual module, and the ResNet-50 is pre-trained on ImageNet.

[0015] In the technical solution, preferably, the output feature of the residual module is the sum of the residual mapping of the input feature after a convolution layer and the input feature, the residual module adopts a 1*1 convolution layer for dimension reduction, a 3*3 convolution layer, and a 1*1 convolution layer for dimension restoration.

[0016] In the technical solution, preferably, the atrous convolution module adopts four parallel branches, wherein the first three branches fuse feature maps of different receptive field sizes in a cascaded manner, and the last branch adds image-level features through global average pooling and reduces the number of channels to the input channel number through convolution operation.

[0017] In the technical solution, preferably, the encoder reduces the image size to a preset ratio through five downsampling operations, and correspondingly increases the receptive field according to the atrous rate of the cascaded branch in the atrous convolution module.

[0018] In the technical solution, preferably, the strip convolution module includes four strip convolutions of horizontal, vertical, left diagonal and right diagonal, the input tensor is sent to four parallel strip convolution branches after 1*1 convolution operation, then the four output tensors are connected and subjected to 1*1 convolution operation, and long-range context information in the feature map is captured from four different directions, so that each pixel in the output feature map can be associated with the pixels in the input feature map in four directions.

[0019] In the technical solution, preferably, the atrous convolution module adjusts the calculation formula of the receptive field of the pixels in the output feature map by adjusting the atrous rate, which is:

[0020] S=(r-1)*(k-1)+k

[0021] Wherein, S is the receptive field, r is the atrous rate, and k is the size of the convolution kernel.

[0022] In the technical solution, preferably, in the atrous convolution module, the receptive field size calculation formula of the cascaded atrous convolution is:

[0023] S'=S1+S2-1

[0024] Wherein, S' is the receptive field size after cascading, S1 and S2 are the receptive field sizes of the two cascaded atrous convolutions respectively.

[0025] The application also provides a road weak information extraction system based on ASU-Net, which applies the road weak information extraction method based on ASU-Net disclosed in any one of the above technical solutions, comprising:

[0026] An extraction model construction module, configured to construct a road extraction network model ASU-Net using an encoder-decoder structure, wherein ResNet-50 is used as the encoder, a dilated convolution module is introduced in the encoder, and a striped convolution module is introduced in the decoder;

[0027] An image sampling and processing module is used to input a remote sensing satellite road image into the ASU-Net model, wherein the encoder performs a downsampling operation on the image and expands the receptive field through parallel cascade branches with a preset dilation rate in the dilated convolution module;

[0028] The road feature extraction module is used to upsample the features learned by the encoder through transposed convolution, and at the same time extract road linear features through the strip convolution module using strip convolutions in four different directions.

[0029] Compared with the prior art, the present invention has the following beneficial effects: by constructing a road extraction network model ASU-Net that integrates strip convolution and dilated convolution, the strip convolution module is used to capture image features, reducing the interference of irrelevant areas in feature learning, and the dilated convolution of the dilated convolution module is used to expand the receptive field. At the same time, multi-scale features are fused through jump connections to better capture global context information and retain lower-level features, thereby improving the performance of road extraction in complex environments, achieving a better balance between precision and recall, and improving the accuracy of road extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of the network structure of the ASU-Net model disclosed in one embodiment of the present invention;

[0031] Figure 2 A schematic diagram of the principle of a residual module disclosed in one embodiment of the present invention;

[0032] Figure 3 This is a schematic structural diagram of a bottleneck module disclosed in one embodiment of the present invention;

[0033] Figure 4 A schematic diagram of the principle of a dilated convolution module disclosed in one embodiment of the present invention;

[0034] Figure 5 A schematic diagram of a dilated convolution process disclosed in an embodiment of the present invention;

[0035] Figure 6 A schematic diagram of a process for increasing the receptive field of a dilated convolution according to an embodiment of the present invention;

[0036] Figure 7 A schematic diagram of the principle of a banded convolution module disclosed in one embodiment of the present invention;

[0037] Figure 8 A schematic diagram of a road image of a Zhouqu road dataset disclosed in an embodiment of the present invention;

[0038] Figure 9 This is a schematic diagram of experimental results comparing different algorithms on the Zhouqu road dataset disclosed in one embodiment of the present invention;

[0039] Figure 10 This is a schematic diagram of ablation experiment results of different models on the Zhouqu road dataset disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0041] The present invention is described in further detail below with reference to the accompanying drawings:

[0042] like Figure 1 As shown, a road weak information extraction method based on ASU-Net provided by the present invention includes:

[0043] The ASU-Net road extraction network model is constructed using an encoder-decoder structure. ResNet-50 is used as the encoder, and a dilated convolution module is introduced in the encoder and a striped convolution module is introduced in the decoder.

[0044] The remote sensing satellite road image is input into the ASU-Net model. The encoder downsamples the image and expands the receptive field through the parallel cascade branches with preset dilation rates in the dilated convolution module.

[0045] The features learned by the encoder are upsampled through transposed convolution, and the road linear features are extracted through the strip convolution module using four strip convolutions in different directions.

[0046] In this embodiment, in the face of the problem that the U-Net model is insufficient in extracting roads, the application constructs a road extraction network model ASU-Net which fuses a strip-shaped convolution and a hollow convolution, adopts a strip-shaped convolution module to capture image features, reduces irrelevant area interference feature learning, expands the receptive field through the hollow convolution of the hollow convolution module, and better captures global context information and retains lower-level features, improves the performance of road extraction in a complex environment, and achieves a better balance between precision and recall, and improves the accuracy of road extraction.

[0047] Specifically, in the ASU-Net model, a hollow convolution module is mainly included to expand the receptive field, and a strip-shaped convolution module is included to align with the shape of the road and extract the linear features of the road. In the encoder part, since ResNet performs well in feature learning, the application adopts ResNet-50 as the encoder.

[0048] In the decoder part, the application up-samples the learned features to the size of the input image through transposed convolution, and extracts the linear features of the road through the strip-shaped convolution module.

[0049] Among them, the hollow convolution, also known as the dilated convolution, is a convolution operation in the convolutional neural network. The initial proposal of the hollow convolution is to solve the image segmentation problem. Common image segmentation algorithms usually use pooling layers to increase the receptive field, but at the same time, the size of the feature map is also reduced. At this time, the size of the image is usually restored by up-sampling, but the process of reducing and enlarging the feature map will cause loss of details. In contrast, the hollow convolution can increase the receptive field of the model without reducing the resolution of the picture. The hollow convolution introduces one or more gaps in the convolution kernel to increase the receptive field, so that the network can capture more extensive context information. Moreover, compared with the ordinary convolution operation, the parameter amount of the hollow convolution remains unchanged, and the model parameter amount is not increased.

[0050] In the above embodiment, preferably, the encoder and the decoder are connected through a jump connection for multi-scale feature information fusion.

[0051] If the decoder only uses the final output feature map of the encoder, it will easily lose pixels and miss small size roads. Therefore, according to the U-Net model, a jump connection is added to realize the low-level information sharing between the encoder and the decoder, so that the low-level features can be retained and multi-scale feature fusion can be realized.

[0052] Through the first four layers of the encoder feature map and the output of the hollow convolution module, the decoder can extract multi-scale context information from the high-level feature map and restore edge details from the low-level feature map, and refine the output of the decoder.

[0053] As Figure 2 and Figure 3 shown in the above embodiments, preferably, the ResNet-50 includes a preset number of convolutional layers, and adopts a bottleneck module as the residual module, and the ResNet-50 is pre-trained on the ImageNet.

[0054] In order to improve feature reuse, the present application uses a residual connection to add the original input tensor to the output tensor to retain feature information.

[0055] In the above embodiments, preferably, the output feature of the residual module is the sum of the residual mapping of the input feature after the convolutional layer and the input feature, the residual module adopts a 1x1 convolutional layer for dimension reduction, a 3x3 convolutional layer, and a 1x1 convolutional layer for dimension restoration.

[0056] Specifically, in the ResNet model, there is a short circuit connection across the neural layers in the residual block, so that the input data can be added to the output data, which makes the network easily learn the identity mapping feature representation, ensures that the learned features can be directly passed to the subsequent layer, but does not introduce too many parameters or computational complexity. Through the design of the residual block, even if a network model increases more neural layers, it can ensure that the performance does not degrade during training.

[0057] As Figure 2 shown, the residual module of ResNet includes two mappings, which are identity mapping and residual mapping. The identity mapping directly passes the input feature to the next layer, while the residual mapping is to perform some nonlinear transformation on the input feature before passing it, which is used to learn the difference between the input and the output. The calculation process of the residual module is shown in the following formula: output=F(x)+x. Where F(x) represents the residual mapping, which is usually composed of a series of convolutional layers.

[0058] The bottleneck module is composed of convolutional layer stacking and jump connection, that is, the input feature is nonlinearly transformed, then the output data is added to the input data, and finally an activation function is applied, which is a complete residual module.

[0059] As Figure 3 shown, the bottleneck module uses a 1x1 convolutional layer for dimension reduction, followed by a 3x3 convolutional layer, and finally a 1x1 convolutional layer for dimension restoration, so that the calculation accuracy is maintained while the computational complexity is reduced. Especially in a deep ResNet network, using a bottleneck module as a residual module can significantly reduce the parameter quantity and computational complexity of the model, and reduce the training cost.

[0060] AsFigure 4 In the above embodiment, preferably, the hollow convolution module adopts four parallel branches, the first three branches fuse feature maps of different receptive field sizes in a cascaded manner, and the last branch adds image-level features through global average pooling and reduces the number of channels to the input channel number through convolution operation.

[0061] Specifically, the output of the encoder has rich high-level features, in order to improve the utilization rate of the encoder output, the hollow convolution module is introduced at the bottom of the encoder to expand the receptive field. Four parallel hollow convolution branches have four receptive field sizes. The size of the feature map receptive field can be modified by adjusting the hollow rate. There is a grid effect in the dilated convolution, which will cause the loss of local information and the decline of small size target detection performance. By cascading dilated convolutions with different hollow rates, it is beneficial to obtain information from a wider range of pixels and avoid the grid effect.

[0062] In the above embodiment, preferably, the encoder reduces the image size to a preset ratio through five downsampling operations, such as reducing the picture size from 512x512 to 16x16. According to the hollow rate of the cascaded branches in the hollow convolution module, the receptive field is correspondingly increased. The hollow rates of the cascaded convolutions in the hollow convolution module are set to 1, 2, and 4 respectively, and the receptive fields of the parallel branches are 3, 7, and 15 respectively. In this way, the entire feature map can be roughly covered. On this basis, an average pooling branch is added to obtain the global information of the image.

[0063] As shown in the formula for adjusting the pixel receptive field in the output feature map of the hollow convolution module: Figure 5

[0064] S=(r-1)×(k-1)+k

[0065] Where S is the receptive field, r is the hollow rate, and k is the convolution kernel size.

[0066] Specifically, the receptive field of 3x3 convolution is 3x3, and the receptive field of 3x3 convolution with a hollow rate of 2 is 5x5.

[0067] In the above embodiment, preferably, in the hollow convolution module, the receptive field size calculation formula of the cascaded hollow convolution is:

[0068] S'=S1+S2-1

[0069] Where S' is the receptive field size after cascading, S1 and S2 are the receptive field sizes of the two cascaded hollow convolutions respectively.

[0070] As shown in the formula for adjusting the pixel receptive field in the output feature map of the hollow convolution module: Figure 6 The increase process of the receptive field of the hollow convolution module is shown.​

[0071] like Figure 7 As shown, in the above embodiment, preferably, the banded convolution module includes four banded convolutions in the horizontal, vertical, left diagonal and right diagonal directions. The input tensor is sent to four parallel banded convolution branches after a 1×1 convolution operation, and then the four output tensors are connected and subjected to a 1×1 convolution operation to capture the long-range context information in the feature map from four different directions, so that each pixel in the output feature map can be connected with the pixels in four directions of the input feature map.

[0072] Specifically, the convolutions in most convolutional neural networks in the current prior art are usually square convolution kernels, and feature maps are learned in square windows. This convolution operation is very suitable for natural objects that are mostly block-shaped. However, the roads studied in the present invention have the characteristics of narrowness, large span, and continuous distribution. If square convolution is used to capture the characteristics of the road, it may require a very large square to cover a road, which will inevitably merge a lot of irrelevant information. Therefore, the strip convolution of the present invention uses a long strip of convolution kernel, which is more in line with the shape of the road and can better capture the characteristics of the road area. The strip convolution will capture local context information along the specified spatial direction and prevent irrelevant areas from interfering with feature learning.

[0073] The strip convolution module uses four strip convolutions in horizontal, vertical, left diagonal and right diagonal directions to capture long-range context information from four different directions, so that each pixel in the output feature map can be connected with multiple pixels in four directions in the input feature map. Let K∈R 2r+1 is a 2r+1 banded convolution kernel, D=(D h ,D w ) is the direction of the strip convolution kernel K, Y D ∈R C×H×W Represents the result of the banded convolution operation. The formula for banded convolution is as follows:

[0074]

[0075] Where X*K represents the convolution operation, and D is the direction vector of the striped convolution. For horizontal, vertical, left diagonal, and right diagonal convolutions, D is (0, 1), (1, 0), (1, 1), and (-1, 1), respectively. For the convolution kernel K, we set r to 4, resulting in 9 parameters for the striped convolution, the same number of parameters as a 3×3 square convolution operation.

[0076] In the banded convolution module, the input tensor undergoes a 1×1 convolution operation and is then fed into four parallel branches, each branch performing a banded convolution in a different direction. The four output tensors are then concatenated and subjected to a 1×1 convolution operation.

[0077] The application further provides a road weak information extraction system based on the ASU-Net.

[0078] The extraction model construction module is configured to construct a road extraction network model ASU-Net by adopting an encoder-decoder structure, wherein ResNet-50 is adopted as the encoder, a hollow convolution module is introduced in the encoder, and a strip convolution module is introduced in the decoder.

[0079] The image sampling processing module is configured to input a remote sensing satellite road image into the ASU-Net model, and the encoder performs downsampling operation on the image, and the receptive field is expanded by parallel cascading branches with a preset hollow rate in the hollow convolution module.

[0080] The road feature extraction module is configured to perform upsampling on the features learned by the encoder by transposed convolution, and perform road linear feature extraction by adopting four strip convolutions with different directions through the strip convolution module.

[0081] The road weak information extraction system based on the ASU-Net disclosed in the above embodiments corresponds to each step of the road weak information extraction method based on the ASU-Net disclosed in the above embodiments, and the functions to be achieved by each module are consistent. In the implementation process, refer to the above embodiments for operation, and details are not repeated here.

[0082] According to the road weak information extraction method and system based on the ASU-Net disclosed in the above embodiments, the process and effect of the method and system are illustrated by the experimental results in the following embodiments.

[0083] Embodiment 1:

[0084] The embodiment uses Zhouqu road dataset to evaluate the performance of the method proposed in the application.

[0085] As shown in Figure 8 , the Zhouqu road dataset is Zhouqu aerial image data, which covers multiple scene road images of the county town, suburbs, rural areas and mountainous areas of Zhouqu County, Gansu Province. To some extent, this dataset can represent the road characteristics of the mountainous county towns in the central and western regions of China.

[0086] The resolution of the dataset is 5 meters, and the image size is 1024x1024. The dataset has a total of 1305 images. According to the ratio of 8:1:1, the dataset is divided into 910 images for training, 195 images for verification, and 200 images for testing. The experiment enhances the training dataset by scaling, rotating, and cropping.

[0087] In order to verify the model proposed in this paper, the results extracted on Zhouqu data set by BiSeNet and LinkNet are compared and analyzed, and the experimental results of evaluating the model are listed in Table 1. The data in the table are all rounded to four decimal places.

[0088] Table 1 Comparison of experimental results of different models on Zhouqu road data set

[0089]

[0090] It can be seen that the ASUNET network structure proposed in the application is superior to the results of the compared BiSeNet and LinkNet models in various indicators.

[0091] Figure 9 The comparison between different algorithms on Zhouqu road data set is shown, wherein (a) is a remote sensing image; (b) is a road label; (c) is BiSeNet; (d) is LinkNet; (e) is ASU-Net. The dry river channel has high similarity with the road, compared with other algorithms, the use of ASU-Net algorithm can avoid the influence of the river channel to a certain extent (I), for the case that the background and road information are highly similar (II the road information is highly similar to the village house; III the road information is highly similar to the surrounding farmland), the algorithm can still extract complete and clear roads. In addition, in the suburban scene (IV), the algorithm can even identify the unmarked road segment, and has high continuity and completeness. In more complex urban scenes (V city edge, VI city center), the proposed road is more consistent with people's visual cognition of the road.

[0092] In order to test the effectiveness of the proposed hollow convolution module and strip convolution module, the application carries out an ablation experiment, using a U-Net model as a basic model. The results of the ablation experiment are shown in Table 2.

[0093] Table 2 Ablation results

[0094]

[0095] With the addition of ResNet-50, hollow convolution and strip convolution, the accuracy is continuously improved, especially after the addition of hollow convolution. After the addition of strip convolution, the accuracy is not greatly improved, and the reason may be that in this group of data, there are similar dry river channels and other ground objects similar to roads, and the strip convolution has limited effect.

[0096] As Figure 10As shown, wherein (a) is a remote sensing image; (b) is a road label; (c) is U-Net+ResNet-50; (d) is U-Net+ResNet-50+empty; (e) is U-Net+ResNet-50+empty+strip (ASU-Net).

[0097] From the Zhouqu road data set, it can be seen that with the addition of ResNet, empty convolution, strip convolution and the like, in the farmland scene (I, II), the road information is more and more clear and complete, and at the same time, the influence of bare soil and bare river on road information extraction is effectively removed; in the urban fringe, the extraction of road information changes obviously, and the extraction effect of urban roads is good.

[0098] In summary, the accuracy, recall rate, F1 score, intersection over union and average intersection over union of the method of the present application on the Zhouqu road data set are better than those of several current mainstream semantic segmentation networks, such as BiSeNet, LinkNet and U-Net, which shows that the ASU-Net model has good performance.

[0099] The ASU-Net model realizes better balance between precision and recall by introducing the empty convolution module and the strip convolution module, thereby improving the F1 score and mIoU. The results are improved when dealing with roads that are obscured by vegetation and similar to the background color.

[0100] To verify the effectiveness of the empty convolution module and the strip convolution module, the present application uses different models for experiments, which are ResU-Net, ResU-Net+empty convolution module, ResU-Net+empty convolution module+strip convolution module. The experimental results show that the introduction of the empty convolution module and the strip convolution module makes the precision and recall better balanced, thereby improving the F1 score and average intersection over union performance, which proves that the two modules can perform more accurate segmentation on road image extraction.

[0101] The above is only a preferred embodiment of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A road weak information extraction method based on ASU-Net, characterized by: include: An encoder-decoder structure is used to construct a road extraction network model ASU-Net, wherein ResNet-50 is used as the encoder, a dilated convolution module is introduced in the encoder, and a striped convolution module is introduced in the decoder; Inputting a remote sensing satellite road image into the ASU-Net model, the encoder downsamples the image and simultaneously expands the receptive field through parallel cascade branches with preset dilation rates in the dilated convolution module; The features learned by the encoder are upsampled by transposed convolution, and the road linear features are extracted by using four different direction strip convolutions through the strip convolution module.

2. The method for extracting weak road information based on ASU-Net according to claim 1, characterized in that: The encoder and the decoder are connected via a skip connection to perform multi-scale feature information fusion.

3. The method for extracting weak road information based on ASU-Net according to claim 1, characterized in that: The ResNet-50 includes a preset number of convolutional layers and uses a bottleneck module as a residual module. The ResNet-50 is pre-trained on ImageNet.

4. The method for extracting weak road information based on ASU-Net according to claim 3, characterized in that: The output feature of the residual module is the sum of the residual map of the input feature after the convolution layer and the input feature. The residual module uses a 1×1 convolution layer for dimensionality reduction, a 3×3 convolution layer, and a 1×1 convolution layer for dimensionality restoration.

5. The method for extracting weak road information based on ASU-Net according to claim 1, characterized in that: The dilated convolution module adopts four parallel branches, among which the first three branches fuse feature maps of different receptive field sizes in a cascade manner, and the last branch adds image-level features through global average pooling, and then reduces the number of channels to the number of input channels through convolution operation.

6. The method for extracting weak road information based on ASU-Net according to claim 1, characterized in that: The encoder reduces the image size to a preset ratio through five downsampling operations, and increases the receptive field accordingly according to the dilation ratio of the cascade branches in the dilated convolution module.

7. The method for extracting weak road information based on ASU-Net according to claim 1, characterized in that: The strip convolution module includes four strip convolutions in the horizontal, vertical, left diagonal and right diagonal directions. The input tensor is fed into four parallel strip convolution branches after a 1×1 convolution operation. The four output tensors are then connected and subjected to a 1×1 convolution operation to capture long-range contextual information in the feature map from four different directions, so that each pixel in the output feature map can be connected to pixels in four directions in the input feature map.

8. The method for extracting weak road information based on ASU-Net according to claim 5, characterized in that: The calculation formula of the hole convolution module to adjust the pixel receptive field in the output feature map by adjusting the hole rate is: S=(r-1)×(k-1)+k Among them, S is the receptive field, r is the dilation rate, and k is the convolution kernel size.

9. The method for extracting weak road information based on ASU-Net according to claim 8, characterized in that: In the dilated convolution module, the calculation formula for the receptive field size of the cascaded dilated convolution is: S'=S1+S2-1 Among them, S' is the receptive field size after cascading, S1 and S2 are the receptive field sizes of the two dilated convolutions in the cascade respectively.

10. A road weak information extraction system based on ASU-Net, characterized by: Applying the ASU-Net-based road weak information extraction method according to any one of claims 1 to 9, comprising: An extraction model construction module is used to construct a road extraction network model ASU-Net using an encoder-decoder structure, wherein ResNet-50 is used as the encoder, a dilated convolution module is introduced in the encoder, and a striped convolution module is introduced in the decoder; An image sampling and processing module is used to input a remote sensing satellite road image into the ASU-Net model, wherein the encoder performs a downsampling operation on the image and expands the receptive field through parallel cascade branches with a preset dilation rate in the dilated convolution module; The road feature extraction module is used to upsample the features learned by the encoder through transposed convolution, and at the same time extract road linear features through the strip convolution module using strip convolutions in four different directions.

Citation Information

Patent Citations

  • Road extraction method based on residual neural network

    CN110781773A

  • Natural image matting method based on deep learning

    CN111161277A

  • Remote sensing image building extraction method based on DCNN boundary guidance

    CN114387523A

  • High-resolution remote sensing image semantic segmentation method based on object enhancement

    CN116543152A

  • Remote sensing image road extraction method and device based on convolutional network

    CN116543304A

Cited By

  • Handwritten form erasing method and device, electronic equipment and storage medium

    CN121617111A