Road extraction method and device combining panchromatic image and multispectral image
By constructing a deep convolutional neural network combining full-color images and multi-spectral images, the spatial details and spectral information of the images are extracted and fused, and the problem of information loss in the image fusion process in the prior art is solved, and the accuracy of road extraction is improved.
Patent Information
- Application Number
- CN202510010107.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-16
AI Technical Summary
The existing road extraction methods for full-color and multi-spectral images are likely to lead to information loss during the image fusion process, which in turn affects the accuracy of the extraction results and increases the risk of image information loss.
By constructing a deep convolutional neural network combining full-color images and multi-spectral images, the branch structures of the feature extraction part and the fusion part are used to extract the spatial details of the full-color images and the spectral information of the multi-spectral images, and fuse them to complete the road segmentation.
It effectively improves the role of spectral information and spatial details in road extraction, reduces information loss, and improves the accuracy of the extraction results.
Smart Images

Figure CN120014438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road extraction and information fusion, and in particular to a road extraction method and device combining panchromatic images and multispectral images. Background Art
[0002] As the infrastructure of transportation, accurate and complete extraction of relevant information is an important basis for establishing a topological network of land features, which is of great significance to autonomous driving, road monitoring and vehicle navigation. Road extraction under the semantic segmentation mode is based on pixel-level interpretation requirements, and the refined segmentation of roads and backgrounds is achieved through an end-to-end network model. In recent years, benefiting from the development and progress of deep learning technology in the field of semantic segmentation and the refinement of the granularity of optical images in spatial and spectral dimensions, current road interpretation research mainly focuses on deep learning methods driven by optical images. This method combines image and label data, establishes a mapping relationship that meets the interpretation requirements in the deep feature space through a network model, and continuously updates and optimizes the mapping parameters according to the set loss function.
[0003] Optical images are based on the electromagnetic wave reflection energy of the ground objects, and visualize the surface information of the ground objects. They have the advantages of high spatial resolution and rich spectral and texture information. At the same time, with the continuous advancement of satellite platforms and sensors, compared with data such as laser point clouds and radar images, optical images have advantages in coverage and visual matching. Therefore, under the premise of guaranteed data quality, optical images are often the focus of research on refined intelligent interpretation of surface information.
[0004] The types of optical image road extraction technology mainly include model-driven method and data-driven method. Model-driven method includes template matching method and knowledge-driven method, which distinguishes road and background information based on shallow features such as shape, color and texture. Data-driven method is based on images and labels, and uses network models to extract deep features of images, so as to effectively associate remote sensing images and road labels. Considering the current refinement of image description granularity, the richness of training data resources and the improvement of hardware computing power, data-driven method has become the mainstream means of road extraction. From the perspective of data type, most of the current technical methods are directly studied around public data sets. Typical road extraction data sets include DeepGlobe, Massachusetts, SpaceNet, CHN6-CUG and LRSNY, whose image data have high spatial resolution and spectral resolution. However, when imaging with optical remote sensing satellites, the sensor will comprehensively weigh the spatial resolution and spectral resolution, that is, the improvement of one parameter will often lead to the decrease of the other parameter. At present, optical remote sensing satellites of various countries (ZY3, GF2, SPOT7, WorldView3, etc.) usually adopt the combined panchromatic and multispectral imaging mode. Panchromatic images have a wider band and can obtain more energy, so the spatial resolution is higher, but there is no color information; the energy receiving band of multispectral images needs to be separated from the panchromatic band, resulting in a decrease in the ability to depict spatial details, but the description level in the spectral dimension is increased, so the spectral resolution is higher. Therefore, there is a certain complementary relationship between panchromatic images and multispectral images in terms of spatial details and spectral information. How to give full play to the advantages of the two types of data through fusion technology is often the focus of back-end applications. The current processing strategy usually divides fusion and extraction into stages, so the extraction accuracy is closely related to both stages, and the image fusion process inevitably causes information loss. Therefore, the above strategy will lead to more complex factors affecting the accuracy of the extraction results, and the risk of image information loss will also be greater. Summary of the invention
[0005] In view of the problem that the existing road extraction methods of panchromatic images and multispectral images have information loss in the image fusion process, which makes the accuracy influencing factors of the extraction results more complicated and the risk of image information loss is greater, the present invention provides a road extraction method that combines panchromatic images and multispectral images. By constructing a deep convolutional neural network structure that combines panchromatic images and multispectral images, the role of spectral information and spatial details in road extraction can be comprehensively improved.
[0006] In a first aspect, the present invention provides a road extraction method combining panchromatic image and multispectral image, comprising:
[0007] A deep convolutional neural network combining panchromatic image and multispectral image is constructed, wherein the deep convolutional network comprises two major parts and three branches. The first part is a feature extraction part, including a panchromatic image branch and a multispectral image branch, which are respectively used to extract features of the panchromatic image and the multispectral image, and output multi-level feature information; the second part is a fusion part, which is a fusion branch, and is used to extract spatial details of the panchromatic image and spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation;
[0008] Road extraction is performed based on the constructed deep convolutional neural network by combining panchromatic and multispectral images.
[0009] Furthermore, the panchromatic image branch and the multispectral image branch have the same structure and independent model parameters, including an encoder and a decoder;
[0010] The encoder uses ResNet34 for feature extraction, and the decoder uses a transposed convolution unit as a basic unit for upsampling, while a jump connection is adopted to add the intermediate layers of the encoder and the decoder pixel by pixel;
[0011] The transposed convolution unit includes a 1×1 convolution, a 3×3 transposed convolution and a 1×1 convolution connected in sequence, and each convolution is connected to a normalization layer and a ReLU nonlinear activation layer.
[0012] Further, the fusion branch includes a polarized self-attention module, a first convolution unit, a polarized self-attention module, and a second convolution unit connected in sequence;
[0013] Correspondingly, the fusion branch is used to fuse the multi-level feature information and complete road segmentation, including:
[0014] Adding attributes of the output of the first transposed convolution unit of the panchromatic image branch and the output of the first transposed convolution unit of the multispectral image branch and fusing the spectral information to obtain a first feature;
[0015] Passing the first feature through the first polarized self-attention module and the first convolution unit to obtain a second feature;
[0016] Perform channel superposition on the output of the third transposed convolution unit of the panchromatic image branch, the output of the third transposed convolution unit of the multispectral image branch, and the spatial detail to obtain a third feature;
[0017] Adding the second feature, the third feature and the spectrum information of the multispectral image to obtain a fourth feature;
[0018] The fourth feature is passed through the second polarized self-attention module and the second convolution unit to obtain the output of the fusion branch.
[0019] Furthermore, extracting the spatial details of the full-color image specifically includes:
[0020] The high-frequency information of the panchromatic image in the horizontal, vertical and diagonal directions is obtained by Haar wavelet transform. The calculation formulas of the low-frequency information and high-frequency information of the Haar wavelet transform are as follows:
[0021]
[0022] Among them, i and j represent the indexes of the rows and columns of the panchromatic image respectively, X() represents the attribute value of the corresponding position in the panchromatic image, and A(i,j) and D(i,j) represent the low-frequency information and high-frequency information of the panchromatic image respectively.
[0023] Furthermore, the extracting spectral information of the multispectral image specifically includes:
[0024] The multispectral image is transformed by HIS color transformation to obtain hue and saturation spectrum information.
[0025] Furthermore, the first convolution unit includes a 1×1 convolution layer, a normalization layer, and a ReLU activation layer;
[0026] The second convolution unit includes a transposed convolution with a convolution kernel of 4 and a step size of 2 and a convolution with a convolution kernel of 3 and a step size of 1. The transposed convolution is connected to a ReLU activation layer, and the convolution is connected to a sigmoid activation layer.
[0027] Furthermore, the polarized self-attention module includes a channel polarized self-attention mechanism and a spatial polarized self-attention mechanism, and the output is the sum of the outputs of the two self-attention mechanisms.
[0028] In a second aspect, the present invention provides a road extraction device combining panchromatic image and multispectral image, comprising:
[0029] A network construction module is used to construct a deep convolutional neural network that combines panchromatic images and multispectral images. The deep convolutional network includes two major parts and three branches. The first part is a feature extraction part, including a panchromatic image branch and a multispectral image branch, which are used to extract features of the panchromatic image and the multispectral image, respectively, and output multi-level feature information; the second part is a fusion part, which is a fusion branch, used to extract the spatial details of the panchromatic image and the spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation;
[0030] The road extraction module is used to extract roads by combining panchromatic images and multispectral images based on the constructed deep convolutional neural network.
[0031] In a third aspect, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described above when executing the program.
[0032] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program executes the method described above when executed by a processor.
[0033] Beneficial effects of the present invention:
[0034] (1) The present invention focuses on the spatial details of full-color images and uses Haar wavelet transform to decompose spatial features in the frequency domain into high-frequency information and low-frequency information, and integrates multi-directional high-frequency information into the structural level shallow features of the fusion branch to make up for the shortcomings of insufficient spatial details of multi-spectral images.
[0035] (2) The present invention targets the spectral information of multispectral images, obtains the hue and saturation of the multispectral images based on the color space transformation in the HIS fusion method, and integrates the above information into the semantic-level deep features and structural-level shallow features of the fusion branch respectively, to make up for the shortcoming of insufficient spectral information of panchromatic images.
[0036] (3) In order to address the risk of information loss in the fusion process, the present invention uses the polarized self-attention mechanism to improve the fusion efficiency of panchromatic images and multispectral images, highlight the key spatial positions and channel dimensions, and comprehensively enhance the role of spectral information and spatial details in road extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of a process of a road extraction method combining panchromatic image and multispectral image provided by an embodiment of the present invention;
[0038] Figure 2 A schematic diagram of the structure of a deep convolutional neural network model provided by an embodiment of the present invention;
[0039] Figure 3 A schematic diagram of generating hue and saturation provided by an embodiment of the present invention;
[0040] Figure 4 A schematic diagram of the structure of a channel polarization self-attention module provided in an embodiment of the present invention;
[0041] Figure 5 A schematic diagram of the structure of a spatial polarization self-attention module provided in an embodiment of the present invention;
[0042] Figure 6 Schematic diagram of the electronic device structure provided by the embodiment of the present invention
[0043] Figure 7 Schematic diagram of road extraction results using different methods provided in embodiments of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] Information fusion is a key technology to achieve complementary advantages of data of different types and structures, and is also the mainstream means of integrating superior information in the current big data environment. In the field of remote sensing, from the perspective of information fusion, there are mainly two forms: "series" and "parallel".
[0046] (1) The implementation path of the "series" form is generally to use a certain data for primary extraction first, and then use another data for post-extraction based on this, so as to achieve a process from coarse to fine, from global to detailed. For example, open source geographic information data and crowdsourcing data are used for primary extraction and sample preparation, and then post-extraction is achieved through deep learning; first, the semantic segmentation confidence map of a single data is used as auxiliary information, and then multiple data are combined as input for fine extraction.
[0047] (2) The implementation strategy of the “parallel” form is to use data or features from different sources as inputs simultaneously to achieve information fusion. The dynamic and static features of trajectory data and the deep features of remote sensing images are used as inputs of the deep convolutional neural network to effectively integrate the superior information.
[0048] However, considering the structural differences of different types of data, the above fusion forms are often too direct and simple, and cannot fully integrate the complementary information of multi-source data. Therefore, feature-level fusion is the current research focus. Simple feature-level fusion includes cascade, summation and attention mechanisms. Subsequently, some researchers have also conducted research on the mechanism design of fusion modules around the two dimensions of channels and space. However, most of the above studies are directly based on the perspective of computer vision, and their own characteristics and relationships in treating fused information need to be further explored.
[0049] Therefore, if Figure 1 As shown, an embodiment of the present invention provides a road extraction method combining panchromatic image and multispectral image, comprising:
[0050] Step 1: Build a deep convolutional neural network that combines panchromatic and multispectral images.
[0051] Specifically, the deep convolutional network consists of two parts and three branches. The first part is the feature extraction part, including the panchromatic image branch and the multispectral image branch, which are used to extract the features of the panchromatic image and the multispectral image respectively, and output multi-level feature information; the second part is the fusion part, which is a fusion branch, used to extract the spatial details of the panchromatic image and the spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation;
[0052] Step 2: Extract roads by combining panchromatic and multispectral images based on the constructed deep convolutional neural network.
[0053] In view of the problem that the existing road extraction methods of panchromatic images and multispectral images have information loss in the image fusion process, which makes the accuracy influencing factors of the extraction results more complicated and the risk of image information loss is greater, the present invention completes the road extraction operation by extracting multi-level feature information of panchromatic images and multispectral images, and fusing the high-frequency information of panchromatic images and the spectral information of multispectral images.
[0054] Based on the above embodiments, the feature extraction part is further designed in the embodiments of the present invention.
[0055] like Figure 2 As shown, the panchromatic image branch and the multispectral image branch of the feature extraction part of this embodiment have the same structure and independent model parameters, including an encoder and a decoder; the encoder uses ResNet34 for feature extraction, and the decoder uses a transposed convolution unit as a basic unit for upsampling, and a jump connection is used to add the intermediate layers of the encoder and decoder pixel by pixel.
[0056] Specifically, "Ei-pan / mss" and "Di-pan / mss" (i = 1, 2, 3) represent the encoding and decoding features of the panchromatic image and multispectral image branches, and "i" reflects the object of the jump connection, such as the object of the E1-pan jump connection is D1-pan, and the object of the E1-mss jump connection is D1-mss. The jump connection is in the form of pixel-by-pixel addition. Using the jump connection form can avoid excessive loss of information.
[0057] Based on the above embodiments, the embodiments of the present invention provide further designs for fusion branches.
[0058] like Figure 2As shown in the figure, the input of the fusion branch comes from three parts. The first is the features of the panchromatic image branch and the multispectral image branch at the decoding layer. To ensure that the input features contain both semantic and detail information, deep and shallow features are used as input respectively. The second is the spectral information of the multispectral image. The hue and saturation generated by the HIS transform are fused with the shallow and deep features in the form of channel superposition. The third is the spatial information of the panchromatic image. The Haar wavelet transform is used to obtain high-frequency information in different directions. Since the high-frequency information needs to retain details, the information is only fused in the shallow features in the channel superposition. In addition, in order to retain more and more information in the fusion stage and enhance the fine-grained perception ability in intensive prediction tasks, the fusion branch uses the polarized self-attention (PSA) module to optimize the features of the channel and spatial dimensions after obtaining the input features of each part.
[0059] Therefore, the fusion branch includes a polarized self-attention module, a first convolution unit, a polarized self-attention module and a second convolution unit connected in sequence; wherein the first convolution unit includes a 1×1 convolution layer, a normalization layer and a ReLU activation layer; the second convolution unit includes a transposed convolution with a convolution kernel of 4 and a stride of 2 and a convolution with a convolution kernel of 3 and a stride of 1, the transposed convolution is connected to the ReLU activation layer, and the convolution is connected to the sigmoid activation layer.
[0060] Correspondingly, the fusion branch processing operation is as follows:
[0061] The output of the first transposed convolution unit of the panchromatic image branch and the output of the first transposed convolution unit of the multispectral image branch are added and the spectral information is fused to obtain the first feature;
[0062] Pass the first feature through the first polarized self-attention module and the first convolution unit to obtain the second feature;
[0063] Channel-superimpose the output of the third transposed convolution unit of the panchromatic image branch, the output of the third transposed convolution unit of the multispectral image branch, and the spatial details to obtain a third feature;
[0064] The second feature, the third feature and the spectral information of the multispectral image are added together to obtain the fourth feature;
[0065] The fourth feature passes through the second polarized self-attention module and the second convolution unit to obtain the output of the fusion branch.
[0066] Furthermore, the spatial details of the panchromatic image are extracted, including:
[0067] The high-frequency information of the panchromatic image in the horizontal, vertical and diagonal directions is obtained by Haar wavelet transform. The calculation formula of the low-frequency information and high-frequency information of Haar wavelet transform is as follows:
[0068]
[0069] Among them, i and j represent the indexes of the rows and columns of the panchromatic image respectively, X() represents the attribute value of the corresponding position in the panchromatic image, and A(i,j) and D(i,j) represent the low-frequency information and high-frequency information of the panchromatic image respectively.
[0070] Specifically, the overall and detail information is obtained by averaging and differencing adjacent pixels, and high-frequency information in three directions is finally obtained based on mutual cross-calculations. Among them, low-pass filtering followed by high-pass filtering obtains high-frequency information in the horizontal direction; high-pass filtering followed by low-pass filtering obtains high-frequency information in the vertical direction; and high-pass filtering twice obtains high-frequency information in the diagonal direction. The calculation formula of A(i,j) is equivalent to low-pass filtering, which obtains low-frequency information, and the calculation formula of D(i,j) is equivalent to high-pass filtering, which obtains high-frequency information.
[0071] It can be understood that the sensing band of panchromatic images is wider than that of multispectral images, and they can sense greater electromagnetic wave energy, so their detailed information is often richer. Haar wavelet transform can obtain high-frequency information of images in the horizontal, vertical and diagonal directions in the frequency domain, and avoid the influence of spatial noise to a certain extent. As a strip-shaped feature formed by artificial paving or long-term rolling of vehicles, the surface material of the road is usually different from that of the adjacent features, while the internal material of the same road is often the same, so the edge of the road is generally high-frequency information, and the interior of the road is low-frequency information. In addition, unlike surface features such as buildings and waters, the shape of roads is mostly strip-shaped, so selecting high-frequency information in the horizontal, vertical and diagonal directions and introducing it into the fusion branch will help further focus on the road target and weaken the information interference of other features.
[0072] Furthermore, the spectral information of the multispectral image is extracted, specifically including:
[0073] The multispectral image is transformed by HIS color transformation to obtain the hue and saturation spectral information.
[0074] Specifically, after HIS color space transformation, three attributes, hue, saturation and brightness, can be obtained. The physical meaning of brightness is consistent with that of full-color images, and full-color images have higher spatial resolution, so the brightness attribute is not considered in this link; hue and saturation are two important attributes related to color. The generation process of the two attributes is as follows Figure 3 As shown in the figure, the size of its attributes is closely related to the electromagnetic spectrum reflection characteristics of the ground object, so adding it to the fusion branch can fully represent the injection of spectral information.
[0075] It can be understood that multispectral images can achieve band-by-band collection of ground object attributes compared to panchromatic images, which are expressed as multi-band grayscale values of pixels. For computers, a higher-dimensional feature space can be established, thereby upgrading the distinction form from lines to surfaces. As a universal way to describe color information, HIS color space can effectively convert spectral information into color information that conforms to visual perception. On this basis, a road information perception path that fits human vision can be built later.
[0076] Furthermore, the polarized self-attention module includes a channel polarized self-attention module and a spatial polarized self-attention module, and the output is the sum of the outputs of the two self-attention modules.
[0077] Specifically, PSA includes a channel polarization self-attention mechanism and a spatial polarization self-attention mechanism, and its final output is the sum of the output results of the two components. Compared with other attention mechanisms, PSA has less compression in the spatial and channel dimensions, so it can effectively reduce information loss. At the same time, in terms of nonlinear activation, it combines the Sigmoid and Softmax functions to enhance fine-grained perception capabilities, making it more suitable for intensive prediction tasks such as road extraction.
[0078] Specifically, the channel polarization self-attention module structure is as follows: Figure 4 As shown, the specific calculation formula is as follows:
[0079]
[0080] Among them, X in represents the input features, X c represents the output of the channel polarization self-attention module, W z1 , W q1 and W v1 is a 1*1 convolution with different weight parameters, W z1 Followed by LayerNorm, F Sig and F Soft Represent the Sigmoid function and the Softmax function respectively, represents matrix multiplication, ⊙ ch It is the channel dimension multiplication, R1 and R2 represent two different Reshape operations;
[0081] Specifically, the spatial polarization self-attention module structure is as follows: Figure 5 As shown, the specific calculation formula is as follows:
[0082]
[0083] Among them, X in represents the input features, X srepresents the output of the spatial polarization self-attention module, R3 and R4 represent two different Reshape operations, and W v2 and W q2 is a 1*1 convolution with different weight parameters, F GAP represents global average pooling, ⊙ sp Represents spatial dimension multiplication.
[0084] The embodiment of the present invention further provides a road extraction device combining panchromatic image and multispectral image, comprising:
[0085] The network construction module is used to construct a deep convolutional neural network that combines panchromatic images and multispectral images. The deep convolutional network consists of two major parts and three branches. The first part is the feature extraction part, including the panchromatic image branch and the multispectral image branch, which are used to extract the features of the panchromatic image and the multispectral image respectively, and output multi-level feature information; the second part is the fusion part, which is a fusion branch, used to extract the spatial details of the panchromatic image and the spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation;
[0086] The road extraction module is used to extract roads by combining panchromatic images and multispectral images based on the constructed deep convolutional neural network.
[0087] like Figure 6 As shown, an embodiment of the present invention further provides an electronic device, including: a processor (processor) 601, a communication interface (Communications Interface) 602, a memory (memory) 603 and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 can call the logic instructions in the memory 603 to execute the method in the above embodiment, and the method includes:
[0088] A deep convolutional neural network that combines panchromatic images and multispectral images is constructed. The deep convolutional network consists of two major parts and three branches. The first part is the feature extraction part, including a panchromatic image branch and a multispectral image branch, which are used to extract the features of panchromatic images and multispectral images respectively, and output multi-level feature information; the second part is the fusion part, which is a fusion branch, used to extract the spatial details of the panchromatic image and the spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation; road extraction is performed based on the constructed deep convolutional neural network combined with panchromatic images and multispectral images.
[0089] In addition, when the logic instructions in the above-mentioned memory 603 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0090] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in the above embodiment is implemented, for example, including:
[0091] A deep convolutional neural network that combines panchromatic images and multispectral images is constructed. The deep convolutional network consists of two major parts and three branches. The first part is the feature extraction part, including a panchromatic image branch and a multispectral image branch, which are used to extract the features of panchromatic images and multispectral images respectively, and output multi-level feature information; the second part is the fusion part, which is a fusion branch, used to extract the spatial details of the panchromatic image and the spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation; road extraction is performed based on the constructed deep convolutional neural network combined with panchromatic images and multispectral images.
[0092] The feasibility and effect of the method provided by the present invention are verified by experiments below:
[0093] 1. Experimental data:
[0094] GF2-FC data set: The remote sensing images of this data set are from GF-2. The image range includes three areas: Shanghai, Taiyuan and Dalian. The road types cover multiple levels such as expressways, ordinary roads, urban roads and rural roads. The present invention selects panchromatic images (spatial resolution: 0.8m) and multispectral images (spatial resolution: 3.2m) for experiments and analysis. The specific method is to cut the original image into 512*512 pixels, select 2100 groups of data (each group of data contains a panchromatic image and a corresponding multispectral image) for training, and 300 groups of data for testing. The multispectral images during training and testing only involve three RGB bands.
[0095] 2. Experimental Details
[0096] The core hardware configuration of the experimental environment is 1 NVIDIA Tesla V100 graphics card with 32G video memory. Adam is selected as the optimizer for network training, and the initial learning rate is set to 2e-4. Whenever the loss value is higher than the current optimal loss value for 3 consecutive times, the learning rate is reduced by 5 times. The training data block size is 16, the iteration epoch value is 100, and the loss function uses the sum of BCE loss and dicecoefficient loss as the basis for supervised training. At the same time, in order to enhance the sample processing, the training data is randomly (50%) flipped vertically, horizontally, diagonally, and radially transformed.
[0097] 3. Index evaluation
[0098] In order to conduct a comprehensive and integrated accuracy evaluation and effectively take into account the needs of road extraction results in practical application scenarios. The present invention performs accuracy statistics of road extraction results at two levels: pixel level and topology level. The pixel-level accuracy evaluation indicators include correctness (P), recall rate (R), F1 score, overall accuracy (OA) and intersection over union (IoU); the topology-level accuracy evaluation indicators include completeness rate (Com) and error rate (Eor).
[0099] 4. Experimental Results
[0100] The experiment mainly uses the selected data level to compare the proposed method with existing classic and advanced methods. There are 10 comparison methods, of which 6 methods support single input, including ASPPUNet, D-LinkNet, MANet, SGCN, UNetFormer and DT-Net; 4 methods support dual input, including MCANe, t,JoiTriNet_e, JoiTriNet_d and CMFNet. The road extraction of the selected data sets is compared and analyzed from the two perspectives of quantitative accuracy and qualitative results.
[0101] Table 1 shows the accuracy statistics of the method of the present invention and the comparative method on the GF2-FC dataset, where a single input only contains a certain type of data. Therefore, in order to comprehensively compare the extraction effects of each method, the present invention independently uses panchromatic images (single input—panchromatic images) and multispectral images (single input—multispectral images) to conduct road extraction experiments.
[0102] The following conclusions can be drawn from the accuracy statistics: (1) Under the single-input condition, the R index, Com index and comprehensive evaluation index (F1 and IoU) of various methods are higher than those of multispectral images, and some methods have obvious advantages, which proves that in the road extraction task, the role of spatial details in mining road features is greater than that of spectral information; (2) Under the single-input condition, the advantages and disadvantages of the P index and Eor index of various methods are not obvious, which proves that the stability of distinguishing roads from similar objects under single-input conditions needs to be improved; (3) The R and Com indicators under the dual-input condition are relatively ideal, but there is a lot of room for improvement in other indicators, indicating that this strategy can extract road information more comprehensively, but also introduces more interfering objects; (4) The five evaluation indicators of P, F1, OA, IoU and Com of the method of the present invention are all the best, and only R and Eor are slightly lower than MANet, which proves that the road extraction results of the method of the present invention are in a leading position in terms of completeness and accuracy.
[0103] Table 1 Comparison of road extraction results accuracy of different methods (unit: %)
[0104]
[0105] In addition to the above quantitative accuracy index comparison analysis, in order to more intuitively compare the road extraction effects of each method, the road extraction results of some test images were selected for comparative analysis. The selected 5 images are from different scenes and basically cover the difficult areas of road extraction. In addition, this part of the analysis no longer covers all the comparison methods, but selects the first 5 methods based on "IoU" as the standard.
[0106] Specific circumstances such as Figure 7As shown in the figure, image (a) is located in a vegetation-covered area and contains three roads. The road in the upper left corner is relatively wide and similar to the background. The five comparison methods did not find this road, while the extraction results of the method of the present invention showed the existence of this road. For the middle road on the left, all methods have a certain degree of mis-extraction and missed extraction, and the missed extraction of the method of the present invention is relatively rare. For the road on the right, MANet, UNetFormer and DT-Net have relatively serious mis-extraction, while SGCN and CMFNet have obvious "broken road" problems. The coverage scene of image (b) is similar to that of image (a). The road information is mainly a horizontal path, which is equivalent to a low-level road, and the left half of the road is obviously covered by vegetation. From the extraction results, the five comparison methods have different degrees of missed extraction problems. Image (c) is located in the suburban area and contains roads of different levels. From the visual effect, MANet, UNetFormer and DT-Net introduced wrong road information and destroyed the topological structure of the road. SGCN, UNetFormer and DT-Net failed to overcome the interference of shadows and did not extract the roads in the middle and lower areas (shadow covered). Image (d) is located in the urban area. The buildings around the road are tall, and the vehicles show obvious point interference in the image. SGCN has obvious problems of missed extraction, and CMFNet has significant misextraction at the edge of the road. The extraction effects of the other methods are ideal. Image (e) is located in the town area and contains two main roads that are almost perpendicular to each other. DT-Net and the method of the present invention have better extraction effects. The other four methods have misextraction, and UNetFormer has certain problems of missed extraction in the left area of the lateral road.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A road extraction method combining panchromatic image and multispectral image, characterized in that: include: A deep convolutional neural network combining panchromatic image and multispectral image is constructed, wherein the deep convolutional network comprises two major parts and three branches. The first part is a feature extraction part, including a panchromatic image branch and a multispectral image branch, which are respectively used to extract features of the panchromatic image and the multispectral image, and output multi-level feature information; the second part is a fusion part, which is a fusion branch, and is used to extract spatial details of the panchromatic image and spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation; Road extraction is performed based on the constructed deep convolutional neural network by combining panchromatic and multispectral images.
2. The road extraction method combining panchromatic image and multispectral image according to claim 1, characterized in that: The panchromatic image branch and the multispectral image branch have the same structure and independent model parameters, including an encoder and a decoder; The encoder uses ResNet34 for feature extraction, and the decoder uses a transposed convolution unit as a basic unit for upsampling, while a jump connection is adopted to add the intermediate layers of the encoder and the decoder pixel by pixel; The transposed convolution unit includes a 1×1 convolution, a 3×3 transposed convolution and a 1×1 convolution connected in sequence, and each convolution is connected to a normalization layer and a ReLU nonlinear activation layer.
3. The road extraction method combining panchromatic image and multispectral image according to claim 1, characterized in that: The fusion branch includes a polarized self-attention module, a first convolution unit, a polarized self-attention module, and a second convolution unit connected in sequence; Correspondingly, the fusion branch is used to fuse the multi-level feature information and complete road segmentation, including: Adding attributes of the output of the first transposed convolution unit of the panchromatic image branch and the output of the first transposed convolution unit of the multispectral image branch and fusing the spectral information to obtain a first feature; Passing the first feature through the first polarized self-attention module and the first convolution unit to obtain a second feature; Perform channel superposition on the output of the third transposed convolution unit of the panchromatic image branch, the output of the third transposed convolution unit of the multispectral image branch, and the spatial detail to obtain a third feature; Adding the second feature, the third feature and the spectrum information of the multispectral image to obtain a fourth feature; The fourth feature is passed through the second polarized self-attention module and the second convolution unit to obtain the output of the fusion branch.
4. The road extraction method combining panchromatic image and multispectral image according to claim 1, characterized in that: The extracting of the spatial details of the panchromatic image specifically includes: The high-frequency information of the panchromatic image in the horizontal, vertical and diagonal directions is obtained by Haar wavelet transform. The calculation formula of the low-frequency information and high-frequency information of the Haar wavelet transform is as follows: Among them, i and j represent the indexes of the rows and columns of the panchromatic image respectively, X() represents the attribute value of the corresponding position in the panchromatic image, and A(i,j) and D(i,j) represent the low-frequency information and high-frequency information of the panchromatic image respectively.
5. The road extraction method combining panchromatic image and multispectral image according to claim 1, characterized in that: The extracting spectral information of the multispectral image specifically includes: The multispectral image is transformed by HIS color transformation to obtain hue and saturation spectrum information.
6. The road extraction method combining panchromatic image and multispectral image according to claim 3 is characterized in that: The first convolution unit includes a 1×1 convolution layer, a normalization layer, and a ReLU activation layer; The second convolution unit includes a transposed convolution with a convolution kernel of 4 and a step size of 2 and a convolution with a convolution kernel of 3 and a step size of 1. The transposed convolution is connected to a ReLU activation layer, and the convolution is connected to a sigmoid activation layer.
7. The road extraction method combining panchromatic image and multispectral image according to claim 3 is characterized in that: The polarized self-attention module includes a channel polarized self-attention mechanism and a spatial polarized self-attention mechanism, and the output is the sum of the outputs of the two self-attention mechanisms.
8. A road extraction device combining panchromatic image and multispectral image, characterized in that: include: A network construction module is used to construct a deep convolutional neural network that combines panchromatic images and multispectral images. The deep convolutional network includes two major parts and three branches. The first part is a feature extraction part, including a panchromatic image branch and a multispectral image branch, which are used to extract features of the panchromatic image and the multispectral image, respectively, and output multi-level feature information; the second part is a fusion part, which is a fusion branch, used to extract the spatial details of the panchromatic image and the spectral information of the multispectral image and fuse them with the multi-level feature information to complete road segmentation; The road extraction module is used to extract roads by combining panchromatic images and multispectral images based on the constructed deep convolutional neural network.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is performed.