High-resolution building extraction method and device based on feature decoupling and recoupling
By using a high-resolution network with feature decoupling and recoupling in remote sensing image processing, the problem of hollowness and insufficient resolution of building extraction results in remote sensing images is solved, and fine extraction of building footprints and high-precision boundary detection are achieved.
Patent Information
- Application Number
- CN202311533421.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-11-16
AI Technical Summary
When extracting buildings in remote sensing images in the prior art, the internal cavity is prone to appear in the interior and the resolution is insufficient, making it difficult to achieve fine extraction of building footprints.
A high-resolution network based on feature decoupling and recoupling is adopted to process remote sensing images through the decoupling high-resolution network, separate boundary features and subject features, and repair and enhance them, ultimately realizing fine extraction of building boundaries and segmentation.
Through the method of feature decoupling and recoupling, the accuracy and resolution of the building extraction results are significantly improved, the holes in the extraction results are reduced, and more accurate and fine extraction of building footprints in remote sensing images is achieved.
Smart Images

Figure CN117671487B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a high-resolution building extraction method and device based on feature decoupling and recoupling. Background Art
[0002] With the increase in application demand and the development of deep learning technology in the field of computer vision, building extraction methods have gradually shifted from traditional methods based on hand-designed features to methods based on deep learning. In recent years, deep learning methods, led by convolutional neural networks, can learn deep features and high-level semantic features of images under data-driven conditions, have stronger generalization capabilities, higher accuracy and efficiency, and have been widely used in the field of computer vision. They have also brought new solutions for faster and more accurate extraction of buildings, especially the network framework for semantic segmentation and saliency detection tasks.
[0003] Although these methods have achieved outstanding results in the field of computer vision, they are not designed for the characteristics of remote sensing images. In particular, compared with the "street view" of natural images, remote sensing images collected from satellites or spacecraft cover a wider range and have more complex information. Compared with natural images, ground objects in remote sensing images may have different semantics at different scales, which makes it easy for holes to appear inside the extraction results of large buildings; and compared with the image resolution used in computer vision, the resolution of high-resolution images is lacking. Summary of the invention
[0004] The present invention provides a high-resolution building extraction method and device based on feature decoupling and recoupling, so as to solve the defects of easy appearance of holes and lack of resolution in the extraction results in the prior art, and realize the fine extraction of building footprints in remote sensing images.
[0005] The present invention provides a high-resolution building extraction method based on feature decoupling and recoupling, comprising:
[0006] Inputting the target remote sensing image into the trained decoupled high-resolution network to obtain a building prediction result of the target remote sensing image;
[0007] Using the building prediction result of the target remote sensing image as the building extraction result of the target remote sensing image;
[0008] Among them, the building prediction result of the target remote sensing image includes the boundary prediction result of the target remote sensing image and the segmentation prediction result of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained according to a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block.
[0009] According to a high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention, the decoupled high-resolution network includes a shallow feature extraction module, a first transition layer, a first multi-scale feature fusion module group, a second transition layer, a second multi-scale feature fusion module group and a prediction result output module connected in sequence;
[0010] The shallow feature extraction module is used to extract features from the input image to obtain first resolution features;
[0011] The first transition layer is used to generate a second resolution feature based on the first resolution feature;
[0012] The first multi-scale feature fusion module group is used to obtain output feature results at multiple scales based on the first resolution feature and the second resolution feature;
[0013] The second transition layer is used to generate a third resolution feature based on the first resolution feature and the second resolution feature after the first multi-scale feature fusion module group;
[0014] The second multi-scale feature fusion module group is used to obtain a multi-scale output feature result based on the first resolution feature, the second resolution feature and the third resolution feature after being fused by the first multi-scale feature fusion module group;
[0015] The prediction result output module is used to fuse and upsample the boundary features and segmentation features based on the output feature results of the first multi-scale feature fusion module group and the output feature results of the second multi-scale feature fusion module group, so as to obtain the boundary prediction result and segmentation prediction result of the input image.
[0016] According to a high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention, the feature decoupling and recoupling module includes:
[0017] Flow field generation module, feature decoupling module, feature enhancement module and feature fusion module;
[0018] The flow field generation module is used to calculate and obtain the flow field after splicing the high-resolution features and the low-resolution features;
[0019] The feature decoupling module is used to obtain boundary features and main features respectively based on the flow field;
[0020] The feature enhancement module is used to repair and enhance the boundary feature and the main feature to obtain an enhanced boundary feature and an enhanced main feature respectively;
[0021] The feature fusion module is used to obtain a boundary information confidence map and a main information confidence map based on the enhanced boundary feature and the enhanced main feature, respectively, and perform feature mosaicking based on the enhanced boundary feature, the enhanced main feature, the boundary information confidence map and the main information confidence map to obtain a fused feature;
[0022] Among them, the output feature result of the multi-scale feature fusion module is obtained according to the enhanced boundary feature and the fusion feature.
[0023] According to a high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention, the boundary features and the main features are repaired and enhanced to obtain enhanced boundary features and enhanced main features, respectively, which specifically includes:
[0024] Using the shallow feature or the enhanced boundary feature output by the previous feature decoupling and recoupling module as a supplement to the boundary feature, repairing the boundary feature to obtain the enhanced boundary feature; and,
[0025] The low-resolution feature is used as a supplement to the main feature, and the main feature is repaired to obtain the enhanced main feature.
[0026] According to a high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention, the training method of the decoupled high-resolution network includes:
[0027] Inputting the sample remote sensing images in the sample remote sensing image training set into the initialized decoupled high-resolution network to obtain the output feature results of each multi-scale feature fusion module and the building prediction results of the sample remote sensing images;
[0028] The output feature results of each multi-scale feature fusion module are deeply supervised, and the decoupled high-resolution network is trained based on the output feature results of each multi-scale feature fusion module, the building prediction results of the sample remote sensing image, the building extraction result labels of the sample remote sensing image and the loss function.
[0029] According to a high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention, for the multi-scale feature fusion module in the second multi-scale feature fusion module group, the convolution block in the basic block is an expanded convolution block.
[0030] The present invention also provides a high-resolution building extraction device based on feature decoupling and recoupling, comprising:
[0031] An input module, used for inputting a target remote sensing image into a trained decoupled high-resolution network to obtain a building prediction result of the target remote sensing image;
[0032] An output module, used for using the building prediction result of the target remote sensing image as the building extraction result of the target remote sensing image;
[0033] Among them, the building prediction result of the target remote sensing image includes the boundary prediction result of the target remote sensing image and the segmentation prediction result of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained according to a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block.
[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the high-resolution building extraction method based on feature decoupling and recoupling as described above is implemented.
[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described high-resolution building extraction methods based on feature decoupling and recoupling.
[0036] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned high-resolution building extraction methods based on feature decoupling and recoupling.
[0037] The high-resolution building extraction method and device based on feature decoupling and recoupling provided by the present invention are designed with a parallel structure, a decoupled high-resolution network including a feature decoupling and recoupling module, and then the decoupled high-resolution network is trained according to a sample remote sensing image training set with building extraction result labels, and the target remote sensing image is input into the trained decoupled high-resolution network to obtain the boundary prediction result and segmentation prediction result of the target remote sensing image, and the high-resolution remote sensing image is feature decomposed, and the footprints and boundaries of the buildings on the ground are respectively learned, so that the segmentation and boundaries in the real ground data can be matched more accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0039] Figure 1 A schematic flow chart of a high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention;
[0040] Figure 2 A schematic diagram of the structure of a decoupled high-resolution network provided by the present invention;
[0041] Figure 3 A schematic diagram of the structure of a characteristic decoupling and recoupling module provided by the present invention;
[0042] Figure 4 A schematic diagram of the structure of a high-resolution building extraction device based on feature decoupling and recoupling provided by the present invention;
[0043] Figure 5 This is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] In order to facilitate understanding of the solution provided by the present invention, a brief introduction to the inventive concept of the present invention is first given.
[0046] Although many building feature extraction methods have achieved good results, they are essentially improving the building extraction results by incorporating more semantic context information, without further in-depth research on the root causes of blurred boundaries and internal holes in building extraction results.
[0047] The core of this problem is that the feature representation extracted by the existing network is coupled and redundant, that is, the feature map obtained by the network is optimized by interfering, competing and repeated information. Specifically, in building extraction, information of different scales performs differently in terms of model learning building boundaries and building bodies.
[0048] Among them, global information or high-level semantic information may help improve the continuity of the overall segmentation of the building, but it will interfere with the model's judgment of fine targets such as boundaries; on the contrary, more complete spatial details are helpful for fine extraction of boundaries, but may interfere with targets that require continuity within a large receptive field. Therefore, contradictory information may ultimately lead to the feature representations finally selected by the network to have weak discriminative ability, confusion or semantic bias, limiting the performance and generalization ability of the model.
[0049] Although many networks in current building extraction methods use a multi-task approach and try to separate boundary features by adding boundary detection tasks, achieving a similar decoupling effect, there is still a key problem: multiple tasks are often entangled and share a feature map, which does not change the nature of network feature extraction by optimizing repeated and contradictory information. In addition, the boundaries and segmentation results extracted from multi-tasks are often separated from each other, ignoring the constraint effect of boundary extraction results on segmentation.
[0050] The present invention attempts to migrate the decoupling operation from remote sensing images to feature maps extracted by a network, and provides a high-resolution building extraction method and device based on feature decoupling and recoupling.
[0051] Figure 1 A schematic diagram of the process of the high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention is shown in FIG. Figure 1 As shown, the method includes:
[0052] Step 100: Input the target remote sensing image into the trained decoupled high-resolution network to obtain the building prediction result of the target remote sensing image.
[0053] Step 101: Using the building prediction result of the target remote sensing image as the building extraction result of the target remote sensing image.
[0054] Among them, the building prediction results of the target remote sensing image include the boundary prediction results of the target remote sensing image and the segmentation prediction results of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained based on a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block.
[0055] Specifically, the target remote sensing image refers to a remote sensing image in which segmentation results and boundary results of buildings need to be analyzed.
[0056] The decoupled high-resolution network provided by the present invention is a network improved based on the high-resolution network (High-Resolution Net, HRNet), which is a parallel structure, and is composed of a shallow feature extraction module, a transition layer (Transition), multiple multi-scale feature fusion (Multi-scale Feature Fusion, MFF) modules and a prediction result output module. Among them, the multi-scale feature fusion module is composed of a feature decoupling-recoupling (Feature Decoupling-Recoupling, FDR) module and a basic block.
[0057] In an embodiment of the present invention, the decoupled high-resolution network can use three stacked convolution blocks (conv blocks) and four stacked residual blocks (bottleneck blocks) as shallow feature extraction modules to extract shallow features. Each conv block is composed of a 3×3 convolution, a batch normalization (Batch Normalization, BN) layer and a linear rectifier function (Rectified Linear Unit, ReLU) activation function in series, and the bottleneck block and basic block are obtained by stacking conv blocks and skip connections (shortcuts).
[0058] The MFF module consists of a basic block for extracting more abstract and discriminative features and an FDR module for enhancing and interacting information at different scales. The FDR module can decouple features into features that focus more on boundaries and features that focus more on subjects, and recouple features. At the same time, a shortcut connection is used in the MFF module to supplement the information of the recoupled features.
[0059] The prediction result output module can fuse the multi-scale features obtained by multiple MFF modules, and finally obtain the boundary prediction results and segmentation prediction results as the building prediction results.
[0060] The decoupled high-resolution network can be trained according to a sample remote sensing image training set with building extraction result labels, and the parameters of the decoupled high-resolution network can be adjusted according to the building extraction result labels and the building prediction results obtained by the decoupled high-resolution network until the conditions for training completion are met, thereby obtaining a trained decoupled high-resolution network.
[0061] After obtaining the trained decoupled high-resolution network, the target remote sensing image can be input into the trained decoupled high-resolution network to obtain the building prediction result of the target remote sensing image, and the building prediction result of the target remote sensing image can be used as the building extraction result of the target remote sensing image.
[0062] The high-resolution building extraction method based on feature decoupling and recoupling provided by the present invention designs a parallel structure, a decoupled high-resolution network including a feature decoupling and recoupling module, and then trains the decoupled high-resolution network according to a sample remote sensing image training set with building extraction result labels, inputs a target remote sensing image into the trained decoupled high-resolution network, obtains a boundary prediction result and a segmentation prediction result of the target remote sensing image, decomposes the features of the high-resolution remote sensing image, and respectively learns the building footprints and boundaries corresponding to the ground, so as to more accurately match the segmentation and boundaries in the real ground data.
[0063] Optionally, the decoupled high-resolution network includes a shallow feature extraction module, a first transition layer, a first multi-scale feature fusion module group, a second transition layer, a second multi-scale feature fusion module group and a prediction result output module connected in sequence;
[0064] The shallow feature extraction module is used to extract features from the input image to obtain first resolution features;
[0065] The first transition layer is used to generate a second resolution feature based on the first resolution feature;
[0066] The first multi-scale feature fusion module group is used to obtain output feature results at multiple scales based on the first resolution feature and the second resolution feature;
[0067] A second transition layer, used for generating a third resolution feature based on the first resolution feature and the second resolution feature after being fused by the first multi-scale feature fusion module group;
[0068] A second multi-scale feature fusion module group is used to obtain a multi-scale output feature result based on the first resolution feature, the second resolution feature and the third resolution feature after being fused by the first multi-scale feature fusion module group;
[0069] The prediction result output module is used to fuse and upsample the boundary features and segmentation features based on the output feature results of the first multi-scale feature fusion module group and the output feature results of the second multi-scale feature fusion module group, so as to obtain the boundary prediction result and segmentation prediction result of the input image.
[0070] Specifically, the decoupled high-resolution network provided by the present invention may include a shallow feature extraction module, a first transition layer, a first multi-scale feature fusion module group, a second transition layer, a second multi-scale feature fusion module group and a prediction result output module connected in sequence.
[0071] Figure 2 The schematic diagram of the structure of the decoupled high-resolution network provided by the present invention is as follows: Figure 2 As shown, the shallow feature extraction module includes three convolution blocks and four residual blocks, which extract features from the input image to obtain the first resolution features. It can be understood that if it is a trained decoupled high-resolution network, the input image is the target remote sensing image, and if it is a decoupled high-resolution network being trained, the input image is a sample remote sensing image in the sample remote sensing image training set.
[0072] It should be noted that Figure 2 This is only a possible example of the decoupled high-resolution network provided by the present invention, and the number of modules does not limit the solution provided by the present invention.
[0073] In some implementations, since rich spatial detail information is crucial to the complete extraction of building boundaries, after weighing complexity and accuracy, the downsampling multiples of the network during shallow feature extraction are reduced in the embodiments of the present invention, so that the highest resolution retained in the network is increased to half the size of the original input image (i.e., 1 / 2). Optionally, the number of feature map channels with the highest resolution can also be set to 48 to ensure the performance of the network.
[0074] The present invention hopes to obtain meaningful boundary detection and segmentation result outputs at different scales. However, the lowest resolution in the parallel branches generated by HRNet is only 1 / 32 of the original image, and the prediction results after interpolation upsampling often become blurred and contain less usable information. At the same time, this low-resolution representation may interfere with multi-scale information fusion. In order to reduce memory and computing costs, an embodiment of the present invention can reduce the number of parallel branches and retain only three parallel branches with different resolutions, namely, first resolution features, second resolution features, and third resolution features.
[0075] In some implementations, the three parallel branches of different resolutions used in the embodiments of the present invention are 1 / 2, 1 / 4 and 1 / 8, that is, the first resolution feature, the second resolution feature and the third resolution feature are 1 / 2 resolution feature, 1 / 4 resolution feature and 1 / 8 resolution feature, respectively.
[0076] The first resolution feature can be passed through the first transition layer (transition layer 1 in the figure) to generate a second resolution feature, and the first resolution feature and the second resolution feature are input into the first multi-scale feature fusion module group, and the first multi-scale feature fusion group includes multiple ( Figure 2 There are 2 stacked MFF modules in the figure, so as to obtain the output feature results of each MFF module.
[0077] Afterwards, the first resolution feature and the second resolution feature that have passed through the first multi-scale feature fusion module group can be passed through the second transition layer (transition layer 2 in the figure) to generate a third resolution feature, and the first resolution feature, the second resolution feature, and the third resolution feature are input into the second multi-scale feature fusion module group. The second multi-scale feature fusion group includes multiple ( Figure 2 There are 4 stacked MFF modules in the figure, so as to obtain the output feature results of each MFF module.
[0078] In some implementations, for the multi-scale feature fusion module in the second multi-scale feature fusion module group, the convolution block in the basic block is an expanded convolution block, thereby ensuring that the network has a sufficiently large receptive field, which can help the network obtain richer global information. When setting the expanded convolution parameters, you can refer to the Multi-grid coefficient of the series network in the DeepLabv3 network.
[0079] After obtaining the output feature results of the first multi-scale feature fusion module group and the output feature results of the second multi-scale feature fusion module group, the boundary features and segmentation features therein can be fused and upsampled respectively through the prediction result output module, and finally the boundary prediction result and segmentation prediction result of the input image are obtained.
[0080] Optionally, the feature decoupling and recoupling module includes:
[0081] Flow field generation module, feature decoupling module, feature enhancement module and feature fusion module;
[0082] The flow field generation module is used to calculate the flow field after splicing high-resolution features and low-resolution features;
[0083] The feature decoupling module is used to obtain boundary features and main features based on the flow field;
[0084] The feature enhancement module is used to repair and enhance the boundary features and the main features to obtain enhanced boundary features and enhanced main features respectively;
[0085] The feature fusion module is used to obtain a boundary information confidence map and a main information confidence map based on the enhanced boundary features and the enhanced main features, respectively, and to perform feature mosaicking based on the enhanced boundary features, the enhanced main features, the boundary information confidence map and the main information confidence map to obtain a fused feature;
[0086] Among them, the output feature results of the multi-scale feature fusion module are obtained based on the enhanced boundary features and the fusion features.
[0087] Specifically, the feature decoupling and recoupling module provided by the present invention may include a flow field generation module, a feature decoupling module, a feature enhancement module and a feature fusion module.
[0088] Figure 3 The structural diagram of the characteristic decoupling and recoupling module provided by the present invention is as follows: Figure 3 As shown, the flow field generation module calculates the flow field after splicing the high-resolution features and the low-resolution features.
[0089] It can be understood that the high-resolution feature is obtained based on the first resolution feature; the low-resolution feature is obtained based on the second resolution feature, or the low-resolution feature is obtained based on the second resolution feature and the third resolution feature. The low-resolution feature map extracted by the network often focuses on the most prominent part of the image, has an abstract semantic representation, but rarely contains boundary detail information. Therefore, the present invention hopes to learn a flow field that mainly points to the inside of the target object from the trend of the high-resolution to low-resolution feature map, and based on the flow field, the boundary information in the high-resolution feature map can be separated.
[0090] In the embodiment of the present invention, the high-resolution representation F maintained in the network is used hr As the high-resolution feature map of flow field learning, the remaining i low-resolution representations in the parallel structure Upsample to F hr The same size is used as the low-resolution feature map in flow field learning. For example, the high-resolution feature map is a 1 / 2 resolution feature, and the other two resolution features are a 1 / 4 resolution feature and a 1 / 8 resolution feature. The 1 / 4 resolution feature and the 1 / 8 resolution feature are sampled 2 times and 4 times respectively and then used as low-resolution features.
[0091] After concatenating the high-resolution features with the low-resolution features, a 3×3 convolution can be used to learn the flow field Φ. The process is shown in the following formula:
[0092]
[0093] in, It is represented by m convolution kernels of size k×k, concat(.) is the channel concatenation operation, Up(.) represents the upsampling operation, and n represents the number of repetitions of the feature decoupling and recoupling module.
[0094] After obtaining the flow field Φ, the feature decoupling module can obtain boundary features and main features based on the flow field.
[0095] Based on the flow field Φ, we can get the offset Φ(p) required for mapping each pixel on the feature map. Then, simply by p+Φ(p), we can map any point p on the original feature map to the feature map containing the main building information. On this basis, we can use the nearest neighbor interpolation to estimate each point p in the main feature. l Finally, we can get the main part (i.e. main feature) of the original feature map containing the overall structure and shape of the building. The whole process is shown in the following formula:
[0096]
[0097] Among them, ω warp represents the bilinear kernel weights on the warped spatial grid obtained by flow field calculation. S Φ (.) indicates p l The four nearest neighboring pixels, each pixel p is mapped by p+Φ(p). Since the coupled feature consists of the boundary part and the main part, the main part F can be subtracted from the original high-resolution feature map F. body , the remaining part is the feature F containing the boundary and details of the building boundary (i.e. boundary features), as shown below:
[0098] F boundary =FF body
[0099] After obtaining the boundary features and the main features, the feature enhancement module can repair and enhance the boundary features and the main features, thereby obtaining enhanced boundary features and enhanced main features respectively.
[0100] Optionally, the boundary features and the main features are repaired and enhanced to obtain enhanced boundary features and enhanced main features, respectively, which specifically includes:
[0101] Using the shallow feature or the enhanced boundary feature output by the previous feature decoupling and recoupling module as a supplement to the boundary feature, repairing the boundary feature to obtain the enhanced boundary feature; and,
[0102] The low-resolution features are used as a supplement to the main features, and the main features are repaired to obtain enhanced main features.
[0103] Specifically, the decoupling process may lead to the loss of global information. In the embodiment of the present invention, the feature map with lower resolution in the parallel branch can be used as a supplement to the feature map of the main information. By fusing multi-scale information, the high-level semantic information and context information contained in the main feature map are enhanced, thereby improving the integrity of the building. The repair process of the main feature is shown in the following formula:
[0104]
[0105] in, is the enhanced main feature after restoration, and C is the number of channels corresponding to the highest resolution feature in the parallel branch.
[0106] For the feature map of boundary information, we can combine the shallow features of the network or the enhanced boundary features output by the previous feature decoupling and recoupling module, and use dense connections to enhance it, so that the boundary information feature map can focus as much as possible on the boundary information mainly based on the building boundary. The repair process is shown in the following formula:
[0107]
[0108] in, represents the enhanced boundary features after the MFF module is repaired at the nth repetition, F shallow It is understood that when n=1, the boundary features are repaired by combining the shallow features of the network; and when n>1, the boundary features are repaired by combining the enhanced boundary features output by the previous feature decoupling and recoupling module.
[0109] Then, the feature fusion module can be used to obtain a boundary information confidence map and a main information confidence map according to the enhanced boundary features and the enhanced main features, respectively, and feature mosaicking is performed according to the enhanced boundary features, the enhanced main features, the boundary information confidence map and the main information confidence map to obtain the fused features.
[0110] In the embodiment of the present invention, a mosaic method based on subject-boundary attention is adopted to fuse the boundary feature map containing rich building edge details and the subject feature map containing the continuous and complete structure of the building.
[0111] Different from the simple addition method, the mosaic method assigns different fusion strengths to different regions by using the confidence maps of the main information and boundary information as a guide. This method can better fuse the boundary information and the main information to a certain extent, thus producing a feature map with stronger semantic information and less redundant information.
[0112] By enhancing the subject feature map and enhanced boundary feature map The subject information confidence map M can be obtained bodyand boundary information confidence map M boundary The specific process is as follows:
[0113]
[0114] Where Sig(.) represents the sigmoid function. body and M boundary The fusion feature F can be obtained as follows fuse :
[0115]
[0116] in, represents the dot product of the two matrices. It can be found that the first term in the equation pays more attention to the boundary area, especially the boundary part, and suppresses the influence of the main area; while the second term is the opposite, paying more attention to the main area. In this way, the interaction between the segmentation result and the boundary detection result can be achieved, and the segmentation result is constrained by the boundary extraction result, thereby promoting a better representation of the building features.
[0117] Finally, the output feature results of the multi-scale feature fusion module can be obtained according to the enhanced boundary features and the fusion features, and the boundary prediction results and the segmentation prediction results can be obtained by the prediction result output module respectively.
[0118] In some implementations, based on the fusion feature F fuse , the shortcut method can be used to supplement the feature information of each resolution. The process is as follows:
[0119]
[0120]
[0121] in, Indicates the highest resolution after correction. represents the corrected i-th low-resolution representation, and Down(.) represents the downsampling operation.
[0122] The supplemented feature F multi The method of obtaining is as follows:
[0123]
[0124] Among them, F multi In this case, the enhanced boundary feature can be used as the boundary feature in the output feature result, and the supplemented feature can be used as the segmentation feature.
[0125] Optionally, the training method of the decoupled high-resolution network includes:
[0126] The sample remote sensing images in the sample remote sensing image training set are input into the initialized decoupled high-resolution network to obtain the output feature results of each multi-scale feature fusion module and the building prediction results of the sample remote sensing images;
[0127] The output feature results of each multi-scale feature fusion module are deeply supervised, and the decoupled high-resolution network is trained based on the output feature results of each multi-scale feature fusion module, the building prediction results of the sample remote sensing image, the building extraction result labels of the sample remote sensing image and the loss function.
[0128] Specifically, according to the above introduction to the decoupled high-resolution network, in the process of training the decoupled high-resolution network, the sample remote sensing images in the sample remote sensing image training set are input into the initialized decoupled high-resolution network, and the output feature results of each multi-scale feature fusion module and the building prediction results of the sample remote sensing images can be obtained.
[0129] Deep supervised learning can accelerate the convergence of the model and improve the expressive power of each layer of the network. It is widely used in building extraction tasks. Therefore, the output feature results of each multi-scale feature fusion module can be deeply supervised, and the decoupled high-resolution network can be trained based on the output feature results of each multi-scale feature fusion module, the building prediction results of the sample remote sensing image, the building extraction result labels of the sample remote sensing image, and the loss function.
[0130] The feature decoupling and recoupling module generates enhanced boundary features, enhanced main features and fusion features at each stage. In order to ensure that the network has sufficient flexibility and expressiveness, better integrate multi-scale information, and ensure the continuity and integrity of the extracted boundaries, in the embodiment of the present invention, only the prediction results of the enhanced boundary features and fusion features (or supplemented features) are deeply supervised at each stage. The following is an example of the deep supervision process using the supplemented features as an example.
[0131] The convolution operation can be used to extract the boundary prediction results of each MFF module from the boundary feature map and the fused feature map. And the segmentation results The specific process is as follows:
[0132]
[0133] in, and represents the loss corresponding to the decoupling-recoupling module at the Nth repetition, L(.) represents the related loss function, and Corresponding to the true value labels of building boundaries and segmentation results respectively.
[0134] Finally, the prediction results of boundaries and segmentation generated at each stage are fused to help the network learn how to combine prediction results from different stages, thereby further improving the performance and accuracy of building extraction. The process is as follows:
[0135]
[0136] in, Represents the corresponding loss after the boundary prediction results are fused, Represents the corresponding loss after the segmentation results are fused.
[0137] Regarding the setting of the relevant loss function L, the present invention adopts binary cross entropy loss and Dice loss to jointly optimize the model's learning of segmentation and boundaries, in the hope that the network can more comprehensively understand the structure and boundaries of the building and improve its performance in practical applications.
[0138]
[0139] in, Represents the prediction result, y represents the true value label, and binary cross entropy, as a common loss function for segmentation, can effectively measure the difference between the prediction result and the true value label of each pixel, while the Dice loss function can evaluate the overlap between the prediction map and the true value label, and can alleviate the problem of class imbalance to a certain extent. The detailed formulas of the two are as follows:
[0140]
[0141] Where N represents the number of pixels in the prediction result, i represents the i-th pixel, and ω is the relevant weight setting. Considering that the building boundary area usually accounts for only a small proportion in the image, in order to balance the weight of the loss function, the present invention sets the weight in the boundary detection task to the ratio of positive and negative samples in the training set. In the semantic segmentation task, the present invention sets the weight to 1.
[0142]
[0143] Wherein, ∈ is a Laplace smoothing term, which is used to avoid the denominator being zero, and the present invention sets it to 1. Considering that the prediction of segmentation and edges gradually changes from coarse to fine, the present invention sets weights of different sizes for the loss function of deep supervision, and gradually increases with the depth of the network.
[0144] Finally, the total loss function of the present invention is the weighted sum of all losses, as follows:
[0145]
[0146] Among them, ωn =[0.3, 0.3, 0.5, 0.5, 0.5, 0.5] is the weight of the deep supervision output at different stages. The decoupled high-resolution network can be trained based on the loss function, the output feature results of each multi-scale feature fusion module, the building prediction results of the sample remote sensing image, and the building extraction result labels of the sample remote sensing image.
[0147] The following is a description of a high-resolution building proposal device based on feature decoupling and recoupling provided by the present invention. The high-resolution building proposal device based on feature decoupling and recoupling described below and the high-resolution building proposal method based on feature decoupling and recoupling described above can be referenced to each other.
[0148] Figure 4 The schematic diagram of the structure of the high-resolution building extraction device based on feature decoupling and recoupling provided by the present invention is as follows: Figure 4 As shown, the device comprises:
[0149] An input module 400 is used to input the target remote sensing image into the trained decoupled high-resolution network to obtain a building prediction result of the target remote sensing image;
[0150] An output module 410, used to use the building prediction result of the target remote sensing image as the building extraction result of the target remote sensing image;
[0151] Among them, the building prediction results of the target remote sensing image include the boundary prediction results of the target remote sensing image and the segmentation prediction results of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained based on a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block.
[0152] Optionally, the decoupled high-resolution network includes a shallow feature extraction module, a first transition layer, a first multi-scale feature fusion module group, a second transition layer, a second multi-scale feature fusion module group and a prediction result output module connected in sequence;
[0153] The shallow feature extraction module is used to extract features from the input image to obtain first resolution features;
[0154] The first transition layer is used to generate a second resolution feature based on the first resolution feature;
[0155] The first multi-scale feature fusion module group is used to obtain output feature results at multiple scales based on the first resolution feature and the second resolution feature;
[0156] A second transition layer, used for generating a third resolution feature based on the first resolution feature and the second resolution feature after being fused by the first multi-scale feature fusion module group;
[0157] A second multi-scale feature fusion module group is used to obtain a multi-scale output feature result based on the first resolution feature, the second resolution feature and the third resolution feature after being fused by the first multi-scale feature fusion module group;
[0158] The prediction result output module is used to fuse and upsample the boundary features and segmentation features based on the output feature results of the first multi-scale feature fusion module group and the output feature results of the second multi-scale feature fusion module group, so as to obtain the boundary prediction result and segmentation prediction result of the input image.
[0159] Optionally, the feature decoupling and recoupling module includes:
[0160] Flow field generation module, feature decoupling module, feature enhancement module and feature fusion module;
[0161] The flow field generation module is used to calculate the flow field after splicing high-resolution features and low-resolution features;
[0162] The feature decoupling module is used to obtain boundary features and main features based on the flow field;
[0163] The feature enhancement module is used to repair and enhance the boundary features and the main features to obtain enhanced boundary features and enhanced main features respectively;
[0164] The feature fusion module is used to obtain a boundary information confidence map and a main information confidence map based on the enhanced boundary features and the enhanced main features, respectively, and to perform feature mosaicking based on the enhanced boundary features, the enhanced main features, the boundary information confidence map and the main information confidence map to obtain a fused feature;
[0165] Among them, the output feature results of the multi-scale feature fusion module are obtained based on the enhanced boundary features and the fusion features.
[0166] Optionally, the boundary features and the main features are repaired and enhanced to obtain enhanced boundary features and enhanced main features, respectively, which specifically includes:
[0167] Using the shallow feature or the enhanced boundary feature output by the previous feature decoupling and recoupling module as a supplement to the boundary feature, repairing the boundary feature to obtain the enhanced boundary feature; and,
[0168] The low-resolution features are used as a supplement to the main features, and the main features are repaired to obtain enhanced main features.
[0169] Optionally, the training method of the decoupled high-resolution network includes:
[0170] The sample remote sensing images in the sample remote sensing image training set are input into the initialized decoupled high-resolution network to obtain the output feature results of each multi-scale feature fusion module and the building prediction results of the sample remote sensing images;
[0171] The output feature results of each multi-scale feature fusion module are deeply supervised, and the decoupled high-resolution network is trained based on the output feature results of each multi-scale feature fusion module, the building prediction results of the sample remote sensing image, the building extraction result labels of the sample remote sensing image and the loss function.
[0172] Optionally, for the multi-scale feature fusion module in the second multi-scale feature fusion module group, the convolution block in the basic block is a dilated convolution block.
[0173] Figure 5 A schematic diagram of the structure of an electronic device provided by the present invention, such as Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the high-resolution building extraction method based on feature decoupling and recoupling provided by the above methods, including:
[0174] Input the target remote sensing image into the trained decoupled high-resolution network to obtain the building prediction result of the target remote sensing image;
[0175] The building prediction result of the target remote sensing image is used as the building extraction result of the target remote sensing image;
[0176] Among them, the building prediction results of the target remote sensing image include the boundary prediction results of the target remote sensing image and the segmentation prediction results of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained based on a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block.
[0177] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0178] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program, the computer program can be stored in a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the high-resolution building extraction method based on feature decoupling and recoupling provided by the above methods, including:
[0179] Input the target remote sensing image into the trained decoupled high-resolution network to obtain the building prediction result of the target remote sensing image;
[0180] The building prediction result of the target remote sensing image is used as the building extraction result of the target remote sensing image;
[0181] Among them, the building prediction results of the target remote sensing image include the boundary prediction results of the target remote sensing image and the segmentation prediction results of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained based on a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block.
[0182] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the high-resolution building extraction method based on feature decoupling and recoupling provided by the above methods, including:
[0183] Input the target remote sensing image into the trained decoupled high-resolution network to obtain the building prediction result of the target remote sensing image;
[0184] The building prediction result of the target remote sensing image is used as the building extraction result of the target remote sensing image;
[0185] Among them, the building prediction results of the target remote sensing image include the boundary prediction results of the target remote sensing image and the segmentation prediction results of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained based on a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block.
[0186] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0187] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A high-resolution building extraction method based on feature decoupling and recoupling, characterized in that: include: Inputting the target remote sensing image into the trained decoupled high-resolution network to obtain a building prediction result of the target remote sensing image; Using the building prediction result of the target remote sensing image as the building extraction result of the target remote sensing image; The building prediction result of the target remote sensing image includes the boundary prediction result of the target remote sensing image and the segmentation prediction result of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained based on a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block; The decoupled high-resolution network includes a shallow feature extraction module, a first transition layer, a first multi-scale feature fusion module group, a second transition layer, a second multi-scale feature fusion module group and a prediction result output module connected in sequence; The shallow feature extraction module is used to extract features from the input image to obtain first resolution features; The first transition layer is used to generate a second resolution feature based on the first resolution feature; The first multi-scale feature fusion module group is used to obtain output feature results at multiple scales based on the first resolution feature and the second resolution feature; The second transition layer is used to generate a third resolution feature based on the first resolution feature and the second resolution feature after the first multi-scale feature fusion module group; The second multi-scale feature fusion module group is used to obtain a multi-scale output feature result based on the first resolution feature, the second resolution feature and the third resolution feature after being fused by the first multi-scale feature fusion module group; The prediction result output module is used to fuse and upsample the boundary features and segmentation features based on the output feature results of the first multi-scale feature fusion module group and the output feature results of the second multi-scale feature fusion module group, so as to obtain the boundary prediction result and segmentation prediction result of the input image.
2. The high-resolution building extraction method based on feature decoupling and recoupling according to claim 1 is characterized in that: The characteristic decoupling and recoupling module comprises: Flow field generation module, feature decoupling module, feature enhancement module and feature fusion module; The flow field generation module is used to calculate and obtain the flow field after splicing the high-resolution features and the low-resolution features; The feature decoupling module is used to obtain boundary features and main features respectively based on the flow field; The feature enhancement module is used to repair and enhance the boundary feature and the main feature to obtain an enhanced boundary feature and an enhanced main feature respectively; The feature fusion module is used to obtain a boundary information confidence map and a main information confidence map based on the enhanced boundary feature and the enhanced main feature, respectively, and perform feature mosaicking based on the enhanced boundary feature, the enhanced main feature, the boundary information confidence map and the main information confidence map to obtain a fused feature; Among them, the output feature result of the multi-scale feature fusion module is obtained according to the enhanced boundary feature and the fusion feature.
3. The high-resolution building extraction method based on feature decoupling and recoupling according to claim 2 is characterized in that: The repairing and enhancing of the boundary features and the main features to obtain enhanced boundary features and enhanced main features respectively specifically includes: Using the shallow feature or the enhanced boundary feature output by the previous feature decoupling and recoupling module as a supplement to the boundary feature, repairing the boundary feature to obtain the enhanced boundary feature; and, The low-resolution feature is used as a supplement to the main feature, and the main feature is repaired to obtain the enhanced main feature.
4. The high-resolution building extraction method based on feature decoupling and recoupling according to claim 1 is characterized in that: The training method of the decoupled high-resolution network includes: Inputting the sample remote sensing images in the sample remote sensing image training set into the initialized decoupled high-resolution network to obtain the output feature results of each multi-scale feature fusion module and the building prediction results of the sample remote sensing images; The output feature results of each multi-scale feature fusion module are deeply supervised, and the decoupled high-resolution network is trained based on the output feature results of each multi-scale feature fusion module, the building prediction results of the sample remote sensing image, the building extraction result labels of the sample remote sensing image and the loss function.
5. The high-resolution building extraction method based on feature decoupling and recoupling according to claim 1 is characterized in that: For the multi-scale feature fusion module in the second multi-scale feature fusion module group, the convolution block in the basic block is a dilated convolution block.
6. A high-resolution building extraction device based on feature decoupling and recoupling, characterized in that: include: An input module, used for inputting a target remote sensing image into a trained decoupled high-resolution network to obtain a building prediction result of the target remote sensing image; An output module, used for using the building prediction result of the target remote sensing image as the building extraction result of the target remote sensing image; The building prediction result of the target remote sensing image includes the boundary prediction result of the target remote sensing image and the segmentation prediction result of the target remote sensing image; the decoupled high-resolution network is a parallel structure, which is composed of a shallow feature extraction module, a transition layer, multiple multi-scale feature fusion modules and a prediction result output module, and is trained based on a sample remote sensing image training set with building extraction result labels; the multi-scale feature fusion module is composed of a feature decoupling and recoupling module and a basic block; The decoupled high-resolution network includes a shallow feature extraction module, a first transition layer, a first multi-scale feature fusion module group, a second transition layer, a second multi-scale feature fusion module group and a prediction result output module connected in sequence; The shallow feature extraction module is used to extract features from the input image to obtain first resolution features; The first transition layer is used to generate a second resolution feature based on the first resolution feature; The first multi-scale feature fusion module group is used to obtain output feature results at multiple scales based on the first resolution feature and the second resolution feature; The second transition layer is used to generate a third resolution feature based on the first resolution feature and the second resolution feature after the first multi-scale feature fusion module group; The second multi-scale feature fusion module group is used to obtain a multi-scale output feature result based on the first resolution feature, the second resolution feature and the third resolution feature after being fused by the first multi-scale feature fusion module group; The prediction result output module is used to fuse and upsample the boundary features and segmentation features based on the output feature results of the first multi-scale feature fusion module group and the output feature results of the second multi-scale feature fusion module group, so as to obtain the boundary prediction result and segmentation prediction result of the input image.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the high-resolution building extraction method based on feature decoupling and recoupling as described in any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the high-resolution building extraction method based on feature decoupling and recoupling as claimed in any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the high-resolution building extraction method based on feature decoupling and recoupling as claimed in any one of claims 1 to 5 is implemented.