Building hierarchical interaction enhancement extraction method based on improved U-Net

By introducing DPM, HIM and SFM modules into the U-Net network, the problems of loss of details and insufficient interactivity in the building extraction method are solved, the accuracy of building classification and extraction is improved, and the feature diversity is enhanced.

CN120236195APending Publication Date: 2025-07-01CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510321252.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing building extraction methods based on U-Net networks have lost details and insufficient interactivity between different levels, resulting in a reduced accuracy of building classification and extraction.

Method used

DPM module, HIM module and SFM module are introduced in the U-Net network. The DPM module is used to reduce the information loss caused by continuous downsampling, the HIM module is used to interact information in the feature interaction and fusion stage, and the SFM module is used to adaptive selection of features at different levels.

Benefits of technology

By introducing these modules, the accuracy of building classification and extraction is improved, the details of spatial distribution are preserved, and the diversity of building characteristics is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236195A_ABST
    Figure CN120236195A_ABST
Patent Text Reader

Abstract

The invention discloses a building hierarchical interaction enhancement extraction method based on improved U-Net. The method comprises the following steps: collecting building remote sensing images to construct a data set; an improved U-net network is constructed; inputting the data set into the improved U-net network to obtain feature maps of different stages, and outputting a final fusion feature map through interaction modules of different levels; and S3, optimizing the fused feature map in the step S3 through a loss function. According to the building hierarchical interaction enhancement extraction method based on the improved U-Net, the DPM module is introduced to reduce information loss caused by continuous downsampling of the U-Net network, and the accuracy of building classification and extraction is improved. The HIM module is introduced for information interaction, and the SFM module is used for adaptive selection of different hierarchical features, so that details of spatial distribution are reserved, and diversity of building features is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of building extraction, and particularly relates to an enhanced extraction method for hierarchical interaction of buildings based on an improved U-Net. Background Art

[0002] Due to production and living needs, buildings have different characteristics in terms of space, shape, and structure. With the continuous expansion of cities, the number of buildings in cities is increasing continuously. The accurate classification and extraction of buildings have important application values in urban planning, disaster rescue, environmental monitoring, and population estimation.

[0003] Among the existing building recognition and extraction technologies, the U-Net network is currently the most commonly used building extraction method. However, the building extraction method based on the U-Net network has problems such as loss of details and insufficient interaction between different levels, resulting in a decrease in the accuracy of building classification and extraction. Summary of the Invention

[0004] The technical problem of the present invention is to propose an enhanced extraction method for hierarchical interaction of buildings based on an improved U-Net, which reduces the loss of details caused by continuous downsampling, improves the insufficient interaction between features at different levels, and improves the accuracy of building classification and extraction.

[0005] The object of the present invention is to solve the above problems. The technical solution adopted by the present invention is: an enhanced extraction method for hierarchical interaction of buildings based on an improved U-Net, which introduces a DPM module, a HIM module, and an SFM module into the U-net network; the DPM module is used to reduce the information loss caused by continuous downsampling; the HIM module is used for information interaction in the feature interaction and fusion stage; the SFM module is used for the adaptive selection of features at different levels; It includes the following steps: S1: Collect remote sensing images of buildings to construct a data set; S2: Construct an improved U-net network; S3: Input the data set into the improved U-net network to obtain feature maps at different stages, and output the final fused feature map through interaction modules at different levels; S4: Optimize the fused feature map in step S3 through a loss function.

[0006] Further, in step S1, it also includes selecting remote sensing images of buildings in different time periods and regions, and performing preprocessing such as shearing and screening, and then dividing the data set into a training set and a test set according to a fixed ratio.

[0007] Further, in step S2, the improved U-net network includes an encoder part and a downsampling part; the DPM module is introduced into the encoder part; the HIM module is introduced into both the encoder part and the downsampling part.

[0008] Preferably, in the encoder part, a series of convolutional and pooling operations are performed to convert the input image into high-level features, and the image information is compressed through a dimensionality reduction operation to obtain feature maps at different stages.

[0009] Further, step S2 includes the following sub-steps: 1) Input the feature map into the DPM module, and perform max pooling and space-to-depth convolution through two branches respectively. The calculation formula is: ; ; ; In the formula, X represents the input feature map, represents the max pooling operation, represents the space-to-depth convolution operation, f1 represents the output after the max pooling operation of the first branch, represents the output after the depth convolution operation of the second branch, f1 and f2 , w represents the width of the feature map, c represents the number of channels of the feature map, h represents the width of the feature map, and R represents the real numbers in this feature map; 2) Analyze the space-to-depth convolution module : The input feature map X is divided into 4 sub-regions, and a 1×1 convolution operation is performed on each sub-region to capture local features, and the output features are concatenated to form a feature map, and then the final output is obtained by performing a convolution operation. The calculation formula is: ; ; ; In the formula, represents passing through a 2×2 merging block, and separating the pixels in the 2×2 spatial region of the input feature map according to the position into i feature maps, i represents the calculation unit, i = 1, 2, 3, 4; represents the 1×1 convolution operation, represents the concatenation operation, represents the feature map obtained from left to right and top left to bottom after being processed, , X1 top left, X2 top right, X3 bottom left, X4 bottom right, represents for after passing through The feature map after operation, ; It represents the feature after enhancing the input feature X through the spatial-to-depth convolution module, ; 3) Convert the output feature into a weight map using an activation function, and perform an element-wise addition operation on the output features of the branches. The calculation formula is: ; In the formula, represents the sigmoid activation function, represents the feature after the input feature is processed by max pooling, represents the feature processed by the SPD-Conv branch, represents the feature processed by the DPM module, , represents the multiplication operation.

[0010] Preferably, the hierarchical interaction module HIM includes respectively inputting adjacent features in the feature extraction stage into the hierarchical interaction module HIM, outputting feature data with different resolutions, and connecting them to obtain the final feature.

[0011] Preferably, inputting adjacent features into the hierarchical interaction module HIM includes performing downsampling, upsampling, and selective fusion module SFM operations on the feature map input into the hierarchical interaction module HIM. The calculation formula is: ; ; ; ; ; In the formula, X1 and X2 represent the input feature maps, represents the downsampling operation, represents the upsampling operation, represents the selective fusion module, and S1 represents the input feature map is downsampled to ensure the same size as X2, , and S2 represents the feature map after upsampling after the feature maps X1 and X2 are subjected to the selective fusion module operation, , and S3 represents the fused feature map after the feature maps X1 and X2 are subjected to the selective fusion module SFM operation, .

[0012] Further, the input of the Selective Fusion Module (SFM) is two different types of features, and the two features are selectively fused; the SFM sums the input features and obtains the weight coefficients of the relevant channels of different feature maps through global average pooling, a fully connected layer, and softmax activation; then, selective fusion is performed according to the weight coefficients, and the calculation formula is: ; ; ; ; In the formula, GA represents the global average pooling operation, represents the fully connected layer operation, represents the activation function, represents the two different types of shallow features of the input, represents the two different types of deep features of the input, represents the fused features of the shallow and deep layers after global average pooling, , represents the fused features after passing through the fully connected layer, , represents the weight coefficients of the relevant channels of different feature maps obtained after softmax activation processing of .

[0013] Further, in step S4, which is included in the feature interaction and fusion stage, upsampling, concatenation, and 1×1 convolution operations are respectively performed, and then the feature maps in the feature extraction stage are fused to generate a prediction probability map for building extraction.

[0014] Preferably, step S3 further includes performing upsampling, DPM, and Concat operations on the feature map generated after the above operations to obtain the final prediction map, and the calculation formula is: ; In the formula, represents 1×1 convolution, represents the concatenation of feature maps, represents the final fused feature map of the shallow and deep layers processed by HIM in the decoding part, represents the feature map of the shallowest layer in the encoding part processed by the DPM module, and Out represents the final prediction map, .

[0015] Compared with the prior art, the beneficial effects of the present invention include: 1) The building hierarchical interaction enhancement extraction method based on the improved U-Net proposed by the present invention introduces a DPM module to reduce the information loss caused by continuous downsampling of the U-net network, and improves the accuracy of building classification and extraction.

[0016] 2) The building hierarchical interaction enhancement extraction method based on the improved U-Net proposed by the present invention introduces information interaction of the HIM module, and uses the SFM module for adaptive selection of features at different levels, retains the details of the spatial distribution, and strengthens the diversity of building features. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below in conjunction with the drawings and embodiments.

[0018] Figure 1 It is the network structure diagram of the building hierarchical interaction enhancement extraction based on the improved U-Net in the embodiment of the present invention; Figure 2 It is the structure diagram of the DPM module based on the improved U-Net in the embodiment of the present invention; Figure 3 It is the structure diagram of the SPD_Conv module based on the improved U-Net in the embodiment of the present invention; Figure 4 It is the structure diagram of the HIM module based on the improved U-Net in the embodiment of the present invention; Figure 5 It is the structure diagram of the SFM module based on the improved U-Net in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] As Figure 1 shown, a building hierarchical interaction enhancement extraction method based on the improved U-Net introduces a DPM module, a HIM module and an SFM module into the U-net network; the DPM module is used to reduce the information loss caused by continuous downsampling; the HIM module is used for information interaction in the feature interaction and fusion stage; the SFM module is used for adaptive selection of features at different levels; It includes the following steps: S1: Collect building remote sensing images to construct a data set; In step S1, it also includes selecting building remote sensing images in different time periods and regions, and performing preprocessing such as shearing and screening, and then dividing the data set into a training set and a test set according to a fixed ratio.

[0020] As Figure 2 and Figure 3 shown, S2: Construct an improved U-net network; In step S2, the improved U-net network includes an encoder part and a downsampling part; the DPM module is introduced into the encoder part; and the HIM module is introduced into both the encoder part and the downsampling part.

[0021] The encoder part includes a series of convolution and pooling operations to convert the input image into high-level features, and compresses the image information through a dimensionality reduction operation to obtain feature maps at different stages.

[0022] Further, step S2 includes the following sub-steps: 1) Input the feature map into the DPM module, and calculate through max pooling and space-to-depth convolution of two branches respectively. The calculation formula is: ; ; ; In the formula, X represents the input feature map, represents the max pooling operation, represents the space-to-depth convolution operation, f1 represents the output after the max pooling operation of the first branch, represents the output after the depth convolution operation of the second branch, f1 and f2 , w represents the width of the feature map, c represents the number of channels of the feature map, h represents the width of the feature map, and R represents the real numbers in this feature map; 2) Analyze the space-to-depth convolution module The input feature map X is divided into 4 sub-regions through, and each sub-region is processed independently to increase the spatial information of the features; for each divided region, a 1×1 convolution operation is performed to capture local features, and the different convolution output features are concatenated to form a detailed feature map that retains the spatial distribution and diversity, and then the concatenated feature map is further processed by performing a 1×1 convolution operation again to obtain the final output. The calculation formula is: ; ; ; In the formula, represents passing through a 2×2 merging block, and separating the pixels in the 2×2 spatial region of the input feature map into i feature maps, i represents the calculation unit, i = 1, 2, 3, 4; represents the 1×1 convolution operation, represents the concatenation operation, represents being The feature map obtained after processing, from left to right and from top left to bottom , X1 at the upper left corner, X2 at the upper right corner, X3 at the lower left corner, X4 at the lower right corner, denotes the feature map after going through the operation, ; denotes the feature after enhancing the input feature X through the spatial-to-depth convolution module, ; 3) Convert the output feature to a weight map using an activation function, and perform an element-wise addition operation on the output features of the branches. The calculation formula is: ; In the formula, denotes the sigmoid activation function, denotes the feature after the input feature undergoes max pooling processing, denotes the feature after being processed by the SPD-Conv branch, denotes the feature after being processed by the DPM module, , denotes the multiplication operation.

[0023] The hierarchical interaction module HIM includes respectively inputting adjacent features in the feature extraction stage into the hierarchical interaction module HIM, outputting feature data with different resolutions, and connecting them to obtain the final feature.

[0024] Inputting adjacent features into the hierarchical interaction module HIM includes performing downsampling, upsampling, and selective fusion module SFM operations on the feature map input into the hierarchical interaction module HIM. The calculation formula is: ; ; ; ; ; In the formula, X1 and X2 denote the input feature maps, denotes the downsampling operation, denotes the upsampling operation, denotes the selective fusion module, and S1 denotes the input feature map after downsampling to ensure the same size as X2, , S2 denotes the feature map after upsampling after the selective fusion module operation on the feature maps X1 and X2, , S3 denotes the fused feature map after the selective fusion module SFM operation on the feature maps X1 and X2, 。

[0025] The input of the Selective Fusion Module (SFM) is two different types of features, and it selectively fuses the two features. The SFM sums the input features and obtains the weight coefficients of the relevant channels of different feature maps through global average pooling, fully connected layers, and softmax activation. Then, selective fusion is performed according to the weight coefficients, and the calculation formula is: ; ; ; ; In the formula, GA represents the global average pooling operation, represents the fully connected layer operation, represents the activation function, represents the two different types of shallow features of the input, represents the two different types of deep features of the input, represents the fused features of the shallow and deep layers after global average pooling, , represents the fused features after passing through the fully connected layer, , represents for the weight coefficients of the relevant channels of different feature maps obtained after softmax activation processing, 。

[0026] S3: Input the dataset into the improved U-net network to obtain feature maps at different stages, and output the final fused feature map through interaction modules at different levels; As shown in Figure 4 and Figure 5 , in step S3, during the feature interaction and fusion stage, upsampling, concatenation, and 1×1 convolution operations are respectively performed, and then the feature maps in the feature extraction stage are fused to generate the prediction probability map for building extraction.

[0027] Step S3 also includes performing upsampling, DPM, and Concat operations on the feature maps generated after the above operations to obtain the final prediction map, and the calculation formula is: ; In the formula, represents 1×1 convolution, represents the concatenation of feature maps, represents the final fused feature map of the shallow and deep layers processed by HIM in the decoding part, It represents the feature map processed by the DPM module at the shallowest layer of the encoding part, and Out represents the final prediction map. 。

[0028] S4: Optimize the fused feature map in step S3 through a loss function.

[0029] To verify the specific effect of the designed hierarchical interaction network structure for building extraction in the present invention, the trained model is experimented on the WHU dataset. The WHU building dataset comes from the GPCV team of Wuhan University and consists of satellite and aerial datasets. The experimental data uses the aerial dataset, with an image size of 512 × 512 pixels and a spatial resolution of 0.3m. In this study, according to the original resolution, each image is further cropped into smaller segments of 256 × 256 pixels without overlap. Then, the parts in the image and annotation that do not contain buildings are removed to obtain the final data.

[0030] To evaluate the method we proposed, we use precision (Pre), recall (Rec), and intersection over union (IoU) as evaluation metrics. The higher the precision value, the better the model performance. The higher the recall, the lower the missed detection rate of the model. The higher the IoU, the better the localization accuracy of the model. The experimental environment is as follows: Windows 10 operating system, 16GB of memory, Ryzen 5 5600G CPU, and NVIDIA GeForce GTX 1660 graphics card. The programming language is Python 3.8, and the deep learning framework is Pytorch 1.12. The test results are compared with the current mainstream building extraction models, and the results on the WHU dataset are shown in Table 1: Table 1 Comparison of Results on the WHU Dataset

[0031] In the present invention, the precision value, recall, and intersection all reach the maximum values, which are 95.36%, 95.28%, and 91.05% respectively. It shows more excellent performance compared with the advanced algorithms of many classic building extraction networks.

[0032] The above description of the embodiments is only for facilitating the understanding and application of the present invention by those skilled in the art. Those skilled in the art can make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by researchers according to the hints of the present invention should be within the protection scope of the present invention.

Claims

1. A hierarchical interactive enhancement extraction method for buildings based on improved U-Net, characterized in that: The DPM module, HIM module and SFM module are introduced into the U-net network; the DPM module is used to reduce the information loss caused by continuous downsampling; the HIM module is used for information interaction in the feature interaction fusion stage; the SFM module is used for adaptive selection of features at different levels; The following steps are involved: S1: Collect remote sensing images of buildings to construct a dataset; S2: Build an improved U-net network; S3: Input the data set into the improved U-net network to obtain feature maps at different stages, and output the final fused feature map through interactive modules at different levels; S4: Optimize the fused feature map of step S3 through the loss function.

2. According to the improved U-Net-based building hierarchical interactive enhancement extraction method of claim 1, it is characterized in that: Step S1 also includes selecting remote sensing images of buildings in different time periods and regions, and performing cropping and screening preprocessing, and then dividing the data set into a training set and a test set according to a fixed ratio.

3. According to the improved U-Net-based building hierarchical interactive enhancement extraction method of claim 1, it is characterized in that: In step S2, the improved U-net network includes an encoder part and a downsampling part; the encoder part introduces a DPM module; and the HIM module is introduced into the encoder part and the downsampling part.

4. According to claim 3, a building hierarchical interactive enhancement extraction method based on improved U-Net is characterized in that: The encoder part includes a series of convolution and pooling operations to convert the input image into high-level features, and compresses the image information through dimensionality reduction operations to obtain feature maps at different stages.

5. According to claim 4, a building hierarchical interactive enhancement extraction method based on improved U-Net is characterized in that: Step S2 includes the following sub-steps: 1) Input the feature map into the DPM module, and pass through the maximum pooling and spatial to depth convolution of the two branches respectively. The calculation formula is: ; ; ; In the formula, X represents the input feature map, represents the max pooling operation, represents the spatial to depth convolution operation, f1 represents the output of the first branch after the maximum pooling operation, Represents the output of the second branch through the depthwise convolution operation, f1 and f2 , w represents the feature map width, c represents the number of feature map channels, h represents the feature map width, and R represents the real number in this feature map; 2) For the spatial to depth convolution module Analysis: The input feature map X is divided into 4 sub-regions, and a 1×1 convolution operation is performed on each sub-region to capture local features. The output features are concatenated to form a feature map, and the final output is obtained by performing a convolution operation. The calculation formula is: ; ; ; In the formula, Indicates that the pixels of the 2×2 spatial region of the input feature map are separated into i In the feature map, i Indicates the unit of calculation, i =1,2,3,4; represents a 1×1 convolution operation, Represents a splicing operation, Indicates being The feature map obtained after processing from left to right, top left to bottom, , X1 upper left corner, X2 upper right corner, X3 lower left corner, X4 lower right corner, Express go through The feature map after operation, ; It represents the feature after the input feature X is enhanced by the spatial to depth convolution module. ; 3) Convert the output features into a weight map using an activation function, and perform element-by-element addition of the output features of the branches. The calculation formula is: ; In the formula, represents the sigmoid activation function, Represents the features after the input features are processed by the maximum pooling. represents the features processed by the SPD-Conv branch, Indicates the features processed by the DPM module. , Represents a multiplication operation.

6. According to claim 5, a building hierarchical interactive enhancement extraction method based on improved U-Net is characterized in that: The hierarchical interaction module HIM includes inputting adjacent features in the feature extraction stage into the hierarchical interaction module HIM respectively, outputting feature data with different resolutions, and obtaining final features after connection.

7. According to claim 6, a building hierarchical interactive enhancement extraction method based on improved U-Net is characterized in that: The adjacent features are input to the hierarchical interaction module HIM, including the feature map input to the hierarchical interaction module HIM for downsampling, upsampling and selective fusion module SFM operations, and the calculation formula is: ; ; ; ; ; Where X1 and X2 represent the input feature maps, represents the downsampling operation, represents the upsampling operation, represents the selective fusion module, S1 represents the input feature map Downsample to ensure that it is the same size as X2. , S2 represents the selective fusion module of feature maps X1 and X2 The feature map after upsampling after operation, , S3 represents the fusion feature map after the feature map X1 and X2 are operated by the selective fusion module SFM, .

8. The building hierarchical interactive enhancement extraction method based on improved U-Net according to claim 7 is characterized in that: The input of the selective fusion module SFM is two different types of features, and the two features are selectively fused; the selective fusion module SFM sums the input features, and obtains the weight coefficients of the relevant channels of different feature maps through global average pooling, fully connected layer and softmax activation; and then selective fusion is performed according to the weight coefficient, and the calculation formula is: ; ; ; ; In the formula, GA represents the global average pooling operation, represents the fully connected layer operation, represents the activation function, Represents the input of two types of shallow features, Represents the input of two types of deep features, Represents the fusion features of the shallow and deep layers after global average pooling. , represents the fusion features after being processed by the fully connected layer, , Express The weight coefficients of the relevant channels of different feature maps obtained after softmax activation processing, .

9. The building hierarchical interactive enhancement extraction method based on improved U-Net according to claim 1, characterized in that: The step S3 includes performing upsampling, cascading, and 1×1 convolution operations in the feature interaction fusion stage, and then fusing the feature maps from the feature extraction stage to generate a prediction probability map for building extraction.

10. The method for extracting hierarchical interactive enhancement of buildings based on improved U-Net according to claim 9, characterized in that: The step S3 further includes performing upsampling, DPM and Concat operations on the feature map generated after the above operations to obtain the final prediction map, and the calculation formula is: ; In the formula, represents 1×1 convolution, Indicates the concatenation of feature maps. It represents the final fusion feature map of the shallow and deep layers after HIM processing in the decoding part. It represents the feature map of the shallowest layer of the encoding part after being processed by the DPM module, and Out represents the final prediction map. .