Intelligent Building Extraction Method Based on Remote Sensing Images
By designing a semantic segmentation network of remote sensing images based on coding-decoding structure, combining global-local feature fusion and multi-scale dense connection hollow convolution attention module, the problem of poor building segmentation performance in remote sensing images is solved, and efficient building extraction and segmentation effects are achieved.
Patent Information
- Application Number
- CN202411502090.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-25
AI Technical Summary
The prior art has challenges in building segmentation in remote sensing images, including large changes in building size, long-tail distribution of roof types and heights, large differences between classes, small differences within classes, and the presence of interference factors such as shadows, shading and clouds, resulting in poor segmentation performance.
A semantic segmentation network of remote sensing image building based on encoding-decoding structure is designed, combining semantic segmentation encoder and decoder, and a global-local feature fusion module and a multi-scale densely connected hollow convolution attention module are introduced. Through multi-task joint optimization of semantic segmentation and DSM estimation tasks, the information complementarity between semantic features and high features is achieved.
It effectively improves the extraction performance of buildings in remote sensing images, improves the aggregation ability of context information and the multi-scale feature extraction ability, enhances the discrimination of building features, and improves the segmentation effect.
Smart Images

Figure CN119445380B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly relates to an intelligent building extraction method based on remote sensing images. Background Technique
[0002] With the rapid development of satellite imaging technology, high-spatial-resolution remote sensing data at home and abroad has become increasingly rich, playing an important role in many fields such as national economic construction, resource and environmental protection. Automatically and accurately extracting information such as the location, boundary, and type of ground objects from high-resolution remote sensing images is a key and difficult point in remote sensing image processing.
[0003] In recent years, with the rise of deep neural networks, remote sensing image semantic segmentation based on deep neural network models has achieved high accuracy on standard datasets and also shown great application potential in aspects such as ground object recognition and information extraction from high-resolution remote sensing images.
[0004] The sizes of buildings in remote sensing images vary greatly, and overall present the characteristics of small and weak targets, bringing very great challenges to the segmentation of buildings in remote sensing images. In addition, there are large inter-class differences and small intra-class differences in remote sensing images. The roof types and heights of buildings in remote sensing images show a long-tail distribution, and the edges of different categories in remote sensing images are similar, such as roofs and roads. In actual scenarios, affected by the shooting time and angle, there are interference factors such as shadows, occlusions, and clouds in remote sensing images. These factors will all make the segmentation of buildings face great difficulties.
[0005] At present, the application of deep semantic segmentation models in the task of remote sensing image semantic segmentation has the following difficulties: on the one hand, the contradiction between the inherent high-level semantics and low-level details of deep convolutional networks causes a large amount of loss of detail information; at the same time, high-resolution remote sensing images are rich in details, and the neural network needs to be able to extract more discriminative detail information and multi-scale information. Summary of the Invention
[0006] In order to solve the above technical problems existing in the prior art, the purpose of the present invention is to provide an intelligent building extraction method, electronic device, and storage medium based on remote sensing images, which can effectively improve the performance of extracting buildings in remote sensing images.
[0007] To achieve the above invention purpose, the present invention provides an intelligent building extraction method based on remote sensing images, including the following steps:
[0008] Step S1, obtain a high-resolution dataset;
[0009] Step S2: Design a remote sensing image building semantic segmentation network based on an encoding-decoding structure. The remote sensing image building semantic segmentation network includes a semantic segmentation encoder and a semantic segmentation decoder. The semantic segmentation encoder is a feature extraction backbone network, and the feature extraction backbone network includes a GLFE module for fusing global features and local features;
[0010] Step S3: Design a remote sensing image building DSM estimation network based on a generative adversarial network. The DSM estimation network includes a DSM generator and a DSM discriminator. The DSM generator includes a DSM generator encoder and a DSM generator decoder, and the DSM generator encoder is the feature extraction backbone network;
[0011] Step S4: Design a feature fusion and enhancement module. The feature fusion and enhancement module is used to realize the output feature fusion of the feature extraction backbone network, and the feature fusion between the semantic segmentation decoder and the DSM generator decoder, including a MASPP module and a DSATT module;
[0012] Step S5: Design a loss function;
[0013] Step S6: According to the high-resolution dataset and the loss function, train and optimize the remote sensing image building intelligent extraction network. The remote sensing image building intelligent extraction network includes the remote sensing image building semantic segmentation network, the remote sensing image DSM estimation network, and the feature fusion and enhancement module;
[0014] Step S7: Perform intelligent extraction of buildings based on remote sensing images through the trained remote sensing image building intelligent extraction network.
[0015] According to a technical solution of the present invention, the step S2 specifically includes the following steps:
[0016] Step S21: Design the feature extraction backbone network;
[0017] Among them, the feature extraction backbone network is based on the ResNet101 network as the basic network, divided into 5 stages, removing the fully connected layer on the network output side, and no downsampling is performed in stage5. It is pre-trained using ImageNet. The input size of the feature extraction backbone network is 3×512×512. Among them, the structures of stage1 and stage2 of the feature extraction backbone network are the same as those of the basic structure of ResNet101, and the GLFE module is included in stage3, stage4, and stage5 of the feature extraction backbone network; Step S22: Design the GLFE module;
[0018] The GLFE module includes a global feature extraction module GFE and a local feature extraction module LFE. The global feature extraction module GFE is implemented by an external attention module with two linear layers, which is expressed as follows:
[0019] F out = Linear2(Norm(Linear1(F in )))
[0020] The local feature extraction module LFE consists of two parallel branches. After the input of the local feature extraction module LFE undergoes channel upsampling through a 1×1 convolution, the channels are evenly divided. Half of the evenly divided channels pass through a branch with a cascaded 1×3 and 3×1 convolution to extract features, which is expressed as follows:
[0021]
[0022] Another branch of the local feature extraction module LFE takes the sum of F 1out and the features of the other half of the evenly divided channels as input, and extracts features through pooling and 1×5 and 5×1 convolutions, which is expressed as follows:
[0023]
[0024] Take F 1out and F 2out Perform a 1×1 convolution on the channel dimension, which is expressed as follows:
[0025] F = conv 1×1 (F 1out + F 2out )
[0026] Step S23: Design the semantic segmentation decoder;
[0027] The semantic segmentation decoder is used to upsample the semantic segmentation features extracted by the feature extraction backbone network to the input size; the semantic segmentation decoder includes two branches. The first branch of the semantic segmentation decoder is a cascaded 3×1 and 1×3 convolution kernel, which is expressed as follows:
[0028] F branch1 = conv 1×3 (conv 3×1 (F dec ))
[0029] where F dec represents the input of the semantic segmentation decoder, and F branch1 represents the output of the first branch of the semantic segmentation decoder;
[0030] The second branch of the semantic segmentation decoder takes the fusion of the output features of the first branch of the semantic segmentation decoder and the input features of the semantic segmentation decoder as input, and outputs after passing through a depthwise separable convolutional layer with a convolution kernel of 3 and a dilation rate of 2 and a densely connected layer with a convolution kernel of 3 and a dilation rate of 3. The output of the second branch of the semantic segmentation decoder is denoted as F branch2 ; then the two branches of the semantic segmentation decoder are concatenated and fused in the channel dimension, and channel information fusion is performed through a convolution with a convolution kernel of 1×1.
[0031] Step S24: Design the fusion module of the semantic segmentation decoder. For each stage of the semantic segmentation decoder, its input includes the output of the previous stage of the semantic segmentation decoder, the output of the corresponding stage of the DSM generator decoder, and the output of the corresponding stage of the feature extraction backbone network.
[0032] According to a technical solution of the present invention, the design of the building DSM estimation network in step S3 specifically includes the following steps:
[0033] Step S31: Design the DSM generator decoder;
[0034] The DSM generator decoder adopts a dual-branch fusion module, and its input is evenly divided into channels after the input channel dimension is increased through a 1×1 conventional convolution, that is, F c / 2,1 、F c / 2,2 ;
[0035] The first branch of the DSM generator decoder is a 3×3 conventional convolution, and the input of the second branch of the DSM generator decoder is the channel fusion feature of F c / 2,2 and the first branch of the DSM generator decoder, and outputs after passing through a 3×3 convolution; finally, the outputs of the two branches of the DSM generator decoder are added and fused;
[0036] The channel dimension of each stage of the DSM generator decoder is half of the corresponding stage of the building semantic segmentation decoder;
[0037] Step S32: Design the fusion module of the DSM generator decoder. For each stage of the DSM generator decoder, its input includes the output of the previous stage of the DSM generator decoder, the output of the corresponding stage of the semantic segmentation decoder, and the output of the corresponding stage of the feature extraction backbone network.
[0038] Step S33: Design the DSM discriminator;
[0039] The DSM discriminator is used to identify the authenticity of the DSM generated by the DSM generator. Using ResNet50 as the basic framework, it includes 5 stages. Among them, the second stage, the fourth stage, and the fifth stage of the DSM discriminator all use a dual-branch feature extraction module. The dual-branch feature extraction module realizes channel dimension increase through a 1×1 conventional convolution for the input. The input of the first branch of the dual-branch feature extraction module is a 1 / 2-channel feature, and feature extraction is realized through a 3×3 depthwise separable convolution. The second branch of the dual-branch feature extraction module adds another 1 / 2-channel feature to the output of the first branch of the dual-branch feature extraction module as the input, and outputs after passing through a depthwise separable dilated convolution with a convolution kernel of 3 and a dilation rate of 2. Finally, the two branches of the dual-branch feature extraction module are fused in channels, and feature extraction is realized through a 1×1 convolution.
[0040] According to a technical solution of the present invention, the design of the feature fusion and enhancement module in step S4 includes the following steps:
[0041] Step S41: Design the MASPP module after the output of the feature extraction backbone network;
[0042] The MASPP module adopts multi-branch parallelism. The input of the MASPP module is F pin , the first branch of the MASPP module is a conventional convolution with a convolution kernel of 1×1, and the output of the first branch of the MASPP module is F p1 ; the second branch of the MASPP module is a conventional convolution with a convolution kernel of 3×3, and the output of the second branch of the MASPP module is F p2 ; the input of the third branch of the MASPP module is the fusion of F pin and F p2 , and the output of the third branch of the MASPP module can be expressed as follows:
[0043] F p3_1 = Aconv 3,2 (AP 3 (F pin + F p2 ))
[0044]
[0045] F p3 = F p3_1 + F p3_2 + AP 3 (F pin + F p2 )
[0046] Wherein, AP3 Denotes average pooling with a kernel size of 3, Aconv 3,2 Denotes dilated convolution with a convolution kernel size of 3 and a dilation rate of 2, Aconv 3,3 Denotes dilated convolution with a convolution kernel size of 3 and a dilation rate of 3 Denotes that features are fused based on channels;
[0047] The input of the fourth branch of the MASPP module is F pin And F p3 The fusion output can be expressed as follows:
[0048] F p4_1 = Aconv 3,4 (AP 5 (F pin + F p3 ))
[0049]
[0050] F p4 = F p4_1 + F p4_2 + AP 5 (F pin + F p3 )
[0051] Where, AP 5 Denotes average pooling with a kernel size of 5, Aconv 3,4 Denotes dilated convolution with a convolution kernel size of 3 and a dilation rate of 4, Aconv 3,5 Denotes dilated convolution with a convolution kernel size of 3 and a dilation rate of 5; the fifth branch of the MASPP module is global average pooling, and the output of the fifth branch of the MASPP module is denoted as F p5 ;
[0052] The output ends of the parallel branches of the MASPP module use the channel fusion method to fuse F p1 , F p2 , F p3 , F p4 And F p5 To achieve information fusion between channels through 1×1 convolution, so that the channel dimension is consistent with the module input, and the output is denoted as F piout ; Strengthen effective features through the channel attention mechanism, and the output of the MASPP module is:
[0053] F pout = F piout + SEB(F piout )
[0054] Among them, SEB(.) represents the fusion channel attention mechanism of average pooling and maximum pooling;
[0055] Step S42: Design the DSATT module;
[0056] The input of the DSATT module is the output feature F of the semantic segmentation decoder seg and the output feature F of the DSM generator decoder at the corresponding stage dsm , and the output process of the DSATT module can be expressed as follows:
[0057] Att seg = S(Linear(ReLU(Linear(GAP(F seg )))))
[0058] Att dsm = S(Linear(ReLU(Linear(GAP(F DSM )))))
[0059]
[0060]
[0061] where F segout 、F dsmout represent the output after passing through the DSATT module, GAP(.) represents passing through global average pooling, Linear(.) represents linear transformation, S(.) represents the activation function, Att seg and Att dsm represent the channel attention weights of the output features of the semantic segmentation decoder and the channel attention weights of the output features of the DSM generator decoder respectively, and the sum of Art seg and Att dsm is 1.
[0062] According to a technical solution of the present invention, in the step S5, it specifically includes:
[0063] Step S51: Design the loss function of the remote sensing image building semantic segmentation network. The total loss of the remote sensing image building semantic segmentation network includes pixel-level loss and regional loss. The calculation formula of the pixel-level loss is as follows:
[0064]
[0065] where yi c represents the true label at position i, represented by one-hot. When the true label is a building, it is 1, otherwise it is 0. C is the number of categories, which is 2 here, and N is the total number of pixels in the prediction map; pic Represents the probability of classifying sample i as class c; w c Is a weight parameter, calculates the number of pixels of buildings in all samples of the training set, and calculates w based on this c , which is expressed as follows:
[0066]
[0067] The calculation formula of the regional loss is as follows:
[0068]
[0069] Among them, y pred Represents the probability of predicting class c, y true Represents the true one-hot label;
[0070] The total loss of the remote sensing image building semantic segmentation network is:
[0071] L seg =αL wce +βL softIoU ;
[0072] Step S52, design the loss function of the building DSM estimation network. The total loss of the building DSM estimation network includes the adversarial loss and the height estimation loss. The adversarial loss is expressed as follows:
[0073] L G =-E(logD(G(x input )))
[0074] L D =-E(logD(x input_dsm ))-E(log(1-D(G(x input ))))
[0075] L adv =L G +L D
[0076] Among them, G and D respectively represent the DSM generator encoder and the DSM generator decoder, E(.) represents the mathematical expectation, x input_dsm Represents the true DSM of the input remote sensing image, x input Represents the input remote sensing image;
[0077] The height estimation loss is expressed as follows:
[0078]
[0079] The total loss of the building DSM estimation network is:
[0080] L dsm = γ 1 L adv + γ 2 L h ;
[0081] where γ 1 and γ 2 respectively represent the loss weights of the adversarial loss and the height estimation loss.
[0082] According to a technical solution of the present invention, in the step S1, the high-resolution dataset includes the Vaihingen dataset and the Potsdam dataset, and each sample in the high-resolution dataset includes an RGB image and its corresponding DSM image; the samples of the Vaihingen dataset and the Potsdam dataset are preprocessed and cropped into slices of 521 pixels × 512 pixels.
[0083] According to an aspect of the present invention, there is provided an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes a method for intelligent extraction of buildings based on remote sensing images as described in any one of the above technical solutions.
[0084] According to an aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, a method for intelligent extraction of buildings based on remote sensing images as described in any one of the above technical solutions is implemented.
[0085] Compared with the prior art, the present invention has the following beneficial effects:
[0086] 1. Through the optimized encoder, the network's ability to aggregate context information and multi-scale feature extraction ability are effectively improved;
[0087] 2. Aiming at the multi-scale problem of ground objects, the present invention adds a multi-scale dense connection atrous convolution attention module between the encoder and the decoder. Through the multi-branch structure, conventional convolution, pooling operation, atrous pyramid convolution, dense connection, and feature flow between branches, the extraction and fusion of features at different scales are realized, enriching the context information.
[0088] 3. Channel attention is added after the multi-scale dense connection atrous pyramid pooling to strengthen the effective features, making the extracted building features more discriminative;
[0089] 4. To address the problem of difficulty in distinguishing similar features, the present invention optimizes multi-task features through multi-task joint. By jointly performing building semantic segmentation task and DSM generation task, information complementarity between semantic features and height features is achieved to enhance the features, and multi-source features are utilized to improve the segmentation performance;
[0090] 5. Regarding the time and cost of DSM acquisition, the present invention uses generative adversarial network technology to generate DSM for a single remote sensing image. When predicting and segmenting buildings, only a single remote sensing image needs to be input to simultaneously obtain the corresponding DSM and segmentation map. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0092] Figure 1 Schematically showing the design flow chart of the building intelligent extraction model based on remote sensing images according to an embodiment of the present invention;
[0093] Figure 2 Schematically showing the flow chart of the building intelligent extraction method based on remote sensing images according to an embodiment of the present invention;
[0094] Figure 3 Schematically showing the overall software architecture diagram of the building intelligent extraction method based on remote sensing images according to an embodiment of the present invention;
[0095] Figure 4 Schematically showing the GLFE module structure diagram according to an embodiment of the present invention;
[0096] Figure 5 Schematically showing the MASPP module structure diagram according to an embodiment of the present invention;
[0097] Figure 6 Schematically showing the DSATT module structure diagram according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0098] The description of the embodiments of this specification should be combined with the corresponding drawings, and the drawings should be regarded as a part of the complete specification. In the drawings, the shape or thickness of the embodiments can be enlarged and simplified or conveniently marked. Furthermore, each part of the structure in the drawings will be described separately. It should be noted that the elements not shown or described in words in the drawings are in the forms known to those of ordinary skill in the art.
[0099] In the description of the embodiments herein, any reference to directions and orientations is for convenience of description only and should not be construed as any limitation on the scope of protection of the present invention. The following description of the preferred embodiments involves combinations of features, which may exist independently or in combination. The present invention is not particularly limited to the preferred embodiments. The scope of the present invention is defined by the claims.
[0100] Based on the method of the present invention, this embodiment provides an intelligent building extraction method based on remote sensing images, as Figures 1 to 6 shown, including the following steps:
[0101] Step S1, obtain a high-resolution dataset;
[0102] According to a technical solution of the present invention, in step S1, the high-resolution dataset includes the Vaihingen dataset and the Potsdam dataset. Each sample in the high-resolution dataset includes an RGB image and its corresponding DSM image; the samples of the Vaihingen dataset and the Potsdam dataset are preprocessed and cropped into slices of 521 pixels × 512 pixels.
[0103] Step S2, design a remote sensing image building semantic segmentation network based on an encoder-decoder structure. The remote sensing image building semantic segmentation network includes a semantic segmentation encoder and a semantic segmentation decoder. The semantic segmentation encoder is a feature extraction backbone network, and the feature extraction backbone network includes a GLFE module for global feature and local feature fusion;
[0104] According to a technical solution of the present invention, step S2 specifically includes the following steps:
[0105] Step S21, design a feature extraction backbone network;
[0106] Among them, the feature extraction backbone network is based on the ResNet101 network as the basic network, divided into 5 stages, removing the fully connected layer on the output side of the network, and no downsampling is performed in stage 5. It is pre-trained using ImageNet; the input size of the feature extraction backbone network is 3×512×512; among them, the basic structures of stage 1 and stage 2 of the feature extraction backbone network are the same as those of ResNet101, and the GLFE module is included in stage 3, stage 4, and stage 5 of the feature extraction backbone network.
[0107] Step S22, design a GLFE module;
[0108] The GLFE module includes a global feature extraction module GFE and a local feature extraction module LFE. The global feature extraction module GFE is implemented by using an external attention module with two linear layers, expressed as follows:
[0109] F ou t = Linear2(Norm(Linear1(F in )))
[0110] The local feature extraction module LFE is composed of two parallel branches. After the input of the local feature extraction module LFE passes through a 1×1 convolution for channel dimension expansion, the channels are evenly divided. Half of the evenly divided channels pass through a branch with a cascaded 1×3 and 3×1 convolution to extract features, which is expressed as follows:
[0111]
[0112] Another branch of the local feature extraction module LFE takes F 1out and the sum of the features of the other half of the evenly divided channels as the input, and extracts features through pooling and 1×5 and 5×1 convolutions, which is expressed as follows:
[0113]
[0114] Take F 1out and F 2out Perform a 1×1 convolution in the channel dimension, which is expressed as follows:
[0115] F = conv 1×1 (F 1out + F 2out );
[0116] The above process adopts a global-local feature extraction branch to extract features with different receptive fields, strengthen information flow, and enrich context information, making the features of the building more robust to different scales.
[0117] Step S23: Design a semantic segmentation decoder;
[0118] The semantic segmentation decoder is used to upsample the semantic segmentation features extracted by the feature extraction backbone network to the input size; the semantic segmentation decoder needs to go through 4 times of upsampling to restore to 3×512×512. The semantic segmentation decoder includes two branches. The first branch of the semantic segmentation decoder is a cascaded 3×1 and 1×3 convolution kernel, which is expressed as follows:
[0119] F branch1 = conv 1×3 (conv 3×1 (F dec ))
[0120] where F dec represents the input of the semantic segmentation decoder, F branch1Represents the output of the first branch of the semantic segmentation decoder;
[0121] The second branch of the semantic segmentation decoder takes as input the fusion of the output features of the first branch of the semantic segmentation decoder and the input features of the semantic segmentation decoder, and outputs after passing through a depthwise separable convolutional layer with a convolutional kernel of 3 and a dilation rate of 2 and a densely connected layer with a convolutional kernel of 3 and a dilation rate of 3. The output of the second branch of the semantic segmentation decoder is denoted as F branch2 ; Then, the two branches of the semantic segmentation decoder are concatenated and fused in the channel dimension, and channel information fusion is performed through a convolution with a convolutional kernel of 1×1.
[0122] This module performs dense connection through depthwise separable convolutions with different branches and different dilation rates, expands the receptive field with fewer parameters, realizes the fusion of features at different scales, and solves the multi-scale characteristics of buildings.
[0123] Step S24, design the fusion module of the semantic segmentation decoder. For each stage of the semantic segmentation decoder, its input includes the output of the previous stage of the semantic segmentation decoder, the corresponding stage output of the DSM generator decoder, and the corresponding stage output of the feature extraction backbone network. The fusion module of the semantic segmentation decoder effectively fuses multi-source features, detailed features of the encoder part, and high-level features of the previous stage of the decoder, making the extracted building features more discriminative.
[0124] Through step S2, the encoding part and the decoding part of the building semantic segmentation are designed. The global-local branch is used to enrich the context features, and dense connection is performed through depthwise separable convolutions with different branches and different dilation rates to realize the fusion of features at different scales, alleviating the impact brought by the multi-scale characteristics of remote sensing images.
[0125] Step S3, design a remote sensing image building DSM estimation network based on a generative adversarial network. The DSM estimation network includes a DSM generator and a DSM discriminator, which are used for DSM generation of buildings in the inference stage. The DSM generator includes a DSM generator encoder and a DSM generator decoder, and the DSM generator encoder is a feature extraction backbone network;
[0126] Step S31, design the DSM generator decoder; In order to alleviate the problems of complex and costly DSM generation process, generative adversarial network technology is used to synthesize DSM. The encoder part of the DSM generator in the present invention uses the feature extraction backbone network of step S21.
[0127] The DSM generator decoder adopts a dual-branch fusion module, and its input is evenly divided into channels after the input channel dimension is increased through a 1×1 conventional convolution, that is, F c / 2,1 、F c / 2,2 ;
[0128] The first branch of the DSM generator decoder is a 3×3 conventional convolution, and the input of the second branch of the DSM generator decoder is F c / 2,2 It is fused with the channel fusion features of the first branch of the DSM generator decoder and output after 3×3 convolution; finally, the outputs of the two branches of the DSM generator decoder are added and fused;
[0129] The channel dimension of each stage of the DSM generator decoder is half of that of the corresponding stage of the building semantic segmentation decoder;
[0130] Step S32: Design the fusion module of the DSM generator decoder. For each stage of the DSM generator decoder, its input includes the output of the previous stage of the DSM generator decoder, the output of the corresponding stage of the semantic segmentation decoder, and the output of the corresponding stage of the feature extraction backbone network. The fusion module of the DSM generator decoder effectively fuses multi-source features, detailed features of the encoder part, and high-level features of the previous stage of the decoder, enriching the DSM features.
[0131] Step S33: Design the DSM discriminator;
[0132] The DSM discriminator is used to identify the authenticity of the DSM generated by the DSM generator. It uses ResNet50 as the basic backbone and includes 5 stages. Among them, the 2nd stage, the 4th stage, and the 5th stage of the DSM discriminator all use a dual-branch feature extraction module. The dual-branch feature extraction module realizes channel upsampling through 1×1 conventional convolution for the input. The input of the first branch of the dual-branch feature extraction module is a 1 / 2-channel feature, and feature extraction is realized through 3×3 depthwise separable convolution. The second branch of the dual-branch feature extraction module adds another 1 / 2-channel feature to the output of the first branch of the dual-branch feature extraction module as the input, and outputs after depthwise separable dilated convolution with a convolution kernel of 3 and a dilation rate of 2. Finally, the two branches of the dual-branch feature extraction module are channel-fused and feature extraction is realized through 1×1 convolution. By adopting different basic modules at different stages, the extraction of features at different scales is realized, making the features of the building DSM more discriminative.
[0133] The DSM estimation network is used to generate DSM for a single RGB remote sensing image to assist in the feature extraction of building semantic segmentation. In order to effectively fuse the features of the common encoder part and the decoding branches of different task branches, increase the information flow, and effectively fuse the semantic segmentation branch and the DSM estimation branch, different fusion modules will be used below.
[0134] Step S4: Design a feature fusion and enhancement module. The feature fusion and enhancement module is used to achieve the feature fusion of the output features of the feature extraction backbone network, the semantic segmentation decoder, and the DSM generator decoder, including the MASPP module and the DSATT module. The MASPP module realizes multi-scale feature fusion, fuses multi-scale features, and enhances effective features. The DSATT module enhances the fusion of different modality features and improves the performance of semantic segmentation.
[0135] According to a technical solution of the present invention, the design of the feature fusion and enhancement module in step S4 includes the following steps:
[0136] Step S41: Design the MASPP module after the output of the feature extraction backbone network.
[0137] As Figure 4 shown. In order to aggregate context information and enrich multi-scale features to improve the segmentation performance, the MASPP module adopts multiple branches in parallel. The input of the MASPP module is F pin . The first branch of the MASPP module is a conventional convolution with a convolution kernel of 1×1, and the output of the first branch of the MASPP module is F p1 . The second branch of the MASPP module is a conventional convolution with a convolution kernel of 3×3, and the output of the second branch of the MASPP module is F p2 . The input of the third branch of the MASPP module is the fusion of F pin and F p2 . The output of the third branch of the MASPP module can be expressed as follows:
[0138] F p3_1 =Aconv 3,2 (AP 3 (F pin +F p2 ))
[0139]
[0140] F p3 =F p3_1 +F p3_2 +AP 3 (F pin +F p2 )
[0141] where, AP 3 represents average pooling with a kernel of 3, and Aconv 3,2 represents dilated convolution with a convolution kernel of 3 and a dilation rate of 2, and Aconv 3,3 represents dilated convolution with a convolution kernel of 3 and a dilation rate of 3. It means that features are fused according to channels; for the third branch and above, the previous branch and the input are used for fusion, and average pooling and dilated dense convolution are combined to achieve the fusion of features at different scales.
[0142] The input of the fourth branch of the MASPP module is F pin For the fusion with F p3 The output can be expressed as follows:
[0143] F p4_1 = Aconv 3,4 (AP 5 (F pin + F p3 ))
[0144]
[0145] F p4 = F p4_1 + F p4_2 + AP 5 (F pin + F p3 )
[0146] Among them, AP 5 represents average pooling with a kernel of 5, Aconv 3,4 represents dilated convolution with a convolution kernel of 3 and a dilation rate of 4, and Aconv 3,5 represents dilated convolution with a convolution kernel of 3 and a dilation rate of 5; the fifth branch of the MASPP module is global average pooling, and the output of the fifth branch of the MASPP module is expressed as F p5 ;
[0147] At the output ends of the parallel branches of the MASPP module, the channels of F p1 , F p2 , F p3 , F p4 and F p5 are fused using the channel fusion method. Information fusion between channels is achieved through 1×1 convolution, so that the channel dimension is consistent with the module input, and the output is expressed as F piout ; effective features are strengthened through the channel attention mechanism, and the output of the MASPP module is:
[0148] F pout = F piout + SEB(F piout )
[0149] Among them, SEB(.) represents the fused channel attention mechanism of average pooling and max pooling;
[0150] The design of the MASPP module integrates parallel branches with different receptive fields, introduces pooling operations and densely connected dilated convolutional pyramids, and realizes multi-scale feature extraction and flow between features. Finally, the effective features are strengthened through the channel attention mechanism.
[0151] Step S42, designing a DSATT module;
[0152] like Figure 5 As shown in Figure 1, the DSATT module effectively integrates features from multiple sources and uses the complementary attention features of the two branches to strengthen the main task and auxiliary tasks, thereby improving the feature recovery function of the decoder. The input of the DSATT module is the output feature F of the semantic segmentation decoder. seg And the output feature F of the DSM generator decoder at the corresponding stage dsm , the output process of the DSATT module can be expressed as follows:
[0153] Att seg =S(Linear(ReLU(Linear(GAP(F seg )))))
[0154] Att dsm =S(Linear(ReLU(Linear(GAP(F DSM )))))
[0155]
[0156]
[0157] Among them, F segout 、F dsmout represents the output after the DSATT module, GAP(.) represents global average pooling, Linear(.) represents linear change, S(.) represents the activation function, Att seg With Att dsm Att denotes the channel attention weights of the output features of the semantic segmentation decoder and the output features of the DSM generator decoder, respectively. seg With Att dsm The sum is 1. The feature enhancement of the main branch is achieved, and the constraints of the auxiliary tasks are added to improve each other.
[0158] The above steps complete the data preprocessing and network design, and the loss function will be designed to train the network later.
[0159] Step S5: design a loss function; constrain the entire network architecture, train and optimize the network, and improve network performance.
[0160] According to a technical solution of the present invention, in step S5, it specifically includes:
[0161] Step S51: Design the loss function of the remote sensing image building semantic segmentation network. Considering that the proportion of buildings in the remote sensing image is small and to solve the problem of class imbalance, the weighted cross-entropy is used to optimize the semantic segmentation network. The total loss of the remote sensing image building semantic segmentation network includes the pixel-level loss and the region loss. The calculation formula of the pixel-level loss is as follows:
[0162]
[0163] where, yi c represents the true label at position i, represented by one-hot. When the true label is a building, it is 1, otherwise it is 0. C is the number of classes, which is 2 here, and N is the total number of pixels in the prediction map; p ic represents the probability of classifying sample i as class c; w c is the weight parameter. Calculate the number of pixels of buildings in all samples of the training set, and calculate w c as follows:
[0164]
[0165] To increase the constraint of buildings, a region loss constraint is added on the basis of the pixel-level loss. The calculation formula of the region loss is as follows:
[0166]
[0167] where, y pred represents the probability of predicting class c, and y true represents the true one-hot label;
[0168] The total loss of the remote sensing image building semantic segmentation network is:
[0169] L seg =αL wce +βL softIoU ;
[0170] Step S52: Design the loss function of the building DSM estimation network. The total loss of the building DSM estimation network includes the adversarial loss and the height estimation loss. The adversarial loss is expressed as follows:
[0171] L G =-E(logD(G(x input )))
[0172] L D =-E(logD(x input_dsm ))-E(log(1-D(G(x input))))
[0173] L adv = L G + L D
[0174] Among them, G and D respectively represent the DSM generator encoder and the DSM generator decoder, E(.) represents the mathematical expectation, and x input_dsm represents the true DSM of the input remote sensing image, and x input represents the input remote sensing image;
[0175] The DSM generator and the discriminator synthesize the true DSM image through game confrontation. In order to constrain the accuracy of DSM synthesis, the Huber loss with δ being 1 is used to reduce the sensitivity to DSM outliers. The height estimation loss is expressed as follows:
[0176]
[0177] The total loss of the building DSM estimation network is:
[0178] L dsm = γ 1 L adv + γ 2 L h ;
[0179] Among them, γ 1 and γ 2 respectively represent the loss weights of the adversarial loss and the height estimation loss.
[0180] Step S6: According to the high-resolution dataset and the loss function, train and optimize the intelligent extraction network for buildings in remote sensing images. The intelligent extraction network for buildings in remote sensing images includes a semantic segmentation network for buildings in remote sensing images, a DSM estimation network for remote sensing images, and a feature fusion and enhancement module;
[0181] Step S7: Perform intelligent extraction of buildings based on remote sensing images through the trained intelligent extraction network for buildings in remote sensing images.
[0182] Through the above steps, the intelligent extraction of buildings from high-resolution remote sensing images can be achieved. In view of the factors such as rich details of ground objects in high-resolution remote sensing images, large scale variations of buildings, and difficulty in distinguishing similar categories, which result in poor building segmentation performance in the actual scenario, the present invention designs multi-task joint to achieve the intelligent extraction of buildings from high-resolution remote sensing images. In view of the limited context information extraction of the deep convolutional neural network, the present invention optimizes the encoder by designing a global-local dual-branch module, realizes the extraction of global information through a simplified self-attention mechanism, and realizes the fusion of features with different receptive fields through convolutional cascading and the fusion of features with different convolutional kernels. Through the optimized encoder, the network's ability to aggregate context information and extract multi-scale features is effectively improved. In view of the multi-scale problem of ground objects, the present invention adds a multi-scale densely connected atrous convolution attention module between the encoder and the decoder, and realizes the extraction and fusion of features with different scales through a multi-branch structure, conventional convolution, pooling operation, atrous pyramid convolution, dense connection, and feature flow between branches, enriching the context information. In addition, channel attention is added after the multi-scale densely connected atrous pyramid pooling to strengthen the effective features, making the extracted building features more discriminative.
[0183] In view of the problem that it is difficult to distinguish similar ground objects, the present invention optimizes multi-task features through multi-task joint. By jointly performing the building semantic segmentation task and the DSM generation task, the information complementarity of semantic features and height features is realized, the features are enhanced, and multi-source features are used to improve the segmentation performance. In view of the time and cost of DSM acquisition, the present invention uses the generative adversarial network technology to generate DSM for a single remote sensing image. In view of the fusion problem of semantic features and height features, the present invention designs a complementary attention mechanism. In the semantic feature branch, the branch attention enhanced feature and the complementary attention feature of the height feature are fused, and in the height feature branch, the branch attention enhanced feature and the complementary attention feature of the semantic feature are fused, so as to increase the features of the main task and supplement with the complementary features of the other branch, effectively strengthening the information flow and fusion. In order to extract multi-scale features and detail information, a dual-branch feature extraction module is introduced into the decoder of the DSM estimation network and the DSM discriminator, and convolutional features with different atrous rates are fused. During the optimization training process, the network is optimized through pre-training methods and various loss functions, including adversarial loss, height estimation loss, semantic segmentation cross-entropy loss, and intersection over union loss. When predicting and segmenting buildings, only a single remote sensing image needs to be input, and the corresponding DSM and segmentation map can be obtained simultaneously.
[0184] According to one aspect of the present invention, there is provided an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes a method for intelligent extraction of buildings based on remote sensing images according to any one of the above technical solutions.
[0185] The processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0186] According to one aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, a method for intelligent extraction of buildings based on remote sensing images according to any one of the above technical solutions is implemented.
[0187] The computer-readable storage medium may include any medium capable of storing or transmitting information. Examples of the computer-readable storage medium include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. Code segments may be downloaded via a computer network such as the Internet, an intranet, etc.
[0188] In addition, it should be noted that the present invention may be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0189] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0190] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0191] It should also be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or terminal device including the said element.
[0192] Finally, it should be noted that the above description is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once they know the basic creative concept of the present invention, several improvements and refinements can be made without departing from the principle described in the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A method for intelligent building extraction based on remote sensing images, characterized in that: The following steps are involved: Step S1, obtaining a high-resolution data set; Step S2, designing a remote sensing image building semantic segmentation network based on an encoding-decoding structure, wherein the remote sensing image building semantic segmentation network includes a semantic segmentation encoder and a semantic segmentation decoder, wherein the semantic segmentation encoder is a feature extraction skeleton network, and the feature extraction skeleton network includes a GLFE module for fusing global features with local features; Step S3, designing a remote sensing image building DSM estimation network based on a generative adversarial network, wherein the DSM estimation network includes a DSM generator and a DSM discriminator, wherein the DSM generator includes a DSM generator encoder and a DSM generator decoder, and the DSM generator encoder is the feature extraction skeleton network; Step S4, designing a feature fusion and enhancement module, the feature fusion and enhancement module is used to realize the output feature fusion of the feature extraction skeleton network, the feature fusion of the semantic segmentation decoder and the DSM generator decoder, and includes a MASPP module and a DSATT module; the MASPP module realizes multi-scale feature fusion, fuses multi-scale features and enhances effective features; The DSATT module strengthens the fusion of different modality features and improves the performance of semantic segmentation; Step S5, designing a loss function; Step S6, training and optimizing a remote sensing image building intelligent extraction network according to the high-resolution data set and the loss function, wherein the remote sensing image building intelligent extraction network includes the remote sensing image building semantic segmentation network, the remote sensing image DSM estimation network and the feature fusion and enhancement module; Step S7: performing intelligent extraction of buildings based on remote sensing images through the trained remote sensing image building intelligent extraction network.
2. The intelligent building extraction method based on remote sensing images according to claim 1 is characterized in that: The step S2 specifically comprises the following steps: Step S21, designing the feature extraction skeleton network; The feature extraction skeleton network uses the ResNet101 network as the basic network, is divided into 5 stages, removes the fully connected layer on the network output side, does not perform downsampling in stage5, and uses ImageNet for pre-training; the input size of the feature extraction skeleton network is 3×512×512; the stage1 and stage2 of the feature extraction skeleton network have the same basic structure as ResNet101, and the stage3, stage4 and stage5 of the feature extraction skeleton network include the GLFE module; Step S22, designing the GLFE module; The GLFE module includes a global feature extraction module GFE and a local feature extraction module LFE. The global feature extraction module GFE is implemented by an external attention module using a two-linear layer, which is expressed as follows: F out =Linear2(Norm(Linear1(F in ))) The local feature extraction module LFE is composed of two parallel branches. The input of the local feature extraction module LFE is subjected to 1×1 convolution for channel dimension increase, and the channel is evenly divided. The half of the channels after even division are extracted through a 1×3, 3×1 convolution cascade branch, which is expressed as follows: Another branch of the local feature extraction module LFE is F 1out The sum of the features of the other half of the channels after equal division is used as input, and features are extracted through pooling and 1×5 and 5×1 convolutions, as shown below: F 1out With F 2out A 1×1 convolution is performed on the channel dimension, which is expressed as follows: F=conv 1×1 (F 1out +F 2out ); Step S23, designing the semantic segmentation decoder; The semantic segmentation decoder is used to upsample the semantic segmentation features extracted by the feature extraction skeleton network to restore them to the input size; the semantic segmentation decoder includes two branches, and the first branch of the semantic segmentation decoder is a convolution kernel of 3×1 and 1×3 cascade, which is expressed as follows: F branch1 =conv 1×3 (conv 3×1 (F dec )) Among them, F dec represents the input of the semantic segmentation decoder, F branch1 represents the output of the first branch of the semantic segmentation decoder; The second branch of the semantic segmentation decoder takes the output features of the first branch of the semantic segmentation decoder and the fusion of the input features of the semantic segmentation decoder as input, and outputs after passing through a depthwise separable convolutional layer with a convolution kernel of 3 and a dilation rate of 2 and a densely connected layer with a convolution kernel of 3 and a dilation rate of 3. The output of the second branch of the semantic segmentation decoder is represented by F branch2 ; Then the two branches of the semantic segmentation decoder are spliced and fused from the channel dimension, and the channel information is fused through a convolution with a convolution kernel of 1×1; Step S24, design a fusion module of the semantic segmentation decoder; for each stage of the semantic segmentation decoder, its input includes the output of the previous stage of the semantic segmentation decoder, the corresponding stage output of the DSM generator decoder, and the output of the corresponding stage of the feature extraction skeleton network.
3. The intelligent building extraction method based on remote sensing images according to claim 2 is characterized in that: The design of the building DSM estimation network in step S3 specifically includes the following steps: Step S31, designing the DSM generator decoder; The DSM generator decoder adopts a dual-branch fusion module, whose input is divided into channels after the input channel dimension is increased by 1×1 regular convolution, that is, F c / 2,1 、F c / 2,2 ; The first branch of the DSM generator decoder is a 3×3 regular convolution, and the input of the second branch of the DSM generator decoder is F c / 2,2 The channel fusion features of the first branch of the DSM generator decoder are output through 3×3 convolution; finally, the outputs of the two branches of the DSM generator decoder are added and fused; The channel dimension of each stage of the DSM generator decoder is half of the corresponding stage of the building semantic segmentation decoder; Step S32, designing a fusion module of the DSM generator decoder, wherein for each stage of the DSM generator decoder, its input includes the output of the previous stage of the DSM generator decoder, the corresponding stage output of the semantic segmentation decoder, and the output of the corresponding stage of the feature extraction skeleton network; Step S33, designing the DSM discriminator; The DSM discriminator is used to identify the authenticity of the DSM generated by the DSM generator, uses ResNet50 as the basic skeleton, and includes 5 stages, wherein the second stage, the fourth stage, and the fifth stage of the DSM discriminator all use a dual-branch feature extraction module, and the dual-branch feature extraction module implements channel dimension upgrading through a 1×1 conventional convolution on the input. The input of the first branch of the dual-branch feature extraction module is a 1 / 2 channel feature, and a 3×3 depth-separable convolution is used to implement feature extraction. The second branch of the dual-branch feature extraction module adds another 1 / 2 channel feature to the output of the first branch of the dual-branch feature extraction module as input, and outputs it after a depth-separable dilated convolution with a convolution kernel of 3 and a dilation rate of 2. Finally, the two branches of the dual-branch feature extraction module are channel-fused, and feature extraction is implemented through a 1×1 convolution.
4. The intelligent building extraction method based on remote sensing images according to claim 3 is characterized in that: The feature fusion and enhancement module design in step S4 includes the following steps: Step S41, designing the MASPP module after the feature extraction skeleton network output; The MASPP module adopts multi-branch parallelism. The input of the MASPP module is F pin , the first branch of the MASPP module is a conventional convolution with a convolution kernel of 1×1, and the output of the first branch of the MASPP module is F p1 ; The second branch of the MASPP module is a conventional convolution with a convolution kernel of 3×3, and the output of the second branch of the MASPP module is F p2 ; The input of the third branch of the MASPP module is F pin With F p2 The output of the third branch of the MASPP module can be expressed as follows: F p3_1 =Aconv 3,2 (AP3(F pin +F p2 )) F p3 =F p3_1 +F p3_2 +AP3(F pin +F p2 ) Among them, AP3 represents the average pooling with a kernel of 3, Aconv 3,2 Aconv represents a dilated convolution with a convolution kernel of 3 and a dilation rate of 2. 3,3 It represents a dilated convolution with a convolution kernel of 3 and a dilation rate of 3. Indicates that features are fused based on channels; The input of the fourth branch of the MASPP module is F pin With F p3 The fusion output can be expressed as follows: F p4_1 =Aconv 3,4 (AP5(F pin +F p3 )) F p4 =F p4_1 +F p4_2 +AP5(F pin +F p3 ) Among them, AP5 represents the average pooling with a kernel of 5, and Aconv 3,4 Aconv represents a dilated convolution with a kernel size of 3 and a dilation rate of 4. 3,5 It is represented as a dilated convolution with a convolution kernel of 3 and a dilation rate of 5; the fifth branch of the MASPP module is global average pooling, and the output of the fifth branch of the MASPP module is represented as F p5 ; The output of the parallel branch of the MASPP module uses channel fusion to combine F p1 、F p2 、F p3 、F p4 and F p5 Fusion is performed to achieve information fusion between channels through 1×1 convolution, so that the channel dimension is consistent with the module input, and the output is represented as F piout ; By strengthening the effective features through the channel attention mechanism, the output of the MASPP module is: F pout =F piout +SEB(F piout ) Where SEB(.) represents the fusion channel attention mechanism of mean pooling and maximum pooling; Step S42, designing the DSATT module; The input of the DSATT module is the output feature F of the semantic segmentation decoder seg And the output feature F of the DSM generator decoder at the corresponding stage dsm , the output process of the DSATT module can be expressed as follows: To seg =S(Linear(ReLU(Linear(GAP(F seg ))))) To dsm =S(Linear(ReLU(Linear(GAP(F DSM ))))) Among them, F segout 、F dsmout represents the output after the DSATT module, GAP(.) represents global average pooling, Linear(.) represents linear change, S(.) represents activation function, Att seg With Att dsm Denote the channel attention weights of the output features of the semantic segmentation decoder and the channel attention weights of the output features of the DSM generator decoder, Att seg With Att dsm The sum is 1.
5. The intelligent building extraction method based on remote sensing images according to claim 4 is characterized in that: In the step S5, it specifically includes: Step S51, designing the loss function of the remote sensing image building semantic segmentation network, the total loss of the remote sensing image building semantic segmentation network includes pixel-level loss and regional loss, and the calculation formula of the pixel-level loss is as follows: Among them, y ic represents the true label at position i, expressed as one-hot. It is 1 when the true label is a building, otherwise it is 0. C is the number of categories, which is 2 here, and N is the total number of pixels in the predicted image. ic represents the probability of classifying sample i as category c; w c is the weight parameter, and the number of pixels of buildings in all samples of the training set is calculated to calculate w c , which is expressed as follows: The calculation formula of the area loss is as follows: Among them, y pred represents the probability of predicting category c, y true Represents the true one-hot label; The total loss of the remote sensing image building semantic segmentation network is: L seg =αL wce +βL softIoU ; Step S52: Design a loss function of the building DSM estimation network. The total loss of the building DSM estimation network includes adversarial loss and height estimation loss. The adversarial loss is expressed as follows: L G =-E(logD(G(x input ))) L D =-E(logD(x input_dsm ))-E(log(1-D(G(x input )))) L adv =L G +L D Where G and D represent the DSM generator encoder and the DSM generator decoder respectively, E(.) represents the mathematical expectation, x input_dsm represents the real DSM of the input remote sensing image, x input Represents the input remote sensing image; The height estimation loss is expressed as follows: The building DSM estimates the total network loss to be: L dsm =γ1L adv +γ2L h ; Among them, γ1 and γ2 represent the loss weights of adversarial loss and height estimation loss, respectively.
6. The intelligent building extraction method based on remote sensing images according to claim 1 is characterized in that: In step S1, the high-resolution dataset includes a Vaihingen dataset and a Potsdam dataset, and each sample in the high-resolution dataset includes an RGB image and its corresponding DSM image; the samples of the Vaihingen dataset and the Potsdam dataset are preprocessed and cropped into slices of 521 pixels×512 pixels.
7. An electronic device, characterized in that: include: One or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the intelligent building extraction method based on remote sensing images as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, implement the intelligent building extraction method based on remote sensing images as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Remote sensing image building extraction method based on multi-scale feature fusion and enhancement
CN114387512A
Remote sensing building semantic segmentation method and system based on multi-scale region attention
CN115205672A