A method for monomerization segmentation of buildings and roads, an electronic device, and a storage medium

By constructing a depth multi-scale feature extraction module, a multi-head attention mechanism and a deep and shallow multi-scale hollow space pyramid module, combined with a multi-loss function optimization model, the problems of fuzzy boundary and low accuracy in monomerization segmentation of buildings and roads are solved, and high-precision monomerization segmentation is achieved.

CN119091303BActive Publication Date: 2025-07-04HARBIN AEROSPACE STAR DATA SYST TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411234274.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-07-04
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

The prior art has problems such as blurred boundary, excessive smoothness and discontinuity, and low accuracy in monolithic segmentation of buildings and roads, especially in the absence of multi-scale feature mining and feature loss and misalignment.

Method used

The building road monolithic segmentation model DSMALnet is used to extract multi-scale feature extraction module DE_MSCAN, multi-head attention mechanism module Mhead_BRN, deep and shallow multi-scale hollow space pyramid module DS_MASPP and multi-loss function to build a building road monolithic segmentation model DSMALnet, and perform multi-level and multi-scale feature information extraction and optimization.

Benefits of technology

It realizes efficient and accurate monolithic segmentation of buildings and roads, improves the accuracy and segmentation capabilities of the model, and is suitable for monolithic segmentation of buildings and roads in multi-time phase satellite remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091303B_ABST
    Figure CN119091303B_ABST
Patent Text Reader

Abstract

A method for segmenting buildings and roads into individual entities, an electronic device, and a storage medium, belonging to the technical field of building entity recognition. To solve the problem of blurred boundaries in the segmentation of buildings and roads into individual entities, the present invention collects GF-1 2-meter multispectral image data of buildings and roads, and uses manual creation to generate a dataset of building and road entity labels; constructs a deep multi-scale feature extraction module; constructs a multi-head attention mechanism module for buildings and roads; forms a feature hierarchical extraction module; constructs a deep and shallow multi-scale atrous spatial pyramid module, and forms a feature decoding module; designs a multi-loss function to guide the optimization of the building-road entity segmentation model; constructs a building-road entity segmentation model DSMALnet, conducts model training, and obtains an optimal building-road entity segmentation model. The method of the present invention efficiently and accurately realizes the segmentation of buildings and roads into individual entities in multi-temporal satellite remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of building monomer recognition, and specifically relates to a method for segmenting buildings and roads into monomers, an electronic device, and a storage medium. Background Art

[0002] With the accelerating urbanization process, construction land has also changed rapidly. Timely and accurately monitoring the changes in construction land can accurately grasp the state of urban development. The relevant information such as the quantity, area, location, and ownership of rural houses is important basic data for applications such as rural land use planning, land use change monitoring, new rural construction, and the renovation of hollow villages. At present, the traditional methods for extracting construction land information from remote sensing images are mainly divided into three categories: The first category is based on the relevant knowledge of the spectral characteristics of typical ground objects or spectral relationships, and uses supervised classification methods, logical discrimination methods, or decision tree methods to extract construction land information. Although these methods can remove bare land and extract construction land information, due to the large differences in spectral relationship characteristics in different regions or different images, the efficiency of extracting construction land information and the universality of the model are restricted. The second category is to extract construction land information based on classification techniques. Due to the spectral heterogeneity of construction land, the extraction accuracy of this method is often not high, and subsequent processing is required to improve the accuracy. The third category is to automatically extract construction land information by establishing a remote sensing index model. This method has inaccurate extraction accuracy for building road monomers.

[0003] With the development of artificial intelligence, higher requirements have been put forward for the high-precision, high-efficiency, and automated monomer extraction of buildings and roads. With the development of deep learning technology, traditional building-level road segmentation methods are gradually fading out of people's vision. Due to the ability of convolutional neural networks in mining context information and multi-scale feature learning, they have become the standard methods for segmenting buildings and roads. However, the widely used encoder-decoder neural network faces two technical obstacles. The first problem is insufficient multi-scale feature mining, which is prone to false detection and missed detection. The second problem is that features will be lost and misaligned during the encoding and decoding processes. During the extraction process, low-level features will be weakened, and at the same time, features at different scales will also be misaligned. Summary of the Invention

[0004] The problem to be solved by the present invention is the problem of fuzzy, overly smooth and discontinuous, and low-precision boundaries in the segmentation of buildings and roads into monomers. A method for segmenting buildings and roads into monomers, an electronic device, and a storage medium are proposed.

[0005] To achieve the above object, the present invention is realized through the following technical solutions:

[0006] A method for segmenting buildings and roads into monomers, comprising the following steps:

[0007] S1. Collect the GF-1 2-meter multispectral image data of buildings and roads, and use manual methods to create a dataset of building and road monomerization labels;

[0008] S2. Construct a deep multi-scale feature extraction module DE_MSCAN;

[0009] S3. Construct a multi-head attention mechanism module Mhead_BRN for buildings and roads;

[0010] S4. Based on the deep multi-scale feature extraction module DE_MSCAN constructed in step S2 and the multi-head attention mechanism module Mhead_BRN for buildings and roads constructed in step S3, form a feature hierarchical extraction module encoder;

[0011] S5. Construct a deep and shallow multi-scale atrous spatial pyramid module DS_MASPP, use the LightHamHead and MLP methods for decoding, and form a feature decoding module decoder;

[0012] S6. Design a multi-loss function to guide the optimization of the building-road monomerization segmentation model;

[0013] S7. Based on the feature hierarchical extraction module encoder constructed in step S4, the feature decoding module decoder constructed in step S5, and the multi-loss function designed in step S6, construct a building-road monomerization segmentation model DSMALnet;

[0014] S8. Input the building and road monomerization label dataset made in step S1 into the building-road monomerization segmentation model constructed in step S7 for model training to obtain the optimal building-road monomerization segmentation model;

[0015] S9. Use the optimal building-road monomerization segmentation model obtained in step S8 to perform building-road segmentation of the target area.

[0016] Further, the specific implementation method of step S2 includes the following steps:

[0017] S2.1. Set the DE_MSCAN module to be composed of 2 batch normalization modules BN, 1 multi-head attention mechanism module Mhead_BRN for buildings and roads, and 1 framework flexible network module FFN. The calculation process expression is:

[0018] s DE_MSCAN = Mhead_BRN(BN(input t )) + input t (1)

[0019] x DE_MSCAN = FFN(BN(sDE_MSCAN )) + s DE_MSCAN (2)

[0020] Among them, input t represents the input feature information, and s DE_MSCAN represents the feature information after BN and Mhead_BRN operations and matrix addition with the input feature information. x DE_MSCAN represents the feature information after BN and FFN operations and matrix addition with s DE_MSCAN . BN(.) represents the batch normalization operation, Mhead_BRN(.) represents the calculation of the attention mechanism module, and FFN(.) represents the calculation of the FFN module;

[0021] S2.2. The expression of the FFN module calculation process is:

[0022] FFN = Conv 1×1 (GELU(Dw_Conv 3×3 (Conv 1×1 (s DE_MSCAN )))) (3)

[0023] Among them, Conv 1×1 (.) represents the 1×1 convolution operation, GELU(.) represents the activation function operation, and Dw_Conv 3×3 (.) represents the 3×3 depth convolution operation.

[0024] Furthermore, the specific implementation method of step S3 includes the following steps:

[0025] S3.1. Set the Mhead_BRN module to be composed of 2 1×1 convolutions, 1 GELU activation function, and 1 multi-attention mechanism Mhead_BR module. The expression of the Mhead_BRN module calculation process is:

[0026] S Mhead_BRN = GELU(Conv 1×1 (input h )) (4)

[0027] x Mhead_BRN = Conv 1×1 (Mhead_BR(S Mhead_BRN )) (5)

[0028] Among them, s Mhead_BRN represents the feature information after 1×1 convolution and GELU activation function operations, and x Mhead_BRN represents the feature information after the Mhead_BR module and 1×1 convolution operations. Conv 1×1(.) represents a 1×1 convolution operation, GELU(.) represents an activation function operation, and Mhead_BR(.) represents the calculation of the Mhead_BR module.

[0029] S3.2. Set the Mhead_BR module to consist of three parts, and the specific structure is as follows:

[0030] S3.2.1. Set the first part as a multi-scale spatial and channel attention mechanism, which consists of seven depth convolutions with different scales. The expression for the calculation process is:

[0031] s dw5 = Dw_Conv 5×5 (s Mhead_BRN )(6)

[0032]

[0033] Among them, s dw5 , s dw7 , s dw11 , s dw21 respectively represent the feature information of depth convolutions passing through 5×5, 7×7, 11×11, and 21×21 convolution kernels. Dw_Conv i×j represents the depth convolution operation of an i×j convolution kernel, where i, j ∈ [1, 5, 7, 11, 21];

[0034] S3.2.2. Set the second part as a multi-scale position attention mechanism, which consists of two average poolings, two max poolings, four normalized attention modules NAM, and one 1×1 convolution. The expression for the calculation process is:

[0035]

[0036]

[0037] H_NAM, W_NAM = split(Conv 1×1 (concat(s h_d , s w_d T ))) (10)

[0038] x pos = H_NAM × W_NAM (11)

[0039] Among them, s h_avgn , s h_maxn , S w_avgn , S w_maxn respectively represent the features obtained by performing NAM calculations after average pooling in the height direction, max pooling in the height direction, average pooling in the width direction, and max pooling in the width direction. sh_d ,S w_d respectively represent the features obtained by matrix dot multiplication of the features in the height and width directions obtained from Formula 8. H_NAM and W_NAM respectively represent the features obtained from Formula 9 after matrix concatenation, 1×1 convolution, and then matrix splitting along the height and width. x pos represents the weight matrix obtained by matrix multiplication of the features obtained from Formula 10. H_AvgPool(.) and W_AvgPool(.) respectively represent the average pooling operations in the height and width directions. H_MaxPool(.) and W_MaxPool(.) respectively represent the max pooling operations in the height and width directions. NAM(.) represents the calculation of the NAM module, Conv 1×1 (.) represents the 1×1 convolution operation, split(.) represents matrix splitting operation, T represents matrix transpose operation, and ⊙ represents matrix dot multiplication calculation;

[0040] Among them, the calculation formula of NAM is as follows:

[0041] Y(BN) = ω λ (BN(Y(F))) (12)

[0042]

[0043] Among them, Y(F) represents the input feature map, λ represents the scale factor of each channel, ω λ represents obtaining the weight, μ B and σ B respectively represent the mean and standard deviation of the mini-batch, r and β respectively represent the trainable finite transformation scale and offset, BN(s) represents the scale factor, ∈ represents the penalty factor, Bin out represents the output feature of the previous layer;

[0044] S3.2.3. Add the 4 weight matrices obtained in the first part and the 1 weight matrix obtained in the second part through matrix addition, and then perform a 1×1 convolution operation. Multiply the obtained weight matrix with the input feature information. The expression of the calculation process is:

[0045] x dw = Conv 1×1 (s dw5 + s dw7 + s dw11 + s dw21 + x pos ) × s Mhead_BRN (15)

[0046] Among them, x dw represents the feature information after the operation of the Mhead_BRN module.

[0047] Further, the specific implementation method of step S4 is to perform a convolution operation with a 3×3 convolution kernel and a stride of 2 on the input data set to obtain the calculation result of the 0th stage, Stage0; then perform calculations in 4 stages on Stage0, that is, perform the DE_MSCAN operation designed in step S2 2 times, 2 times, 4 times, and 2 times. After each stage, connect a 3×3 convolution kernel and a convolution operation with a stride of 2 for downsampling to obtain the calculation results of the 4 stages, which are the calculation result of the 1st stage, Stage1, the calculation result of the 2nd stage, Stage2, the calculation result of the 3rd stage, Stage3, and the calculation result of the 4th stage, Stage4, respectively.

[0048] Further, the DS_MASPP module in step S5 is divided into two parts, and the specific implementation method is as follows:

[0049] S5.1. Set the first part of the DS_MASPP module as shallow multi-scale information fusion, which consists of 1 3×3 convolution kernel, 1 batch normalization BN, and 1 Relu6 activation function. The expression of the calculation process is:

[0050] Stage0_1 = Relu6(BN(Conv 3×3 (concat(Stage0, bicubic(Stage1)))))(16)

[0051] Among them, Stage0_1 represents the shallow multi-scale information fusion feature, bicubic(.) represents the bicubic interpolation operation, concat(.,.) represents the matrix connection operation, Conv 3×3 (.) represents the 3×3 convolution operation, BN(.) represents the batch normalization operation, and Relu6(.) represents the Relu6 activation function operation;

[0052] S5.2. Set the second part of the DS_MASPP module as medium and deep multi-scale information fusion, which consists of 3 3×3 dilated convolutions with dilation rates of 3, 5, and 7 respectively, 4 batch normalizations BN, 4 relu6 activation functions, and 2 3×3 convolutions. The expression of the calculation process is:

[0053] Stage2_4 = Relu6(BN(Conv 3×3 concat(Stage2, bicubic(Stage3, Stage4))))(17)

[0054]

[0055] Stage d = Conv 3×3 (concat(srate3 , s rate5 , s rate7 )) (19)

[0056] Among them, Stage2_4 represents the feature after connecting the middle and deep features, s rate3 , s rate5 , s rate7 respectively represent the feature information after the dilated convolution operations with dilation rates of 3, 5, and 7. Stage d represents the feature after fusing the middle and deep multi-scale information. D_Conv rate=3 , D_Conv rate=5 , D_Conv rate=7 respectively represent the dilated convolution operations with dilation rates of 3, 5, and 7;

[0057] S5.3. Store the Stage0_1 obtained in step S5.1 and the Stage d obtained in step 5.2 in the List list, and then input them into the LightHamHead and MLP for feature decoding to obtain the predicted building road result y out .

[0058] Furthermore, the specific implementation method of step S6 is to perform multi-class cross-entropy loss label calculation, Dice loss function out calculation, and edge detection loss function calculation on the manually made building and road monomerization label data y obtained in step S1 and the y obtained in step S5 respectively, and then add the losses to obtain the multi-loss function to guide the model optimization. The calculation expression is:

[0059]

[0060] The calculation formula of the multi-class cross-entropy loss is as follows:

[0061]

[0062] Among them, M represents the number of categories; y o,c represents a binary indicator, which is 1 if category c is the correct category of the object, and 0 otherwise; p o,c represents the probability that the model predicts that the object belongs to category c;

[0063] The calculation formula of the Dice loss function is as follows:

[0064]

[0065] Among them, F p represents the predicted label set; F gt represents the true label set;

[0066] Edge detection loss function The calculation formula is as follows:

[0067]

[0068] Among them, the Laplace operator is used to perform edge extraction on the label y label and the predicted y out The background is 0 and the edge is 1. The calculation formula is as follows:

[0069]

[0070] Among them, f(k, u) is the pixel point in the image, and f(k - 1, u), f(k + 1, u), f(k, u - 1), and f(k, u + 1) are the four adjacent pixel points around f(k, u) respectively.

[0071] Furthermore, step S7 for constructing the building road monomerization segmentation model DSMALnet is divided into two stages: an encoder and a decoder; the encoder consists of 5 3×3 convolutions for downsampling and 10 DE_MSCANs. After the first four downsamplings, 2, 2, 4, and 2 DE_MSCAN calculations are performed respectively to achieve feature information extraction in multiple scales, multiple spaces, and multiple positions; the decoder inputs the feature information of each downsampling into the DS_MASPP module, and then inputs it into the lightweight Hamburger and the multi-layer perceptron MLP for feature decoding to obtain the prediction result; the loss function uses a multi-loss function to evaluate the model loss.

[0072] Furthermore, in step S8, the building and road monomerization label data set made in step S1 is input into the building road monomerization segmentation model DSMALnet constructed in step S7 for model training to obtain the optimal building road monomerization segmentation model; the average intersection over union mIoU, class average pixel accuracy mPA, overall classification accuracy Accuracy, precision Precision, recall Recall, and F1 score are used to evaluate the segmentation accuracy of the model.

[0073] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the building and road monomerization segmentation method are implemented.

[0074] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the described method for segmenting buildings and roads into individual entities is implemented.

[0075] Advantages of the present invention:

[0076] For the method for segmenting buildings and roads into individual entities described in the present invention, a depth multi-scale feature extraction module DE_MSCAN is designed to achieve multi-level and multi-scale feature information extraction and a more effective correlation between feature information and target information; a multi-head attention mechanism module Mhead_BRN for buildings and roads is designed to obtain and share context information in multiple dimensions and scales, effectively retain high-quality information, and more accurately guide the collection of model feature information; a deep and shallow multi-scale atrous spatial pyramid module DS_MASPP is designed to improve the model's segmentation ability for features of different sizes and shapes; and multiple loss functions are designed to guide the model to perform high-precision and high-efficiency optimization. The method of the present invention can efficiently and accurately segment buildings and roads in multi-temporal satellite remote sensing images, providing certain technical support for many fields such as urban planning, rural renovation, non-agricultural and non-grain, and ecological environmental protection. Description of the Drawings

[0077] Figure 1 is a flowchart of the method for segmenting buildings and roads into individual entities described in the present invention;

[0078] Figure 2 is an architecture diagram of the building and road segmentation model DSMALnet described in the present invention;

[0079] Figure 3 is an architecture diagram of the depth multi-scale feature extraction module DE_MSCAN described in the present invention;

[0080] Figure 4 is an architecture diagram of the multi-head attention mechanism module Mhead_BRN for buildings and roads described in the present invention;

[0081] Figure 5 is an architecture diagram of the deep and shallow multi-scale atrous spatial pyramid module DS_MASPP described in the present invention;

[0082] Figure 6 is an accuracy evaluation diagram of the optimal building and road segmentation model described in the present invention, where (a) is the mIoU evaluation diagram, (b) is the mPA evaluation diagram, (c) is the Accuracy evaluation diagram, (d) is the Precision evaluation diagram, (e) is the Recall evaluation diagram, and (f) is the F1-score evaluation diagram;

[0083] Figure 7This is a comparison chart between the achievement map of the building road monomerization segmentation model described in the present invention and the achievement maps of current mainstream models. Detailed implementation manners

[0084] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in combination with the accompanying drawings and specific implementation manners. It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described specific implementation manners are only a part of the implementation manners of the present invention, rather than all the specific implementation manners. Usually, the components of the specific implementation manners of the present invention described and shown in the accompanying drawings herein can be arranged and designed in various different configurations, and the present invention can also have other implementation manners.

[0085] Therefore, the detailed description of the specific implementation manners of the present invention provided in the accompanying drawings below is not intended to limit the scope of the present invention to be protected, but only represents the selected specific implementation manners of the present invention. All other specific implementation manners obtained by those skilled in the art based on the specific implementation manners of the present invention without creative efforts belong to the scope of protection of the present invention.

[0086] To further understand the content, features and effects of the present invention, the following specific implementation manners are exemplified and are accompanied by Figure 1 - Appendix Figure 7 The details are described as follows:

[0087] Example 1:

[0088] A method for segmenting buildings and roads into monomers includes the following steps:

[0089] S1. Collect the GF-1 2-meter multispectral image data of buildings and roads, and use manual methods to create a label dataset for building and road monomerization;

[0090] Further, in step S1, based on the GF-1 2-meter multispectral image data of the main urban area of Harbin, use ArcGIS 10.7 to perform vector annotation for building and road monomerization, and then use the vector-to-raster function to convert the annotated vector data into tiff raster data as label data, where 0 represents the background, 1 represents the road, and 2 represents the building. Cut the GF-1 2-meter multispectral image data and the created label data into 512×512 sizes, and use 70% as the training dataset and 30% as the test dataset, and finally complete the production of the label dataset for building and road monomerization;

[0091] S2. Construct a deep multi-scale feature extraction module DE_MSCAN;

[0092] Further, the specific implementation method of step S2 includes the following steps:

[0093] S2.1. Set the DE_MSCAN module to consist of 2 batch normalization modules BN, 1 multi-head attention mechanism module Mhead_BRN for building roads, and 1 frame flexible network module FFN. The calculation process expression is:

[0094] s DE_MSCAN = Mhead_BRN(BN(input t )) + input t (1)

[0095] x DE_MSCAN = FFN(BN(s DE_MSCAN )) + s DE_MSCAN (2)

[0096] Among them, input t represents the input feature information, S DE_MSCAN represents the feature information after BN and Mhead_BRN operations and matrix addition with the input feature information, x DE_MSCAN represents the feature information after BN and FFN operations and matrix addition with s DE_MSCAN ; BN(.) represents the batch normalization operation, Mhead_BRN(.) represents the calculation of the attention mechanism module, and FFN(.) represents the calculation of the FFN module;

[0097] S2.2. The calculation process expression of the FFN module is:

[0098] FFN = Conv 1×1 (GELU(Dw_Conv 3×3 (Conv 1×1 (s DE_MACAN )))) (3)

[0099] Among them, Conv 1×1 (.) represents the 1×1 convolution operation, GELU(.) represents the activation function operation, and Dw_Conv 3×3 (.) represents the 3×3 depth convolution operation.

[0100] S3. Construct the multi-head attention mechanism module Mhead_BRN for building roads;

[0101] Furthermore, the specific implementation method of step S3 includes the following steps:

[0102] S3.1. Set the Mhead_BRN module to consist of 2 1×1 convolutions, 1 GELU activation function, and 1 multi-attention mechanism Mhead_BR module. The calculation process expression of the Mhead_BRN module is:

[0103] s Mhead_BRN = GELU(Conv 1×1 (input h )) (4)

[0104] x Mhead_BRN = Conv 1×1 (Mhead_BR(s Mhead_BRN )) (5)

[0105] Among them, s Mhead_BRN represents the feature information after the 1×1 convolution and GELU activation function operations, and x Mhead_BRN represents the feature information after the 1×1 convolution operation of the Mhead_BR module, Conv 1×1 (.) represents the 1×1 convolution operation, GELU(.) represents the activation function operation, and Mhead_BR(.) represents the calculation of the Mhead_BR module;

[0106] S3.2. Set the Mhead_BR module to consist of 3 parts, and the specific structure is as follows:

[0107] S3.2.1. Set the first part as the multi-scale spatial and channel attention mechanism, which consists of 7 depth convolutions of different scales, and the expression of the calculation process is:

[0108] s dw5 = Dw_Conv 5×5 (s Mhead_BRN ) (6)

[0109]

[0110] Among them, s dw5 、s dw7 、s dw11 、s dw21 respectively represent the feature information of the depth convolution with 5×5, 7×7, 11×11, and 21×21 convolution kernels, and Dw_Conv i×j represents the depth convolution operation of the i×j convolution kernel, where i, j ∈ [1, 5, 7, 11, 21];

[0111] S3.2.2. Set the second part as the multi-scale position attention mechanism, which consists of 2 average poolings, 2 max poolings, 4 normalized attention modules NAM, and 1 1×1 convolution, and the expression of the calculation process is:

[0112]

[0113] H_NAM, W_NAM = split(Conv 1×1 (concat(sh_d , s w_d T ))) (10)

[0114] x pos = H_NAM × W_NAM (11)

[0115] where s h_avgn , s h_maxn , s w_avgn , s w_maxn respectively represent the features obtained by performing average pooling in the height direction, max pooling in the height direction, average pooling in the width direction, and max pooling in the width direction followed by NAM calculation. s h_d , s w_d respectively represent the features obtained by performing matrix dot multiplication on the features in the height and width directions obtained from Equation 8. H_NAM and W_NAM respectively represent the features obtained by performing matrix concatenation, 1×1 convolution on the features obtained from Equation 9, and then performing matrix splitting along the height and width. x pos represents the weight matrix obtained by performing matrix multiplication on the features obtained from Equation 10. H_AvgPool(.) and W_AvgPool(.) respectively represent the average pooling operations in the height and width directions. H_MaxPool(.) and W_MaxPool(.) respectively represent the max pooling operations in the height and width directions. NAM(.) represents the NAM module calculation. Conv 1×1 (.) represents the 1×1 convolution operation. split(.) represents the matrix splitting operation. T represents the matrix transpose operation. ⊙ represents the matrix dot multiplication calculation;

[0116] where the NAM calculation formula is as follows:

[0117] Y(BN) = ω λ (BN(Y(F))) (12)

[0118]

[0119] where Y(F) represents the input feature map. λ represents the scale factor for each channel. ω λ represents obtaining the weight. μ B and σ B respectively represent the mean and standard deviation of the mini-batch. r and β respectively represent the trainable finite transformation scale and offset. BN(s) represents the scale factor. ∈ represents the penalty factor. Bin out represents the output feature of the previous layer;

[0120] S3.2.3. Add the four weight matrices obtained in the first part and the one weight matrix obtained in the second part through matrix addition, and then perform a 1×1 convolution operation. Multiply the obtained weight matrix with the input feature information. The expression of the calculation process is:

[0121] x dw =Conv 1×1 (s dw5 +s dw7 +s dw11 +s dw21 +x pos )×s Mhead_BRN (15)

[0122] Among them, x dw represents the feature information after the operation of the Mhead_BRN module.

[0123] S4. Based on the deep multi-scale feature extraction module DE_MSCAN constructed in step S2 and the multi-head attention mechanism module Mhead_BRN for building roads constructed in step S3, form a feature hierarchical extraction module encoder;

[0124] Furthermore, the specific implementation method of step S4 is to perform a convolution operation with a 3×3 convolution kernel and a stride of 2 on the input data set to obtain the calculation result Stage0 of the 0th stage; perform 4 stages of calculations on Stage0, that is, perform the DE_MSCAN operations designed in step S2 2 times, 2 times, 4 times, and 2 times. After each stage, connect a convolution operation with a 3×3 convolution kernel and a stride of 2 for downsampling to obtain the calculation results of the 4 stages, which are the calculation result Stage1 of the 1st stage, the calculation result Stage2 of the 2nd stage, the calculation result Stage3 of the 3rd stage, and the calculation result Stage4 of the 4th stage;

[0125] S5. Construct a deep and shallow multi-scale atrous spatial pyramid module DS_MASPP, and use the LightHamHead and MLP methods for decoding to form a feature decoding module decoder;

[0126] Furthermore, the DS_MASPP module in step S5 is divided into two parts, and the specific implementation method is as follows:

[0127] S5.1. Set the first part of the DS_MASPP module as shallow multi-scale information fusion, which consists of 1 3×3 convolution kernel, 1 batch normalization BN, and 1 Relu6 activation function. The expression of the calculation process is:

[0128] Stage0_1 = Relu6(BN(Conv 3×3(concat(Stage0,bicubic(Stage1))))) (16)

[0129] Among them, Stage0_1 represents the shallow multi-scale information fusion feature, bicubic(.) represents the bicubic interpolation operation, concat(.,.) represents the matrix concatenation operation, Conv 3×3 (.) represents the 3×3 convolution operation, BN(.) represents the batch normalization operation, and Relu6(.) represents the Relu6 activation function operation;

[0130] S5.2. Set the second part of the DS_MASPP module as the middle and deep multi-scale information fusion, which consists of 3×3 dilated convolutions with dilation rates of 3, 5, and 7 respectively, 4 batch normalizations BN, 4 relu6 activation functions, and 2 3×3 convolutions. The expression of the calculation process is:

[0131] Stage2_4 = Relu6(BN(Conv 3×3 concat(Stage2,bicubic(Stage3, Stage4)))) (17)

[0132]

[0133] Stage d = Conv 3×3 (concat(s rate3 ,s rate5 ,s rate7 )) (19)

[0134] Among them, Stage2_4 represents the feature after the connection of the middle and deep features, s rate3 ,s rate5 ,s rate7 respectively represent the feature information after the dilated convolution operations with dilation rates of 3, 5, and 7. Stage d represents the feature after the middle and deep multi-scale information fusion. D_Conv rate=3 , D_Conv rate=5 , D_Conv rate=7 respectively represent the dilated convolution operations with dilation rates of 3, 5, and 7;

[0135] S5.3. Store the Stage0_1 obtained in step S5.1 and the Stage d obtained in step 5.2 in the List list, and then input them into the LightHamHead and MLP for feature decoding to obtain the predicted building road result y out .

[0136] S6. Design multiple loss functions to guide the optimization of the building-road monomerization segmentation model;

[0137] Further, the specific implementation method of step S6 is to perform multi-class cross-entropy loss label calculation on the manually made building and road monomerization label data y out obtained in step S1 and the y obtained in step S5 respectively, calculate the Dice loss function calculate the edge detection loss function calculate, and then add the losses to obtain the multiple loss function to guide model optimization. The calculation expression is:

[0138]

[0139] The calculation formula of the multi-class cross-entropy loss is as follows:

[0140]

[0141] where M represents the number of classes; y o,c represents a binary indicator, which is 1 if class c is the correct class of the object, otherwise 0; p o,c represents the probability that the model predicts that the object belongs to class c;

[0142] The calculation formula of the Dice loss function is as follows:

[0143]

[0144] where F p represents the predicted label set; F gt represents the true label set;

[0145] The calculation formula of the edge detection loss function is as follows:

[0146]

[0147] where the Laplace operator is used to perform edge extraction on the label y label and the predicted y out . The background is 0 and the edge is 1. The calculation formula is as follows:

[0148]

[0149] Among them, f(k, u) is a pixel point in the image, and f(k - 1, u), f(k + 1, u), f(k, u - 1), and f(k, u + 1) are the four adjacent pixel points around f(k, u), respectively.

[0150] S7. Based on the feature hierarchical extraction module encoder established in step S4, the feature decoding module decoder established in step S5, and the multi-loss function designed in step S6, construct the building-road monomerization segmentation model DSMALnet;

[0151] Furthermore, the construction of the building-road monomerization segmentation model DSMALnet in step S7 is divided into two stages: an encoder and a decoder; the encoder consists of 5 3×3 convolutions for downsampling and 10 DE_MSCANs. After the first four downsamplings, 2, 2, 4, and 2 DE_MSCAN calculations are performed respectively to achieve the extraction of feature information in multiple scales, multiple spaces, and multiple positions; the decoder inputs the feature information of each downsampling into the DS_MASPP module, and then inputs it into the lightweight Hamburger and the multi-layer perceptron MLP for feature decoding to obtain the prediction result; the loss function uses the multi-loss function to evaluate the model loss.

[0152] S8. Input the building and road monomerization label dataset produced in step S1 into the building-road monomerization segmentation model constructed in step S7 for model training to obtain the optimal building-road monomerization segmentation model;

[0153] Furthermore, in step S8, input the building and road monomerization label dataset produced in step S1 into the building-road monomerization segmentation model DSMALnet constructed in step S7 for model training to obtain the optimal building-road monomerization segmentation model; use the mean intersection over union mIoU, class average pixel accuracy mPA, overall classification accuracy Accuracy, precision Precision, recall Recall, and F1 score to evaluate the segmentation accuracy of the model.

[0154] Furthermore, the specific calculation formulas are as follows:

[0155]

[0156] Among them, k represents the number of classifications, p ij represents that i is predicted as j, which is a false negative FN; p ji represents that j is predicted as i, which is a false positive FP; p ii represents that i is predicted as i, which is a true positive TP;

[0157]

[0158] Among them, k represents the number of classifications, i represents the classification category, and CPA represents the category pixel accuracy rate;

[0159]

[0160] Among them, k represents the number of classifications, i represents the actual classification category, j represents the predicted classification category, and p ij represents that i is predicted as j, which is a false negative; p ii represents that i is predicted as i, which is a true positive TP;

[0161]

[0162] Among them, TP represents the positive sample with correct predicted classification, that is, the true positive, and FP represents the negative sample with incorrect predicted classification, that is, the false positive;

[0163]

[0164] Among them, TP represents the positive sample with correct predicted classification, that is, the true positive, and FN represents the positive sample with incorrect predicted classification, that is, the false negative;

[0165]

[0166] Among them, Precision represents the accuracy rate, and Recall represents the recall rate.

[0167] S9. Use the optimal building-road monomerization segmentation model obtained in step S8 to perform building-road segmentation on the target area.

[0168] For the convenience of understanding the technical effects of the present invention, the comparison of the method of the embodiment of the present invention with the prior art is as follows in the following table:

[0169] Table 1 Comparison table of the method of the invention embodiment and the prior art effects

[0170]

[0171] It can be seen from Table 1 that compared with other publicly disclosed methods, the model accuracy of the method proposed by the present invention has been improved well, and all evaluation indicators are optimal. Among them, the overall classification accuracy OA reaches 95.19%, the average intersection over union mIoU of building monomerization segmentation reaches 73.12%, and the average intersection over union mIoU of road monomerization segmentation reaches 77.87%; the average pixel accuracy rate mPA of building monomerization reaches 85.56%, and the average pixel accuracy rate mPA of road monomerization reaches 88.20%; the accuracy rate Precision of building monomerization reaches 83.41%, and the accuracy rate Precision of road monomerization reaches 86.93%;

[0172] The recall rate of building monomerization reached 85.56%, and the recall rate of road monomerization reached 88.20%; the F1 score of building monomerization reached 84.49%, and the F1 score of road monomerization reached 87.56%.

[0173] For the convenience of understanding the technical effects of each module in the design of the present invention, ablation experiments were used to conduct comparative experiments between modules. The experimental results are shown in the following table:

[0174] Table 2 Comparative table of ablation experiments of the method of the invention embodiment

[0175]

[0176]

[0177] As can be seen from Table 2, the ablation experiment was carried out on the method proposed by the present invention. Based on the proposed deep and shallow multi-scale atrous spatial pyramid module DS_MASPP, the multi-head attention mechanism module Mhead_BRN for buildings and roads, and the multi-loss function The experimental accuracy was improved compared with the base when tested alone and in pairs, and the highest accuracy was obtained based on the combination of all the extracted modules, indicating that the model containing all the modules achieved the best performance and was helpful for the extraction of building and road monomerization.

[0178] In summary: This embodiment makes full use of the designed deep multi-scale feature extraction module DE_MSCAN to realize multi-level and multi-scale feature information extraction, and the correlation relationship between more effective feature information and target information; designs the multi-head attention mechanism module Mhead_BRN for buildings and roads to obtain and share context information in multiple dimensions and scales, effectively retain high-quality information, and more accurately guide the collection of model feature information; designs the deep and shallow multi-scale atrous spatial pyramid module DS_MASPP to improve the model's segmentation ability for features of different sizes and shapes; designs a multi-loss function to guide the model to be optimized with high precision and high efficiency. It solves the problems of fuzzy segmentation boundaries, over-smoothing and discontinuity, and low accuracy in building and road monomerization segmentation.

[0179] Embodiment 2:

[0180] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for segmenting buildings and roads into monomers described in Embodiment 1 are implemented.

[0181] The computer device of the present invention may include devices such as a processor and a memory, for example, a single-chip microcomputer including a central processing unit. Moreover, when the processor executes the computer program stored in the memory, the steps of a method for segmenting a building and a road into individual units as described above are implemented.

[0182] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0183] The memory mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0184] Embodiment 3:

[0185] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, a method for segmenting a building and a road into individual units as described in Embodiment 1 is implemented.

[0186] The computer-readable storage medium of the present invention may be any form of storage medium readable by the processor of the computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of a method for segmenting a building and a road into individual units as described above can be implemented.

[0187] The computer program includes computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0188] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0189] Although the present application has been described above with reference to specific embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for segmenting buildings and road monomers, characterized in that, It includes the following steps: S1. Collect the GF-1 2-meter multispectral image data of buildings and roads, and use manual methods to create a dataset of building and road monomerization labels; S2. Construct a deep multi-scale feature extraction module DE_MSCAN; S3. Construct a multi-head attention mechanism module Mhead_BRN for buildings and roads; S4. Based on the deep multi-scale feature extraction module DE_MSCAN constructed in step S2 and the multi-head attention mechanism module Mhead_BRN for buildings and roads constructed in step S3, form a feature hierarchical extraction module encoder; S5. Construct a deep and shallow multi-scale atrous spatial pyramid module DS_MASPP, use the LightHamHead and MLP methods for decoding, and form a feature decoding module decoder; The DS_MASPP module in step S5 is divided into two parts, and the specific implementation method is as follows: S5.

1. Set the first part of the DS_MASPP module as shallow multi-scale information fusion, which consists of 1 3×3 convolutional kernel, 1 batch normalization BN, and 1 Relu6 activation function. The expression for the calculation process is: Stage0_1 = Relu6(BN(Conv 3×3 (concat(Stage0, bicubic(Stage1)))))(16) Among them, Stage0_1 represents the shallow multi-scale information fusion feature, bicubic(.) represents the bicubic interpolation operation, concat(.,.) represents the matrix concatenation operation, Conv 3×3 (.) represents the 3×3 convolution operation, BN(.) represents the batch normalization operation, and Relu6(.) represents the Relu6 activation function operation; S5.

2. Set the second part of the DS_MASPP module as medium and deep multi-scale information fusion, which consists of 3 3×3 atrous convolutions with dilation rates of 3, 5, and 7 respectively, 4 batch normalization BNs, 4 relu6 activation functions, and 2 3×3 convolutions. The expression for the calculation process is: Stage2_4 = concat(Stage2, bicubic(Stage3, Stage4)) (17) Stage d = Conv 3×3 (concat(S rate3 ,S rate5 ,S rate7 )) (19) Among them, Stage2_4 represents the feature after the connection of medium and deep features, S rate3 , S rate5 , S rate7 respectively represent the feature information after the dilated convolution operations with dilation rates of 3, 5, and 7. Stage d represents the feature after the fusion of medium and deep multi-scale information. D_Conv rate=3 , D_Conv rate=5 , D_Conv rate=7 respectively represent the dilated convolution operations with dilation rates of 3, 5, and 7; S5.

3. Store the Stage0_1 obtained in step S5.1 and the Stage obtained in step 5.2 in a List, and then input them into LightHamHead and MLP for feature decoding to obtain the predicted building road result y d ; out ; S6. Design a multi-loss function to guide the optimization of the building-road monomerization segmentation model; S7. Based on the feature hierarchical extraction module encoder formed in step S4, the feature decoding module decoder formed in step S5, and the multi-loss function designed in step S6, construct a building-road monomerization segmentation model DSMALnet; S8. Input the building and road monomerization label dataset made in step S1 into the building-road monomerization segmentation model constructed in step S7 for model training to obtain the optimal building-road monomerization segmentation model; S9. Use the optimal building-road monomerization segmentation model obtained in step S8 to perform building-road segmentation in the target area.

2. The method for monomerization segmentation of a building and a road according to claim 1, wherein The specific implementation method of step S2 includes the following steps: S2.

1. Set the DE_MSCAN module to consist of 2 batch normalization modules BN, 1 multi-head attention mechanism module Mhead_BRN for buildings and roads, and 1 feed-forward neural network FFN. The expression for the calculation process is: s DE_MSCAN = Mhead_BRN(BN(input t )) + input t (1) x DE_MSCAN = FFN(BN(s DE_MSCAN )) + s DE_MSCAN (2) Among them, input t represents the input feature information, s DE_MSCAN represents the feature information after BN and Mhead_BRN operations and matrix addition with the input feature information, x DE_MSCAN represents the feature information after BN and FFN operations and matrix addition with s DE_MSCAN BN(.) represents the batch normalization operation, Mhead_BRN(.) represents the calculation of the attention mechanism module, and FFN(.) represents the calculation of the FFN module; S2.

2. The expression for the calculation process of the FFN module is: FFN = Conv 1×1 (GELU(Dw_Conv 3×3 (Conv 1×1 (s DE_MSCAN )))) (3) Among them, Conv 1×1 (.) represents a 1×1 convolution operation, and GELU(.) represents an activation function operation. Dw_Conv 3×3 (.) represents a 3×3 depth convolution operation.

3. A method for monomeric segmentation of buildings and roads according to claim 2, characterized in that, The specific implementation method of step S3 includes the following steps: S3.

1. Set the Mhead_BRN module to consist of 2 1×1 convolutions, 1 GELU activation function, and 1 multi-attention mechanism Mhead_BR module. The expression for the calculation process of the Mhead_BRN module is: s Mhead_BRN = GELU(Conv 1×1 (input h )) (4) x Mhead_BRN = Conv 1×1 (Mhead_BR(s Mhead_BRN )) (5) Among them, s Mhead_BRN represents the feature information after the 1×1 convolution and GELU activation function operations, and x Mhead_BRN represents the feature information after the Mhead_BR module and 1×1 convolution operations. Conv 1×1 (.) represents the 1×1 convolution operation, GELU(.) represents the activation function operation, and Mhead_BR(.) represents the calculation of the Mhead_BR module; S3.

2. Set the Mhead_BR module to consist of 3 parts, and the specific structure is as follows: S3.2.

1. Set the first part as the multi-scale spatial and channel attention mechanism, which consists of 7 depth convolutions with different scales. The expression for the calculation process is: s dw5 = Dw_Conv 5×5 (s Mhead_BRN )(6) Among them, s dw5 represents the feature information of depth convolution through a 5×5 convolution kernel, s dw7 represents the feature information of depth convolution through 1×7 and 7×1 convolution kernels, s dw11 represents the feature information of depth convolution through 1×11 and 11×1 convolution kernels, s dw21 represents the feature information of depth convolution through 1×21 and 21×1 convolution kernels, Dw_Conv i×j represents the depth convolution operation of an i×j convolution kernel, where i, j ∈ [1, 5, 7, 11, 21]; S3.2.

2. Set the second part as the multi-scale position attention mechanism, which consists of 2 average poolings, 2 max poolings, 4 normalized attention modules NAM, and 1 1×1 convolution. The expression for the calculation process is: H_NAM, W_NAM = split(Conv 1×1 (concat(s h_d , s w_d T ))) (10) x pos = H_NAM × W_NAM (11) Among them, s h_avgn , s h_maxn , s w_avgn , s w_maxn respectively represent the features obtained by performing NAM calculation after average pooling in the height direction, max pooling in the height direction, average pooling in the width direction, and max pooling in the width direction, s h_d , s w_d respectively represent the features obtained by matrix dot multiplication of the features in the height and width directions obtained by formula (8). H_NAM and W_NAM respectively represent the features obtained by matrix concatenation, 1×1 convolution, and then matrix splitting along the height and width of the features obtained by formula (9). x pos represents the weight matrix obtained by matrix multiplication of the features obtained by formula (10). H_AvgPool(.) and W_AvgPool(.) respectively represent the average pooling operations in the height and width directions. H_MaxPool(.) and W_MaxPool(.) respectively represent the max pooling operations in the height and width directions. NAM(.) represents the NAM module calculation. Conv 1×1 (.) represents the 1×1 convolution operation. split(.) represents the matrix splitting operation. T represents the matrix transpose operation. ⊙ represents the matrix dot multiplication calculation; Among them, the calculation formula of NAM is as follows: Y(BN) = ω λ (BN(Y(F))) (12) Among them, Y(F) represents the input feature map, λ represents the scale factor for each channel, and ω λ represents obtaining weights, and μ B and σ B represent the mean and standard deviation of the mini-batch respectively, r and β represent the trainable finite transformation scale and offset respectively, BN(s) represents the scale factor, ∈ represents the penalty factor, and Bin out represents the output feature of the previous layer; S3.2.

3. Matrix-add the 4 weight matrices obtained from the first part and the 1 weight matrix obtained from the second part, and then perform a 1×1 convolution operation. Multiply the obtained weight matrix with the input feature information. The expression for the calculation process is: x dw = Conv 1×1 (s dw5 + s dw7 + s dw11 + s dw21 + x pos ) × s Mhead_BRN (15) Among them, x dw represents the feature information after the operation of the Mhead_BRN module.

4. A method for monomerization segmentation of buildings and roads according to claim 3, characterized in that, The specific implementation method of step S4 is to perform a convolution operation with a 3×3 convolution kernel and a stride of 2 on the input data set to obtain the calculation result of the 0th stage, Stage0; then perform 4 stages of calculations on Stage0, that is, perform the DE_MSCAN operation designed in step S2 2 times, 2 times, 4 times, and 2 times. After each stage, connect a convolution operation with a 3×3 convolution kernel and a stride of 2 for downsampling to obtain the calculation results of the 4 stages, which are the calculation result of the 1st stage, Stage1, the calculation result of the 2nd stage, Stage2, the calculation result of the 3rd stage, Stage3, and the calculation result of the 4th stage, Stage4.

5. A method for monomerization segmentation of buildings and roads according to claim 4, characterized in that The specific implementation method of step S6 is to perform multi-class cross-entropy loss calculation on the manually produced building and road monomerization label data y obtained in step S1 label and the y obtained in step S5 out respectively, calculate the Dice loss function calculate the edge detection loss function calculate the edge detection loss function and then add the losses to obtain a multi-loss function to guide model optimization. The calculation expression is as follows: Multi-class cross-entropy loss The calculation formula is as follows: where M represents the number of classes; y o,c represents a binary indicator that is 1 if class c is the correct class of the object and 0 otherwise; p o,c represents the probability that the model predicts the object belongs to class c; Dice loss function The calculation formula is as follows: Among them, F p represents the predicted label set; F gt represents the true label set; Edge detection loss function The calculation formula is as follows: Among them, the Laplace operator is used to perform label y label and the predicted y out for edge extraction. The background is 0 and the edge is 1. The calculation formula is as follows: Among them, f(k,u) is the pixel point in the image, and f(k - 1,u), f(k + 1,u), f(k,u - 1), and f(k,u + 1) are the four adjacent pixel points around f(k,u) respectively.

6. A method for monomerization segmentation of buildings and roads according to claim 5, characterized in that, Step S7 constructs the building-road monomerization segmentation model DSMALnet, which is divided into two stages: an encoder and a decoder; the encoder consists of 5 3×3 convolutions for downsampling and 10 DE_MSCANs. After the first four downsamplings, 2, 2, 4, and 2 DE_MSCAN calculations are performed respectively to realize the extraction of feature information with multiple scales, multiple spaces, and multiple positions; the decoder inputs the feature information of each downsampling into the DS_MASPP module, and then inputs it into the lightweight Hamburger and the multi-layer perceptron MLP for feature decoding to obtain the prediction result; the loss function uses a multi-loss function to evaluate the model loss.

7. A method for monomeric segmentation of buildings and roads according to claim 6, characterized in that, In step S8, input the building and road monomerization label data set made in step S1 into the building-road monomerization segmentation model DSMALnet constructed in step S7 for model training to obtain the optimal building-road monomerization segmentation model; use the mean intersection over union mIoU, the mean pixel accuracy of classes mPA, the overall classification accuracy Accuracy, the precision Precision, the recall Recall, and the F1 score to evaluate the segmentation accuracy of the model.

8. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it realizes the steps of a building and road monomerization segmentation method according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements a method for segmenting a building and a road into individual units according to any one of claims 1-7.

Citation Information

Patent Citations

  • Instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement

    CN115797629A

  • Comprehensive evaluation method for land classification, electronic equipment and storage medium

    CN117113189A