Building height and contour collaborative extraction method
By constructing a collaborative extraction network model and using a dual-branch encoding module and a spatial frequency fusion module to fuse the SAR image and optical image data features, the problem of insufficient accuracy in building height and contour extraction in the existing technology is solved, and efficient building height and contour extraction is achieved.
Patent Information
- Application Number
- CN202510729374.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing technology for extracting building height and outline has problems such as insufficient interaction between data modal complementary features and high computational overhead, resulting in insufficient extraction accuracy.
A collaborative extraction method for building height and outline is adopted. By constructing a collaborative extraction network model, a dual-branch encoding module is used to extract features from SAR image and optical image data, and the dual-domain features are fused through the spatial-frequency fusion module. Finally, the results are output through the dual-branch decoding module.
The accuracy of building height and outline extraction is improved, the complementary features of optical image and SAR image data are fully integrated, and the computational overhead is reduced.
Smart Images

Figure CN120673080A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban planning and construction, and in particular to a method for collaboratively extracting building heights and outlines. Background Art
[0002] As a crucial parameter of a city's three-dimensional spatial structure, accurate extraction of building height has become a key technological breakthrough in modern urban science. As two-dimensional geographic information increasingly fails to meet the demands of smart city development, the acquisition of three-dimensional building data not only marks a revolution in urban surveying and mapping technology but also profoundly impacts the transformation and upgrading of urban governance paradigms. Building height data supports the quantitative study of urban morphology, offering new solutions to traditional challenges such as the urban heat island effect and visual corridor control. Furthermore, research on the correlation between building height and energy consumption reveals the laws of vertical urban development, providing decision-making support for sustainable urban development. At the juncture of urban digital transformation, the value of building height extraction technology has transcended the boundaries of traditional surveying and mapping, reshaping the dimensions of urban cognition. Building height data has evolved beyond a single geographic information attribute to become a crucial parameter for the operation of complex urban systems, providing a new methodology for addressing urban ills and optimizing resource allocation.
[0003] As the new urbanization strategy advances, urban built-up areas are expanding, and the proportion of high-rise buildings continues to climb. Traditional manual measurement methods are no longer sufficient to meet the data collection needs of complex building complexes in megacities. Advances in Earth observation technology have made it possible to automatically extract building heights over large areas. Early approaches primarily involved shadow-based, stereo-pair-based, and Synthetic Aperture Radar (InSAR) imagery-based methods. Shadow-based methods estimate building heights based on the spatial geometric relationship between the sun, buildings, and shadows; stereo-pair-based methods estimate building heights using photogrammetric analysis of optical stereo image pairs; and InSAR image-based methods utilize high-resolution TerraSAR-X / TanDEM-X satellite-based data to obtain ground object heights through interferometry and processing. Although some institutions have generated a series of building height products using these methods, they have limitations when estimating building heights over large areas, such as inaccurate extraction due to building shadows and the difficulty in obtaining ultra-high-resolution stereo-pair data, which means InSAR methods require DSM ancillary data. Therefore, efficient and low-cost building height extraction methods are needed to address these issues. With the development of machine learning technology, machine learning methods based on open-source data have gradually replaced traditional height extraction methods. These methods can be primarily categorized by data type, including those based on optical imagery and those based on synthetic aperture radar (SAR) data. Optical imagery, with its advantages of high-resolution texture, multispectral information, and historical data, is suitable for identifying building boundaries and analyzing shadows in good weather conditions. However, its height inversion relies on indirect methods and is limited by vegetation obstruction and weather conditions. SAR, leveraging the all-weather capability and penetration characteristics of microwaves, can produce high-resolution, all-weather, and light-independent data. Its backscatter values exhibit a strong correlation with building height, opening up new avenues for building height estimation. To this end, some studies have explored the potential of Sentinel-1 SAR data for building height extraction by constructing mapping metrics and spatial regression models. However, SAR imagery lacks fine spatial detail, making it difficult to accurately depict building outlines and structures.
[0004] Optical imagery has the ability to depict building outlines, while SAR imagery offers advantages in all-weather operation and structural perception. These two methods complement each other in the problem of building height extraction. Recent studies have explored ways to collaboratively extract building outlines and heights by fusing Sentinel-1 SAR imagery with Sentinel-2 optical imagery, leveraging traditional machine learning models such as support vector machines and random forests, or deep learning models. These methods have effectively improved large-scale building height extraction. However, SAR and optical imagery exhibit significant heterogeneity and vastly different feature information. Existing studies have typically employed loosely coupled feature concatenation or data concatenation after spatial feature extraction, resulting in insufficient interaction between the complementary features of the two data modalities and prone to information redundancy. Earlier methods, however, employed convolutional neural networks to integrate multimodal information, but these methods often exhibited local sensitivity and lacked long-range dependencies, limiting their ability to integrate relevant features from both modalities for building height extraction. In contrast, Transformer-based models, characterized by their large receptive field and global sensitivity, often surpass convolutional neural networks in capturing extensive contextual information. However, these models suffer from significant computational overhead due to the quadratic growth of resources relative to the sequence length. Summary of the Invention
[0005] The present invention provides a method for collaboratively extracting building height and outline, the purpose of which is to improve the accuracy of extracting building height and outline.
[0006] In order to achieve the above object, the present invention provides a method for collaboratively extracting building height and outline, comprising:
[0007] Step 1: Acquire a building height profile extraction dataset, which includes vertical polarization band data and vertical-horizontal polarization band data in the SAR image data of the building, four band data in the optical image data of the building, reference height data and reference profile data of the building;
[0008] Step 2: Using the building height profile extraction dataset to train the constructed collaborative extraction network model to obtain a trained collaborative extraction network model;
[0009] Step 3: Input the SAR image data and optical image data of the target building into the trained collaborative extraction network model to extract information, and obtain the height extraction result and the outline extraction result of the target building;
[0010] The collaborative extraction network model includes: a dual-branch encoding module for feature extraction of SAR image data and optical image data, a spatial frequency fusion module for fusing dual-domain features, and a dual-branch decoding module for outputting building height extraction results and building contour extraction results.
[0011] Specifically, the dual-branch encoding module includes:
[0012] A first coding unit, a second coding unit, a third coding unit, a fourth coding unit, a fifth coding unit, a sixth coding unit, a seventh coding unit, and an eighth coding unit;
[0013] a first downsampling unit, a second downsampling unit, a third downsampling unit, a fourth downsampling unit, a fifth downsampling unit, and a sixth downsampling unit;
[0014] A first feature fusion unit, a second feature fusion unit, and a third feature fusion unit;
[0015] The input end of the first encoding unit and the input end of the fifth encoding unit are both input ends of the dual-branch encoding module;
[0016] The first output end of the first encoding unit and the first output end of the fifth encoding unit are both connected to the input end of the first feature fusion unit, the output end of the first feature fusion unit is connected to the second input end of the dual-branch decoding module, the second output end of the first encoding unit is connected to the input end of the first downsampling unit, the output end of the first downsampling unit is connected to the input end of the second encoding unit, the second output end of the fifth encoding unit is connected to the input end of the fourth downsampling unit, and the output end of the fourth downsampling unit is connected to the input end of the sixth encoding unit;
[0017] The first output end of the second encoding unit and the first output end of the sixth encoding unit are both connected to the input end of the second feature fusion unit, the output end of the second feature fusion unit is connected to the third input end of the dual-branch decoding module, the second output end of the second encoding unit is connected to the input end of the second downsampling unit, the output end of the second downsampling unit is connected to the input end of the third encoding unit, the second output end of the sixth encoding unit is connected to the input end of the fifth downsampling unit, and the output end of the fifth downsampling unit is connected to the input end of the seventh encoding unit;
[0018] The first output end of the third encoding unit and the first output end of the seventh encoding unit are both connected to the input end of the third feature fusion unit, the output end of the third feature fusion unit is connected to the fourth input end of the dual-branch decoding module, the second output end of the third encoding unit is connected to the input end of the third downsampling unit, the output end of the third downsampling unit is connected to the input end of the fourth encoding unit, the second output end of the seventh encoding unit is connected to the input end of the sixth downsampling unit, and the output end of the sixth downsampling unit is connected to the input end of the eighth encoding unit;
[0019] The output end of the fourth encoding unit and the output end of the eighth encoding unit are both connected to the input end of the spatial frequency fusion module.
[0020] Furthermore, the first coding unit and the fifth coding unit have the same structure;
[0021] The first coding unit includes a first convolution block, a first half-instance normalized residual block, a second half-instance normalized residual block, and a third half-instance normalized residual block connected in sequence;
[0022] The input end of the first convolution block is the input end of the first coding unit, and the output end of the third half-instance normalized residual block is the first output end and the second output end of the first coding unit.
[0023] Furthermore, the first feature fusion unit has the same structure as the second feature fusion unit and the third feature fusion unit;
[0024] The first feature fusion unit includes a splicing block, a second convolution block, a batch normalization block, and an activation function block that are connected in sequence.
[0025] Furthermore, the second coding unit has the same structure as the third coding unit, the fourth coding unit, the sixth coding unit, the seventh coding unit, and the eighth coding unit;
[0026] The second encoding unit includes a patch embedding block, a first Mamba block for feature learning, a patch inverse embedding block, and a third convolution block connected in sequence;
[0027] The input end of the patch embedding block is the input end of the second encoding unit, and the third convolution block is the output end of the second encoding unit.
[0028] Specifically, the spatial-frequency fusion module includes:
[0029] Spatial domain fusion unit, frequency domain fusion unit, adaptive spatial frequency fusion unit;
[0030] The input end of the spatial domain fusion unit is connected to the output end of the fourth encoding unit and the output end of the eighth encoding unit respectively, and the output end of the spatial domain fusion unit is connected to the input end of the adaptive spatial frequency fusion module;
[0031] The input end of the frequency domain fusion unit is connected to the output end of the fourth encoding unit and the output end of the eighth encoding unit respectively, and the output end of the frequency domain fusion unit is connected to the input end of the adaptive space-frequency fusion module;
[0032] The output end of the adaptive space-frequency fusion unit is connected to the first input end of the dual-branch decoding module.
[0033] Specifically, the spatial domain fusion unit includes:
[0034] a first cross-attention block consisting of a first normalization layer, a second normalization layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, a first reshaping layer, a second reshaping layer, a third reshaping layer, a fourth reshaping layer, a first matrix multiplier, a second matrix multiplier, and a first adder;
[0035] A spatial feedforward network consisting of a third normalization layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a first GELU activation layer, a second adder, and a third adder;
[0036] An input end of the first normalization layer is connected to an output end of the fourth encoding unit, an output end of the first normalization layer is connected to an input end of the first convolutional layer, an output end of the first convolutional layer is connected to an input end of the fourth convolutional layer, an output end of the fourth convolutional layer is connected to an input end of the first reshaping layer, and an output end of the first reshaping layer is connected to a first input end of the first matrix multiplier;
[0037] The input end of the second normalization layer is connected to the output end of the eighth encoding unit and the first input end of the first adder respectively, the output end of the second normalization layer is connected to the input end of the second convolutional layer and the input end of the third convolutional layer respectively, the output end of the second convolutional layer is connected to the input end of the fifth convolutional layer, the output end of the fifth convolutional layer is connected to the input end of the second reshaping layer, the output end of the second reshaping layer is connected to the second input end of the first matrix multiplier, and the output end of the first matrix multiplier is connected to the first input end of the second matrix multiplier;
[0038] The output end of the third convolutional layer is connected to the input end of the sixth convolutional layer, the output end of the sixth convolutional layer is connected to the input end of the third reshaping layer, the output end of the third reshaping layer is connected to the second input end of the second matrix multiplier, the output end of the second matrix multiplier is connected to the input end of the fourth reshaping layer, the output end of the fourth reshaping layer is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the second input end of the first adder, the first output end of the first adder is connected to the first input end of the third adder, and the second output end of the first adder is connected to the input end of the third normalization layer;
[0039] The output end of the third normalization layer is connected to the input end of the eighth convolutional layer and the input end of the ninth convolutional layer respectively. The output end of the eighth convolutional layer is connected to the input end of the tenth convolutional layer. The output end of the tenth convolutional layer is connected to the first input end of the second adder.
[0040] The output end of the ninth convolutional layer is connected to the input end of the eleventh convolutional layer, the output end of the eleventh convolutional layer is connected to the input end of the first GELU activation layer, the output end of the first GELU activation layer is connected to the second input end of the second adder, the output end of the second adder is connected to the input end of the twelfth convolutional layer, the output end of the twelfth convolutional layer is connected to the second input end of the third adder, and the output end of the third adder is connected to the input end of the adaptive spatial frequency fusion unit.
[0041] Specifically, the frequency domain fusion unit includes:
[0042] a second cross attention block consisting of a fourth normalization layer, a fifth normalization layer, a sixth normalization layer, a seventh normalization layer, a thirteenth convolution layer, a fourteenth convolution layer, a fifteenth convolution layer, a sixteenth convolution layer, a seventeenth convolution layer, an eighteenth convolution layer, a nineteenth convolution layer, a fifth reshaping layer, a sixth reshaping layer, a seventh reshaping layer, an eighth reshaping layer, a first Fourier transform layer, a second Fourier transform layer, a first inverse Fourier transform layer, a first element multiplier, a second element multiplier, and a fourth adder;
[0043] a feedforward network consisting of an eighth normalization layer, a ninth reshaping layer, a tenth reshaping layer, a third Fourier transform layer, a second inverse Fourier transform layer, a twentieth convolutional layer, a twenty-first convolutional layer, a twenty-second convolutional layer, a twenty-third convolutional layer, a twenty-fourth convolutional layer, a second GELU activation layer, a third element multiplier, a fourth element multiplier, and a fifth adder;
[0044] The input end of the fourth normalization layer is connected to the output end of the eighth encoding unit, the output end of the fourth normalization layer is connected to the input end of the thirteenth convolutional layer, the output end of the thirteenth convolutional layer is connected to the input end of the sixteenth convolutional layer, the output end of the sixteenth convolutional layer is connected to the input end of the fifth reshaping layer, the output end of the fifth reshaping layer is connected to the input end of the first Fourier transform layer, and the output end of the first Fourier transform layer is connected to the first input end of the first element multiplier;
[0045] An input end of the fifth normalization layer is connected to the output end of the fourth encoding unit and the first input end of the fourth adder respectively, an output end of the fifth normalization layer is connected to the input end of the fourteenth convolutional layer and the input end of the fifteenth convolutional layer respectively, an output end of the fourteenth convolutional layer is connected to the input end of the seventeenth convolutional layer, an output end of the seventeenth convolutional layer is connected to the input end of the sixth reshaping layer, an output end of the sixth reshaping layer is connected to the input end of the second Fourier transform layer, an output end of the second Fourier transform layer is connected to the second input end of the first element multiplier, an output end of the first element multiplier is connected to the input end of the first inverse Fourier transform layer, an output end of the first inverse Fourier transform layer is connected to the input end of the eighth reshaping layer, an output end of the eighth reshaping layer is connected to the input end of the sixth normalization layer, and an output end of the sixth normalization layer is connected to the first input end of the second element multiplier;
[0046] The output end of the fifteenth convolutional layer is connected to the input end of the eighteenth convolutional layer, the output end of the eighteenth convolutional layer is connected to the input end of the seventh reshaping layer, the output end of the seventh reshaping layer is connected to the second input end of the second element multiplier, the output end of the second element multiplier is connected to the input end of the nineteenth convolutional layer, and the output end of the nineteenth convolutional layer is connected to the second input end of the fourth adder;
[0047] The output end of the fourth adder is respectively connected to the input end of the seventh normalization layer and the first input end of the fifth adder, the output end of the seventh normalization layer is connected to the input end of the ninth reshaping layer, the output end of the ninth reshaping layer is connected to the input end of the third Fourier transform layer, the output end of the third Fourier transform layer is connected to the first input end of the third element multiplier, the second input end of the third element multiplier is connected to the learnable parameter matrix, the output end of the third element multiplier is connected to the input end of the second inverse Fourier transform layer, the output end of the second inverse Fourier transform layer is connected to the input end of the tenth reshaping layer, the output end of the tenth reshaping layer is respectively connected to the input end of the twentieth convolutional layer and the input end of the twenty-first convolutional layer, the output end of the twentieth convolutional layer is connected to the input end of the twenty-second convolutional layer, and the output end of the twenty-second convolutional layer is connected to the first input end of the fourth element multiplier;
[0048] The output end of the twenty-first convolutional layer is connected to the input end of the twenty-third convolutional layer, the output end of the twenty-third convolutional layer is connected to the input end of the second GELU activation layer, the output end of the second GELU activation layer is connected to the second input end of the fourth element multiplier, the output end of the fourth element multiplier is connected to the input end of the twenty-fourth convolutional layer, and the output end of the twenty-fourth convolutional layer is connected to the input end of the adaptive spatial frequency fusion unit.
[0049] Specifically, the dual-branch decoding module includes:
[0050] A first decoding unit, a second decoding unit, a third decoding unit, a fourth decoding unit, a fifth decoding unit, a sixth decoding unit, a seventh decoding unit, an eighth decoding unit, a height regression unit, and a contour extraction unit;
[0051] The first input end of the first decoding unit and the first input end of the second decoding unit are both connected to the output end of the adaptive spatial frequency fusion unit, and the second input end of the first decoding unit and the second input end of the second decoding unit are both connected to the output end of the third feature fusion unit;
[0052] The first input end of the third decoding unit is connected to the output end of the first decoding unit, the first input end of the fourth decoding unit is connected to the output end of the second decoding unit, and the second input end of the third decoding unit and the second input end of the fourth decoding unit are both connected to the output end of the second feature fusion unit;
[0053] The first input end of the fifth decoding unit is connected to the output end of the third decoding unit, the first input end of the sixth decoding unit is connected to the output end of the fourth decoding unit, and the second input end of the fifth decoding unit and the second input end of the sixth decoding unit are both connected to the output end of the third feature fusion unit;
[0054] The input end of the seventh decoding unit is connected to the output end of the fifth decoding unit, the input end of the eighth decoding unit is connected to the output end of the sixth decoding unit, the output end of the seventh decoding unit is connected to the input end of the height regression unit, the output end of the height regression unit is the first output end of the dual-branch decoding module, the output end of the eighth decoding unit is connected to the input end of the contour extraction unit, and the output end of the contour extraction unit is the second output end of the dual-branch decoding module.
[0055] Furthermore, the first decoding unit has the same structure as the second decoding unit, the third decoding unit, the fourth decoding unit, the fifth decoding unit, the sixth decoding unit, the seventh decoding unit, and the eighth decoding unit;
[0056] The first decoding unit includes a bilinear interpolation block, a third convolution block, and a second Mamba block connected in sequence;
[0057] The input end of the bilinear interpolation block is the input end of the first decoding unit, and the output end of the second Mamba block is the output end of the first decoding unit.
[0058] The above solution of the present invention has the following beneficial effects:
[0059] The present invention trains a constructed collaborative extraction network model using an acquired building height profile extraction data set to obtain a trained collaborative extraction network model; inputs SAR image data and optical image data of a target building into the trained collaborative extraction network model for information extraction to obtain a target building height extraction result and a target building profile extraction result; the collaborative extraction network model comprises: a dual-branch encoding module for extracting features from the SAR image data and the optical image data, a spatial-frequency fusion module for fusing dual-domain features, and a dual-branch decoding module for outputting a building height extraction result and a building profile extraction result; compared with the prior art, the present invention introduces feature extraction from the spatial domain and the frequency domain, promotes the full fusion of complementary features of the optical image data and the SAR image data through the spatial-frequency fusion module, and improves the accuracy of collaborative extraction of the building height and profile.
[0060] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 Schematic diagram of a flow chart of an embodiment of the present invention;
[0062] Figure 2 A schematic diagram of the structure of a collaborative extraction network model in an embodiment of the present invention;
[0063] Figure 3 2 is a schematic structural diagram of a spatial domain fusion unit in an embodiment of the present invention;
[0064] Figure 4 Schematic diagram of the structure of the frequency domain fusion unit in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] To make the technical problems, technical solutions, and advantages to be solved by the present invention more clear, the following is a detailed description with reference to the accompanying drawings and specific embodiments. It is obvious that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0066] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0067] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to a locking connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0068] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0069] In view of the existing problems, the present invention provides a method for collaboratively extracting building height and outline.
[0070] like Figure 1 As shown, an embodiment of the present invention provides a method for collaboratively extracting building height and outline, comprising:
[0071] Step 1: Acquire a building height profile extraction dataset, which includes vertical polarization band data and vertical-horizontal polarization band data in the SAR image data of the building, four band data in the optical image data of the building, reference height data and reference profile data of the building;
[0072] Step 2: Using the building height profile extraction dataset to train the constructed collaborative extraction network model to obtain a trained collaborative extraction network model;
[0073] Step 3: Input the SAR image data and optical image data of the target building into the trained collaborative extraction network model to extract information, and obtain the height extraction result and the outline extraction result of the target building.
[0074] Specifically, step 1 includes:
[0075] Obtain building height contour extraction dataset X from the existing dataset SAR ,IOpt ,H,L}, the building height profile extraction dataset includes Sentinel-1SAR image data, Sentinel-2 optical image data, and the reference height data and reference profile data of the corresponding buildings. The reference height data and reference profile data are both raster data with a resolution of 10m and are divided into 256×256 pixels. For Sentinel-1SAR image data, its vertical polarization (VV, Vertical-vertical) and vertical-horizontal polarization (VH, Vertical-horizontal) band data are taken. in, For Sentinel-2 optical image data, take the 2nd, 3rd, 4th and 8th band data. in, And reference altitude data and reference profile data
[0076] The building height profile extraction dataset X is divided into a training set and a validation set in a ratio of 8:2.
[0077] Specifically, if Figure 2 As shown, the collaborative extraction network model constructed in the embodiment of the present invention includes: a dual-branch encoding module for extracting features from SAR image data and optical image data, a spatial frequency fusion module for fusing dual-domain features, and a dual-branch decoding module for outputting building height extraction results and building contour extraction results.
[0078] Specifically, step 2 includes:
[0079] The training set is input into the collaborative extraction network model for extraction, and the building height extraction result F is obtained. h and the i-th reference building height
[0080] Use the mean square error loss function to calculate the building height extraction result F h and the i-th reference building height The loss value between
[0081] Use binary cross entropy loss function and Dice loss function to calculate the building outline extraction result F f and the i-th reference building outline The loss value between
[0082] The AdamW optimizer, fixed step learning rate scheduler StepLR and loss function are used to train the building height and outline collaborative extraction network model. The initial learning rate is set to 0.001, batch_size is set to 16, epoch is set to 100, the step size of the fixed step learning rate scheduler is set to 8, and gamma is set to 0.95 to obtain the optimized building height and outline collaborative extraction network model.
[0083] Specifically, the dual-branch encoding module includes:
[0084] A first coding unit, a second coding unit, a third coding unit, a fourth coding unit, a fifth coding unit, a sixth coding unit, a seventh coding unit, and an eighth coding unit;
[0085] a first downsampling unit, a second downsampling unit, a third downsampling unit, a fourth downsampling unit, a fifth downsampling unit, and a sixth downsampling unit;
[0086] A first feature fusion unit, a second feature fusion unit, and a third feature fusion unit;
[0087] The input end of the first encoding unit and the input end of the fifth encoding unit are both input ends of the dual-branch encoding module;
[0088] The first output end of the first encoding unit and the first output end of the fifth encoding unit are both connected to the input end of the first feature fusion unit, the output end of the first feature fusion unit is connected to the second input end of the dual-branch decoding module, the second output end of the first encoding unit is connected to the input end of the first downsampling unit, the output end of the first downsampling unit is connected to the input end of the second encoding unit, the second output end of the fifth encoding unit is connected to the input end of the fourth downsampling unit, and the output end of the fourth downsampling unit is connected to the input end of the sixth encoding unit;
[0089] The first output end of the second encoding unit and the first output end of the sixth encoding unit are both connected to the input end of the second feature fusion unit, the output end of the second feature fusion unit is connected to the third input end of the dual-branch decoding module, the second output end of the second encoding unit is connected to the input end of the second downsampling unit, the output end of the second downsampling unit is connected to the input end of the third encoding unit, the second output end of the sixth encoding unit is connected to the input end of the fifth downsampling unit, and the output end of the fifth downsampling unit is connected to the input end of the seventh encoding unit;
[0090] The first output end of the third encoding unit and the first output end of the seventh encoding unit are both connected to the input end of the third feature fusion unit, the output end of the third feature fusion unit is connected to the fourth input end of the dual-branch decoding module, the second output end of the third encoding unit is connected to the input end of the third downsampling unit, the output end of the third downsampling unit is connected to the input end of the fourth encoding unit, the second output end of the seventh encoding unit is connected to the input end of the sixth downsampling unit, and the output end of the sixth downsampling unit is connected to the input end of the eighth encoding unit;
[0091] The output end of the fourth encoding unit and the output end of the eighth encoding unit are both connected to the input end of the spatial frequency fusion module.
[0092] Specifically, the first coding unit and the fifth coding unit have the same structure;
[0093] The first coding unit includes a first convolution block, a first half-instance normalized residual block, a second half-instance normalized residual block, and a third half-instance normalized residual block connected in sequence;
[0094] The input end of the first convolution block is the input end of the first coding unit, and the output end of the third half-instance normalized residual block is the first output end and the second output end of the first coding unit.
[0095] In this embodiment of the present invention, the first convolution block in the first coding unit and the fifth coding unit is composed of a convolution layer with a convolution kernel of 3, a stride of 2, and a padding of 1; the first half-instance normalized residual block, the second half-instance normalized residual block, and the third half-instance normalized residual block in the first coding unit and the fifth coding unit are all composed of a first convolution layer, a first LeakyReLU activation function, an instance normalization layer, a second convolution layer, and a second LeakyReLU activation function connected in sequence, wherein the convolution kernels of the first convolution layer and the second convolution layer are both 3, the stride is 1, and the padding is 1;
[0096] The i-th optical image in the optical image data Input the first convolution block in the first encoding unit for convolution processing to obtain the first optical feature map SAR image data Input the first convolution block in the fifth coding unit for convolution processing to obtain the first SAR image feature map The dimensions are 1×4×w×h, The size of is 1×2×w×h, and the size of w and h are both 256;
[0097] The first optical characteristic graph The first half instance normalized residual block, the second half instance normalized residual block, and the third half instance normalized residual block in the first coding unit are sequentially input for processing to obtain a second optical feature map The first SAR image feature map The first half instance normalized residual block, the second half instance normalized residual block, and the third half instance normalized residual block in the fifth coding unit are sequentially input for processing to obtain the second SAR image feature map. The size is 1×128×
[0098]
[0099] Specifically, the structure of the first feature fusion unit is the same as that of the second feature fusion unit and the third feature fusion unit;
[0100] The first feature fusion unit includes a splicing block, a second convolution block, a batch normalization block, and an activation function block connected in sequence. The data processing process is as follows:
[0101] The feature maps output by the first and fifth encoding units are Splicing along the channel dimension, and then passing the convolution block to obtain the first fusion feature F1, whose size is
[0102] Since the second feature fusion unit and the third feature fusion unit have the same structure as the first feature fusion unit, the data processing process of the second feature fusion unit, the third feature fusion unit and the first feature fusion unit is also the same. Therefore, in order to simplify the description, the embodiment of the present invention directly gives the input and output data of the second feature fusion unit and the third feature fusion unit. The input of the second feature fusion unit is The output is F2, whose dimensions are The input of the third feature fusion unit is The output is F3, whose dimensions are
[0103] In this embodiment of the present invention, the first downsampling unit includes a convolution layer and a batch normalization layer. The convolution kernel of the convolution layer is 1 and the stride is 2. The data processing process is:
[0104] The feature map Input the first downsampling unit to get the feature map Its size is
[0105] Since the second downsampling unit, the third downsampling unit, the fourth downsampling unit, the fifth downsampling unit, and the sixth downsampling unit have the same structure as the first downsampling unit, their data processing processes are also the same. Therefore, in order to simplify the description, the embodiment of the present invention directly provides the input and output data of the second downsampling unit, the third downsampling unit, the fourth downsampling unit, the fifth downsampling unit, and the sixth downsampling unit. The input of the second downsampling unit is The output is Its size becomes The input of the third downsampling unit is The output is Its size is The input of the fourth downsampling unit is The output is Its size becomes The input of the fifth downsampling unit is The output is Its size becomes The input of the sixth downsampling unit is The output is Its size is
[0106] Specifically, the second coding unit has the same structure as the third coding unit, the fourth coding unit, the sixth coding unit, the seventh coding unit, and the eighth coding unit;
[0107] The second encoding unit includes a patch embedding block, a first Mamba block for feature learning, a patch inverse embedding block, and a third convolution block connected in sequence;
[0108] The input end of the patch embedding block is the input end of the second encoding unit, and the third convolution block is the output end of the second encoding unit.
[0109] In an embodiment of the present invention, the patch embedding block consists of a convolutional layer with a convolution kernel of 1 and a stride of 2, a tensor flattening layer, and a normalization layer connected in sequence; the first Mamba block consists of a Mamba layer, a normalization layer, and a ReLU activation function connected in sequence; the patch inverse embedding block consists of a transpose layer, and the third convolution block consists of a convolutional layer with a convolution kernel of 1 and a stride of 1, a layer normalization layer, and a ReLU activation function connected in sequence.
[0110] In the embodiment of the present invention, the principle of the second encoding unit is:
[0111] The second SAR image feature map Input the patch embedding block to obtain the third SAR image feature map Its size is
[0112] Then the third SAR image feature map Input the Mamba block to get the fourth SAR image feature map Its size is
[0113] Then the fourth SAR image feature map Input the patch inverse embedding block to obtain the fifth SAR image feature map Its size is
[0114] Finally, the fifth SAR image feature map Input the convolution block to obtain the sixth SAR image feature map Its size is
[0115] Similarly, the working principle of the sixth coding unit is:
[0116] The second optical characteristic graph Input patch is embedded into the block to obtain the third optical feature map Its size is
[0117] Then the third optical characteristic diagram Input the Mamba block to get the fourth optical feature map Its size is
[0118] Then the fourth optical characteristic diagram Input the patch inverse embedding block to obtain the fifth optical feature map Its size is
[0119] Finally, the fifth optical characteristic diagram Input the convolution block to get the sixth optical feature map Its size is
[0120] Similarly, since the second encoding unit has the same structure as the third encoding unit, the fourth encoding unit, the sixth encoding unit, the seventh encoding unit, and the eighth encoding unit, the second encoding unit has the same process of processing data as the third encoding unit, the fourth encoding unit, the sixth encoding unit, the seventh encoding unit, and the eighth encoding unit. Therefore, in order to simplify the description, the embodiment of the present invention directly provides the input data and output data of the third encoding unit, the fourth encoding unit, the sixth encoding unit, the seventh encoding unit, and the eighth encoding unit. The input of the third encoding unit is The output is the feature map Size The input of the fourth encoding unit is The output is Size The input of the seventh coding unit is The output is Size The input of the eighth coding unit is The output is Size
[0121] Specifically, the spatial frequency fusion module includes:
[0122] Spatial domain fusion unit, frequency domain fusion unit, adaptive spatial frequency fusion unit;
[0123] The input end of the spatial domain fusion unit is connected to the output end of the fourth encoding unit and the output end of the eighth encoding unit respectively, and the output end of the spatial domain fusion unit is connected to the input end of the adaptive spatial frequency fusion module;
[0124] The input end of the frequency domain fusion unit is connected to the output end of the fourth encoding unit and the output end of the eighth encoding unit respectively, and the output end of the frequency domain fusion unit is connected to the input end of the adaptive space-frequency fusion module;
[0125] The output end of the adaptive space-frequency fusion unit is connected to the first input end of the dual-branch decoding module.
[0126] Specifically, if Figure 3 As shown, the spatial domain fusion unit includes:
[0127] a first cross-attention block consisting of a first normalization layer, a second normalization layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, a first reshaping layer, a second reshaping layer, a third reshaping layer, a fourth reshaping layer, a first matrix multiplier, a second matrix multiplier, and a first adder;
[0128] A spatial feedforward network consisting of a third normalization layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a first GELU activation layer, a second adder, and a third adder;
[0129] An input end of the first normalization layer is connected to an output end of the fourth encoding unit, an output end of the first normalization layer is connected to an input end of the first convolutional layer, an output end of the first convolutional layer is connected to an input end of the fourth convolutional layer, an output end of the fourth convolutional layer is connected to an input end of the first reshaping layer, and an output end of the first reshaping layer is connected to a first input end of the first matrix multiplier;
[0130] The input end of the second normalization layer is connected to the output end of the eighth encoding unit and the first input end of the first adder respectively, the output end of the second normalization layer is connected to the input end of the second convolutional layer and the input end of the third convolutional layer respectively, the output end of the second convolutional layer is connected to the input end of the fifth convolutional layer, the output end of the fifth convolutional layer is connected to the input end of the second reshaping layer, the output end of the second reshaping layer is connected to the second input end of the first matrix multiplier, and the output end of the first matrix multiplier is connected to the first input end of the second matrix multiplier;
[0131] The output end of the third convolutional layer is connected to the input end of the sixth convolutional layer, the output end of the sixth convolutional layer is connected to the input end of the third reshaping layer, the output end of the third reshaping layer is connected to the second input end of the second matrix multiplier, the output end of the second matrix multiplier is connected to the input end of the fourth reshaping layer, the output end of the fourth reshaping layer is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the second input end of the first adder, the first output end of the first adder is connected to the first input end of the third adder, and the second output end of the first adder is connected to the input end of the third normalization layer;
[0132] The output end of the third normalization layer is connected to the input end of the eighth convolutional layer and the input end of the ninth convolutional layer respectively. The output end of the eighth convolutional layer is connected to the input end of the tenth convolutional layer. The output end of the tenth convolutional layer is connected to the first input end of the second adder.
[0133] The output end of the ninth convolutional layer is connected to the input end of the eleventh convolutional layer, the output end of the eleventh convolutional layer is connected to the input end of the first GELU activation layer, the output end of the first GELU activation layer is connected to the second input end of the second adder, the output end of the second adder is connected to the input end of the twelfth convolutional layer, the output end of the twelfth convolutional layer is connected to the second input end of the third adder, and the output end of the third adder is connected to the input end of the adaptive spatial frequency fusion unit.
[0134] In this embodiment of the present invention, the feature map output by the fourth encoding unit is As the first input of the cross attention block, it is used to generate the features of the query sequence Q, and the feature map output by the eighth encoding unit As the second input of the cross attention block, it is used to generate the features of the key sequence K and the value sequence V. The specific process is as follows: the feature map output by the fourth encoding unit After normalization and convolution processing through the first normalization layer, the first convolution layer, and the fourth convolution layer, the query sequence Q is obtained, and the feature map output by the eighth encoding unit is After normalization and convolution processing through the second normalization layer, the second convolution layer, and the fifth convolution layer, the key sequence K is obtained, and the feature map output by the eighth encoding unit is After normalization and convolution processing through the second normalization layer, the third convolution layer, and the sixth convolution layer, the value sequence V is obtained; then the query sequence Q, the key sequence K, and the value sequence V are input in parallel to the first reshaping layer, the second reshaping layer, and the third reshaping layer to be flattened into non-overlapping M×M blocks, respectively. The total number of patches is Then perform the reshape patch operation to get the reshape result The calculation formula of the cross attention block is:
[0135]
[0136] Among them, Att Spa represents the result of cross attention calculation, Reshape represents the reshaping patch operation, Softmax(·) represents the activation function, τ∈R M×M×1×1 represents the learnable temperature matrix, and B represents the relative position deviation;
[0137] For the spatial feedforward network, the cross attention calculation result is doubled by the feature channel through the third normalization layer and divided into two parallel branches. The two branches are convolved and added together, and then convolved through the twelfth convolutional layer and added together with the cross attention calculation result to obtain the spatial domain fusion feature map F spa , whose size is
[0138] Specifically, if Figure 4 As shown, the frequency domain fusion unit includes:
[0139] a second cross attention block consisting of a fourth normalization layer, a fifth normalization layer, a sixth normalization layer, a seventh normalization layer, a thirteenth convolution layer, a fourteenth convolution layer, a fifteenth convolution layer, a sixteenth convolution layer, a seventeenth convolution layer, an eighteenth convolution layer, a nineteenth convolution layer, a fifth reshaping layer, a sixth reshaping layer, a seventh reshaping layer, an eighth reshaping layer, a first Fourier transform layer, a second Fourier transform layer, a first inverse Fourier transform layer, a first element multiplier, a second element multiplier, and a fourth adder;
[0140] a feedforward network consisting of an eighth normalization layer, a ninth reshaping layer, a tenth reshaping layer, a third Fourier transform layer, a second inverse Fourier transform layer, a twentieth convolutional layer, a twenty-first convolutional layer, a twenty-second convolutional layer, a twenty-third convolutional layer, a twenty-fourth convolutional layer, a second GELU activation layer, a third element multiplier, a fourth element multiplier, and a fifth adder;
[0141] The input end of the fourth normalization layer is connected to the output end of the eighth encoding unit, the output end of the fourth normalization layer is connected to the input end of the thirteenth convolutional layer, the output end of the thirteenth convolutional layer is connected to the input end of the sixteenth convolutional layer, the output end of the sixteenth convolutional layer is connected to the input end of the fifth reshaping layer, the output end of the fifth reshaping layer is connected to the input end of the first Fourier transform layer, and the output end of the first Fourier transform layer is connected to the first input end of the first element multiplier;
[0142] An input end of the fifth normalization layer is connected to the output end of the fourth encoding unit and the first input end of the fourth adder respectively, an output end of the fifth normalization layer is connected to the input end of the fourteenth convolutional layer and the input end of the fifteenth convolutional layer respectively, an output end of the fourteenth convolutional layer is connected to the input end of the seventeenth convolutional layer, an output end of the seventeenth convolutional layer is connected to the input end of the sixth reshaping layer, an output end of the sixth reshaping layer is connected to the input end of the second Fourier transform layer, an output end of the second Fourier transform layer is connected to the second input end of the first element multiplier, an output end of the first element multiplier is connected to the input end of the first inverse Fourier transform layer, an output end of the first inverse Fourier transform layer is connected to the input end of the eighth reshaping layer, an output end of the eighth reshaping layer is connected to the input end of the sixth normalization layer, and an output end of the sixth normalization layer is connected to the first input end of the second element multiplier;
[0143] The output end of the fifteenth convolutional layer is connected to the input end of the eighteenth convolutional layer, the output end of the eighteenth convolutional layer is connected to the input end of the seventh reshaping layer, the output end of the seventh reshaping layer is connected to the second input end of the second element multiplier, the output end of the second element multiplier is connected to the input end of the nineteenth convolutional layer, and the output end of the nineteenth convolutional layer is connected to the second input end of the fourth adder;
[0144] The output end of the fourth adder is respectively connected to the input end of the seventh normalization layer and the first input end of the fifth adder, the output end of the seventh normalization layer is connected to the input end of the ninth reshaping layer, the output end of the ninth reshaping layer is connected to the input end of the third Fourier transform layer, the output end of the third Fourier transform layer is connected to the first input end of the third element multiplier, the second input end of the third element multiplier is connected to the learnable parameter matrix, the output end of the third element multiplier is connected to the input end of the second inverse Fourier transform layer, the output end of the second inverse Fourier transform layer is connected to the input end of the tenth reshaping layer, the output end of the tenth reshaping layer is respectively connected to the input end of the twentieth convolutional layer and the input end of the twenty-first convolutional layer, the output end of the twentieth convolutional layer is connected to the input end of the twenty-second convolutional layer, and the output end of the twenty-second convolutional layer is connected to the first input end of the fourth element multiplier;
[0145] The output end of the twenty-first convolutional layer is connected to the input end of the twenty-third convolutional layer, the output end of the twenty-third convolutional layer is connected to the input end of the second GELU activation layer, the output end of the second GELU activation layer is connected to the second input end of the fourth element multiplier, the output end of the fourth element multiplier is connected to the input end of the twenty-fourth convolutional layer, and the output end of the twenty-fourth convolutional layer is connected to the input end of the adaptive spatial frequency fusion unit.
[0146] In this embodiment of the present invention, the second cross attention block also converts the feature map output by the fourth encoding unit into Feature map output by the eighth encoding unit Generate the features of query sequence Q, key sequence K and value sequence V. In order to fully improve the potential of frequency domain features in attention interaction, the first Fourier transform layer is used on its basis to reshape the results Converted into frequency domain features, after two element multipliers, it is converted into spatial domain feature maps through the first inverse Fourier transform layer. The calculation formula is:
[0147]
[0148] Among them, Att Fre represents the result of cross attention calculation, represents the Fourier transform operation, represents the inverse Fourier transform operation;
[0149] In order to avoid excessive noise in the frequency domain information extracted by the cross-attention block, the embodiment of the present invention also adds Fourier transform and inverse Fourier transform to the feedforward neural network, and introduces a learnable parameter matrix to update the weights to adapt to the frequency domain characteristics. The calculation process is as follows:
[0150]
[0151] FFN Fre (f')=Gating(f'W f′ )·(f'W f ′)
[0152] Where f' is the fast Fourier transform f (the above ) results, which are represented by the learnable parameter matrix Finally, the frequency domain fusion feature map F is obtained by adding the frequency domain information and the above cross attention calculation results. fre , whose size is 1×1024×
[0153] Specifically, the adaptive spatial frequency fusion unit is composed of a sequentially connected convolutional layer, an exchange fusion block and a convolutional layer, which is used to fuse the spatial domain feature map Fspa And the frequency domain fusion feature map F fre Fusion is performed to obtain the spatial domain and frequency domain feature fusion result F4, whose size is
[0154] Specifically, the dual-branch decoding module includes:
[0155] A first decoding unit, a second decoding unit, a third decoding unit, a fourth decoding unit, a fifth decoding unit, a sixth decoding unit, a seventh decoding unit, an eighth decoding unit, a height regression unit, and a contour extraction unit;
[0156] The first input end of the first decoding unit and the first input end of the second decoding unit are both connected to the output end of the adaptive spatial frequency fusion unit, and the second input end of the first decoding unit and the second input end of the second decoding unit are both connected to the output end of the third feature fusion unit;
[0157] The first input end of the third decoding unit is connected to the output end of the first decoding unit, the first input end of the fourth decoding unit is connected to the output end of the second decoding unit, and the second input end of the third decoding unit and the second input end of the fourth decoding unit are both connected to the output end of the second feature fusion unit;
[0158] The first input end of the fifth decoding unit is connected to the output end of the third decoding unit, the first input end of the sixth decoding unit is connected to the output end of the fourth decoding unit, and the second input end of the fifth decoding unit and the second input end of the sixth decoding unit are both connected to the output end of the third feature fusion unit;
[0159] The input end of the seventh decoding unit is connected to the output end of the fifth decoding unit, the input end of the eighth decoding unit is connected to the output end of the sixth decoding unit, the output end of the seventh decoding unit is connected to the input end of the height regression unit, the output end of the height regression unit is the first output end of the dual-branch decoding module, the output end of the eighth decoding unit is connected to the input end of the contour extraction unit, and the output end of the contour extraction unit is the second output end of the dual-branch decoding module.
[0160] Specifically, the first decoding unit has the same structure as the second decoding unit, the third decoding unit, the fourth decoding unit, the fifth decoding unit, the sixth decoding unit, the seventh decoding unit, and the eighth decoding unit;
[0161] The first decoding unit includes a bilinear interpolation block, a third convolution block, and a second Mamba block connected in sequence;
[0162] The input end of the bilinear interpolation block is the input end of the first decoding unit, and the output end of the second Mamba block is the output end of the first decoding unit.
[0163] In this embodiment of the present invention, the spatial frequency fusion feature map F4 is input into the first decoding unit and the second decoding unit respectively, and first passes through the bilinear interpolation layer to obtain an upsampled feature map, which is then spliced with the fusion feature F3 to obtain a feature map F'4, whose size is Input the concatenated feature map into the second Mamba block to obtain the feature map F' 4-1 , whose size is
[0164] Since the structure of the first decoding unit in the embodiment of the present invention is the same as that of the second decoding unit, the third decoding unit, the fourth decoding unit, the fifth decoding unit, the sixth decoding unit, the seventh decoding unit, and the eighth decoding unit, the process of processing data by the first decoding unit and the second decoding unit, the third decoding unit, the fourth decoding unit, the fifth decoding unit, the sixth decoding unit, the seventh decoding unit, and the eighth decoding unit is also the same, so in order to shorten the length of the description, the embodiment of the present invention directly gives the input and output data of the third decoding unit, the fourth decoding unit, the fifth decoding unit, the sixth decoding unit, the seventh decoding unit, and the eighth decoding unit, and the characteristic map F' 4-1 The data are input into the third decoding unit and the fourth decoding unit respectively, and after bilinear interpolation, they are spliced with the feature map F2 and then input into the second Mamba block to obtain the feature map F'. 3-1 , whose size is The feature map F' 3-1 Input the fifth decoding unit and the sixth decoding unit, after bilinear interpolation, it is spliced with the feature map F1 and then input into the second Mamba block to obtain the feature map F' 2-1 , whose size is The feature map F' 2-1 The data is input into the seventh decoding unit and the eighth decoding unit, and then input into the second Mamba block after bilinear interpolation to obtain the feature map F'1, whose size is 1×16×w×h.
[0165] In the embodiment of the present invention, the height regression unit and the contour extraction unit respectively use a convolution layer with a convolution kernel of 1 to obtain the height extraction result of the target building and the contour extraction result of the target building. and
[0166] The present invention trains a constructed collaborative extraction network model using an acquired building height profile extraction data set to obtain a trained collaborative extraction network model; inputs SAR image data and optical image data of a target building into the trained collaborative extraction network model for information extraction to obtain a target building height extraction result and a target building profile extraction result; the collaborative extraction network model comprises: a dual-branch encoding module for extracting features from the SAR image data and the optical image data, a spatial-frequency fusion module for fusing dual-domain features, and a dual-branch decoding module for outputting a building height extraction result and a building profile extraction result; compared with the prior art, the present invention introduces feature extraction from the spatial domain and the frequency domain, promotes the full fusion of complementary features of the optical image data and the SAR image data through the spatial-frequency fusion module, and improves the accuracy of collaborative extraction of the building height and profile.
[0167] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for collaboratively extracting building height and outline, characterized in that: include: Step 1: Acquire a building height profile extraction dataset, wherein the building height profile extraction dataset includes vertical polarization band data and vertical-horizontal polarization band data in the SAR image data of the building, four band data in the optical image data of the building, reference height data and reference profile data of the building; Step 2: Using the building height profile extraction dataset to train the constructed collaborative extraction network model to obtain a trained collaborative extraction network model; Step 3: Inputting the SAR image data and optical image data of the target building into the trained collaborative extraction network model to extract information, thereby obtaining a height extraction result and a contour extraction result of the target building; The collaborative extraction network model includes: a dual-branch encoding module for extracting features from SAR image data and optical image data, a spatial frequency fusion module for fusing dual-domain features, and a dual-branch decoding module for outputting building height extraction results and building outline extraction results.
2. The building height and outline collaborative extraction method according to claim 1, characterized in that: The dual-branch encoding module includes: A first coding unit, a second coding unit, a third coding unit, a fourth coding unit, a fifth coding unit, a sixth coding unit, a seventh coding unit, and an eighth coding unit; a first downsampling unit, a second downsampling unit, a third downsampling unit, a fourth downsampling unit, a fifth downsampling unit, and a sixth downsampling unit; A first feature fusion unit, a second feature fusion unit, and a third feature fusion unit; The input end of the first encoding unit and the input end of the fifth encoding unit are both input ends of the dual-branch encoding module; The first output end of the first encoding unit and the first output end of the fifth encoding unit are both connected to the input end of the first feature fusion unit, the output end of the first feature fusion unit is connected to the second input end of the dual-branch decoding module, the second output end of the first encoding unit is connected to the input end of the first downsampling unit, the output end of the first downsampling unit is connected to the input end of the second encoding unit, the second output end of the fifth encoding unit is connected to the input end of the fourth downsampling unit, and the output end of the fourth downsampling unit is connected to the input end of the sixth encoding unit; The first output end of the second encoding unit and the first output end of the sixth encoding unit are both connected to the input end of the second feature fusion unit, the output end of the second feature fusion unit is connected to the third input end of the dual-branch decoding module, the second output end of the second encoding unit is connected to the input end of the second downsampling unit, the output end of the second downsampling unit is connected to the input end of the third encoding unit, the second output end of the sixth encoding unit is connected to the input end of the fifth downsampling unit, and the output end of the fifth downsampling unit is connected to the input end of the seventh encoding unit; The first output end of the third encoding unit and the first output end of the seventh encoding unit are both connected to the input end of the third feature fusion unit, the output end of the third feature fusion unit is connected to the fourth input end of the dual-branch decoding module, the second output end of the third encoding unit is connected to the input end of the third downsampling unit, the output end of the third downsampling unit is connected to the input end of the fourth encoding unit, the second output end of the seventh encoding unit is connected to the input end of the sixth downsampling unit, and the output end of the sixth downsampling unit is connected to the input end of the eighth encoding unit; The output end of the fourth encoding unit and the output end of the eighth encoding unit are both connected to the input end of the spatial frequency fusion module.
3. The method for collaboratively extracting building height and outline according to claim 2, characterized in that: The first encoding unit has the same structure as the fifth encoding unit; The first coding unit includes a first convolution block, a first half-instance normalized residual block, a second half-instance normalized residual block, and a third half-instance normalized residual block connected in sequence; The input end of the first convolution block is the input end of the first coding unit, and the output end of the third half-instance normalized residual block is the first output end and the second output end of the first coding unit.
4. The method for collaboratively extracting building height and outline according to claim 3, characterized in that: The first feature fusion unit has the same structure as the second feature fusion unit and the third feature fusion unit; The first feature fusion unit includes a splicing block, a second convolution block, a batch normalization block, and an activation function block that are connected in sequence.
5. The method for collaboratively extracting building height and outline according to claim 4, characterized in that: The second encoding unit has the same structure as the third encoding unit, the fourth encoding unit, the sixth encoding unit, the seventh encoding unit, and the eighth encoding unit; The second encoding unit includes a patch embedding block, a first Mamba block for feature learning, a patch de-embedding block and a third convolution block connected in sequence; The input end of the patch embedding block is the input end of the second encoding unit, and the third convolution block is the output end of the second encoding unit.
6. The method for collaboratively extracting building height and outline according to claim 5, characterized in that: The spatial frequency fusion module includes: Spatial domain fusion unit, frequency domain fusion unit, adaptive spatial frequency fusion unit; The input end of the spatial domain fusion unit is connected to the output end of the fourth encoding unit and the output end of the eighth encoding unit respectively, and the output end of the spatial domain fusion unit is connected to the input end of the adaptive spatial frequency fusion module; The input end of the frequency domain fusion unit is connected to the output end of the fourth encoding unit and the output end of the eighth encoding unit respectively, and the output end of the frequency domain fusion unit is connected to the input end of the adaptive spatial frequency fusion module; The output end of the adaptive space-frequency fusion unit is connected to the first input end of the dual-branch decoding module.
7. The method for collaboratively extracting building height and outline according to claim 6, characterized in that: The spatial domain fusion unit includes: a first cross-attention block consisting of a first normalization layer, a second normalization layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, a first reshaping layer, a second reshaping layer, a third reshaping layer, a fourth reshaping layer, a first matrix multiplier, a second matrix multiplier, and a first adder; A spatial feedforward network consisting of a third normalization layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a first GELU activation layer, a second adder, and a third adder; An input end of the first normalization layer is connected to an output end of the fourth encoding unit, an output end of the first normalization layer is connected to an input end of the first convolutional layer, an output end of the first convolutional layer is connected to an input end of the fourth convolutional layer, an output end of the fourth convolutional layer is connected to an input end of the first reshaping layer, and an output end of the first reshaping layer is connected to a first input end of the first matrix multiplier; The input end of the second normalization layer is connected to the output end of the eighth encoding unit and the first input end of the first adder respectively, the output end of the second normalization layer is connected to the input end of the second convolutional layer and the input end of the third convolutional layer respectively, the output end of the second convolutional layer is connected to the input end of the fifth convolutional layer, the output end of the fifth convolutional layer is connected to the input end of the second reshaping layer, the output end of the second reshaping layer is connected to the second input end of the first matrix multiplier, and the output end of the first matrix multiplier is connected to the first input end of the second matrix multiplier; The output end of the third convolutional layer is connected to the input end of the sixth convolutional layer, the output end of the sixth convolutional layer is connected to the input end of the third reshaping layer, the output end of the third reshaping layer is connected to the second input end of the second matrix multiplier, the output end of the second matrix multiplier is connected to the input end of the fourth reshaping layer, the output end of the fourth reshaping layer is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the second input end of the first adder, the first output end of the first adder is connected to the first input end of the third adder, and the second output end of the first adder is connected to the input end of the third normalization layer; The output end of the third normalization layer is connected to the input end of the eighth convolutional layer and the input end of the ninth convolutional layer respectively, the output end of the eighth convolutional layer is connected to the input end of the tenth convolutional layer, and the output end of the tenth convolutional layer is connected to the first input end of the second adder; The output end of the ninth convolutional layer is connected to the input end of the eleventh convolutional layer, the output end of the eleventh convolutional layer is connected to the input end of the first GELU activation layer, the output end of the first GELU activation layer is connected to the second input end of the second adder, the output end of the second adder is connected to the input end of the twelfth convolutional layer, the output end of the twelfth convolutional layer is connected to the second input end of the third adder, and the output end of the third adder is connected to the input end of the adaptive spatial frequency fusion unit.
8. The method for collaboratively extracting building height and outline according to claim 7, characterized in that: The frequency domain fusion unit includes: a second cross attention block consisting of a fourth normalization layer, a fifth normalization layer, a sixth normalization layer, a seventh normalization layer, a thirteenth convolution layer, a fourteenth convolution layer, a fifteenth convolution layer, a sixteenth convolution layer, a seventeenth convolution layer, an eighteenth convolution layer, a nineteenth convolution layer, a fifth reshaping layer, a sixth reshaping layer, a seventh reshaping layer, an eighth reshaping layer, a first Fourier transform layer, a second Fourier transform layer, a first inverse Fourier transform layer, a first element multiplier, a second element multiplier, and a fourth adder; a feedforward network consisting of an eighth normalization layer, a ninth reshaping layer, a tenth reshaping layer, a third Fourier transform layer, a second inverse Fourier transform layer, a twentieth convolutional layer, a twenty-first convolutional layer, a twenty-second convolutional layer, a twenty-third convolutional layer, a twenty-fourth convolutional layer, a second GELU activation layer, a third element multiplier, a fourth element multiplier, and a fifth adder; An input end of the fourth normalization layer is connected to an output end of the eighth encoding unit, an output end of the fourth normalization layer is connected to an input end of the thirteenth convolutional layer, an output end of the thirteenth convolutional layer is connected to an input end of the sixteenth convolutional layer, an output end of the sixteenth convolutional layer is connected to an input end of the fifth reshaping layer, an output end of the fifth reshaping layer is connected to an input end of the first Fourier transform layer, and an output end of the first Fourier transform layer is connected to a first input end of the first element-wise multiplier; The input end of the fifth normalization layer is connected to the output end of the fourth encoding unit and the first input end of the fourth adder respectively. The output end of the fifth normalization layer is connected to the input end of the fourteenth convolutional layer and the input end of the fifteenth convolutional layer respectively. The output end of the fourteenth convolutional layer is connected to the input end of the seventeenth convolutional layer. The output end of the seventeenth convolutional layer is connected to the input end of the sixth reshaping layer. The output end of the sixth reshaping layer is connected to the input end of the second Fourier transform layer. The output end of the second Fourier transform layer is connected to the second input end of the first element multiplier. The output end of the first element multiplier is connected to the input end of the first inverse Fourier transform layer. The output end of the first inverse Fourier transform layer is connected to the input end of the eighth reshaping layer. The output end of the eighth reshaping layer is connected to the input end of the sixth normalization layer. The output end of the sixth normalization layer is connected to the first input end of the second element multiplier. The output end of the fifteenth convolutional layer is connected to the input end of the eighteenth convolutional layer, the output end of the eighteenth convolutional layer is connected to the input end of the seventh reshaping layer, the output end of the seventh reshaping layer is connected to the second input end of the second element multiplier, the output end of the second element multiplier is connected to the input end of the nineteenth convolutional layer, and the output end of the nineteenth convolutional layer is connected to the second input end of the fourth adder; The output end of the fourth adder is respectively connected to the input end of the seventh normalization layer and the first input end of the fifth adder, the output end of the seventh normalization layer is connected to the input end of the ninth reshaping layer, the output end of the ninth reshaping layer is connected to the input end of the third Fourier transform layer, the output end of the third Fourier transform layer is connected to the first input end of the third element multiplier, the second input end of the third element multiplier is connected to the learnable parameter matrix, the output end of the third element multiplier is connected to the input end of the second inverse Fourier transform layer, the output end of the second inverse Fourier transform layer is connected to the input end of the tenth reshaping layer, the output end of the tenth reshaping layer is respectively connected to the input end of the 20th convolutional layer and the input end of the 21st convolutional layer, the output end of the 20th convolutional layer is connected to the input end of the 22nd convolutional layer, and the output end of the 22nd convolutional layer is connected to the first input end of the fourth element multiplier; The output end of the twenty-first convolutional layer is connected to the input end of the twenty-third convolutional layer, the output end of the twenty-third convolutional layer is connected to the input end of the second GELU activation layer, the output end of the second GELU activation layer is connected to the second input end of the fourth element multiplier, the output end of the fourth element multiplier is connected to the input end of the twenty-fourth convolutional layer, and the output end of the twenty-fourth convolutional layer is connected to the input end of the adaptive spatial frequency fusion unit.
9. The method for collaboratively extracting building height and outline according to claim 8, characterized in that: The dual-branch decoding module includes: A first decoding unit, a second decoding unit, a third decoding unit, a fourth decoding unit, a fifth decoding unit, a sixth decoding unit, a seventh decoding unit, an eighth decoding unit, a height regression unit, and a contour extraction unit; The first input end of the first decoding unit and the first input end of the second decoding unit are both connected to the output end of the adaptive spatial frequency fusion unit, and the second input end of the first decoding unit and the second input end of the second decoding unit are both connected to the output end of the third feature fusion unit; The first input end of the third decoding unit is connected to the output end of the first decoding unit, the first input end of the fourth decoding unit is connected to the output end of the second decoding unit, and the second input end of the third decoding unit and the second input end of the fourth decoding unit are both connected to the output end of the second feature fusion unit; The first input end of the fifth decoding unit is connected to the output end of the third decoding unit, the first input end of the sixth decoding unit is connected to the output end of the fourth decoding unit, and the second input end of the fifth decoding unit and the second input end of the sixth decoding unit are both connected to the output end of the third feature fusion unit; The input end of the seventh decoding unit is connected to the output end of the fifth decoding unit, the input end of the eighth decoding unit is connected to the output end of the sixth decoding unit, the output end of the seventh decoding unit is connected to the input end of the height regression unit, the output end of the height regression unit is the first output end of the dual-branch decoding module, the output end of the eighth decoding unit is connected to the input end of the contour extraction unit, and the output end of the contour extraction unit is the second output end of the dual-branch decoding module.
10. The method for collaboratively extracting building height and outline according to claim 9, characterized in that: The first decoding unit has the same structure as the second decoding unit, the third decoding unit, the fourth decoding unit, the fifth decoding unit, the sixth decoding unit, the seventh decoding unit, and the eighth decoding unit; The first decoding unit includes a bilinear interpolation block, a third convolution block, and a second Mamba block connected in sequence; The input end of the bilinear interpolation block is the input end of the first decoding unit, and the output end of the second Mamba block is the output end of the first decoding unit.
Citation Information
Patent Citations
Farmland building height prediction method based on remote sensing image
CN116503464A
Multi-stage fusion optical and SAR image combined building extraction method and system
CN117975286A
SAR and optical image fusion and intelligent classification method and system, storage medium and computer program product
CN118887529A
Building height prediction method based on wavelet transform and convolutional neural network
CN119904508A
Cited By
Frequency-space combined image fusion method and device and electronic equipment
CN121353094A