Building three-dimensional information extraction method, device and equipment based on multi-modal data
By fusion of heterologous multimodal data and using multimodal semantic segmentation networks and height inversion networks to predict building plane and height information, the problems of slow speed, inability to monitor large areas and data redundancy in the existing technology are solved, and efficient building three-dimensional information extraction is achieved.
Patent Information
- Application Number
- CN202510571919.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing three-dimensional information extraction methods for building three-dimensional information have problems such as slow speed, inability to monitor large areas, and data redundancy, making it difficult to quickly and effectively extract three-dimensional information of building.
The building three-dimensional information extraction method based on multimodal data is adopted, and the building plane information and height information are predicted respectively by obtaining heterologous multimodal data. The building three-dimensional information is finally generated by a pre-trained multimodal semantic segmentation network and a height inversion network.
It significantly improves the data processing speed, realizes large-area monitoring, reduces data redundancy, improves data utilization, and can efficiently extract three-dimensional information of the building.
Smart Images

Figure CN120088659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing monitoring, and in particular to a method, device and equipment for extracting three-dimensional information of buildings based on multi-modal data. Background Art
[0002] The three-dimensional information of buildings has extensive application value in urban planning, architectural design, military simulation, and map navigation. At present, the methods for extracting the three-dimensional information of buildings are roughly divided into three categories: LiDAR point cloud, oblique photogrammetry, and multi-source data fusion method. Among them, the LiDAR point cloud method obtains high-precision point cloud data through LiDAR technology, and uses the geometric features and spatial distribution of the point cloud to extract the three-dimensional information of buildings. However, this method has a slow processing speed and high cost; the multi-angle oblique photogrammetry method uses drones to obtain stereo image pairs, thereby constructing the three-dimensional information of buildings. However, this method has complex data processing and is easily restricted by airspace, and large-scale monitoring cannot be achieved; the method based on multi-source data fusion comprehensively utilizes the advantages of various data such as LiDAR, satellite images, and drones to achieve the purpose of obtaining the three-dimensional information of buildings. However, this method has the characteristics of diverse data sources, multi-dimensions, and easy redundancy, and cannot quickly and effectively extract information.
[0003] In summary, the above methods have problems such as slow speed, inability to monitor large areas, and data redundancy. Therefore, there is an urgent need for a method that can quickly and widely apply the three-dimensional information of buildings. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, device and equipment for extracting three-dimensional information of buildings based on multi-modal data, which can significantly improve the problems existing in the prior art, such as slow speed, inability to detect large areas, and data redundancy.
[0005] In the first aspect, the present invention provides a method for extracting three-dimensional information of buildings based on multi-modal data, including: Obtaining heterogeneous multi-modal data corresponding to the research area, and performing data fusion processing on the heterogeneous multi-modal data to obtain multi-modal fusion data corresponding to the research area. The heterogeneous multi-modal data includes high-resolution multi-spectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data; Based on the multi-modal fusion data, predicting the building plane information corresponding to the research area through a pre-trained multi-modal semantic segmentation network; Based on the multi-modal fusion data and the building plane information, predicting the building height information corresponding to the research area through a pre-trained height inversion network; Generating the three-dimensional information of the building corresponding to the research area based on the building height information and the building plane information.
[0006] In one embodiment, data fusion processing is performed on heterogeneous multimodal data to obtain multimodal fusion data corresponding to the study area, including: Preprocessing, geometric registration processing, and resampling processing are performed on high-resolution multispectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data to obtain processed land surface reflectance, SAR backscattering coefficient data, digital surface model data, digital elevation model data, and land surface temperature data; Band combination processing, normalization processing, and principal component analysis processing are performed on the processed land surface reflectance, the SAR backscattering coefficient data, and the land surface temperature data to obtain multimodal fusion data.
[0007] In one embodiment, the multimodal semantic segmentation network includes a chunking layer, a plurality of Mix Transform encoding layers, a feature extraction layer, and a decoding layer; based on the multimodal fusion data, the multimodal semantic segmentation network obtained by pre-training predicts the building plane information corresponding to the study area, including: Through the chunking layer, the multimodal fusion data is segmented into a plurality of multimodal fusion sub-data, and the multimodal fusion sub-data is converted into one-dimensional feature vectors; Through a plurality of Mix Transform encoding layers, hierarchical feature vectors corresponding to the one-dimensional feature vectors are extracted; Through the feature extraction layer, downsampling feature extraction processing and upsampling feature fusion processing are performed on the hierarchical feature vectors to obtain target feature vectors; Through the decoding layer, the target feature vectors are decoded to predict the building plane information corresponding to the study area.
[0008] In one embodiment, based on the multimodal fusion data and the building plane information, a height inversion network obtained by pre-training predicts the building height information corresponding to the study area, including: Using the building plane information as a mask, the multimodal fusion data corresponding to each building unit within the study area is extracted from the multimodal fusion data; For any building unit, band mean processing is performed on the multimodal fusion data corresponding to the building unit to obtain the multimodal value corresponding to the building unit; Through the height inversion network obtained by pre-training, based on the multimodal value corresponding to each building unit, the building height information corresponding to each building unit within the study area is predicted.
[0009] In one embodiment, based on the building height information and the building plane information, building three-dimensional information corresponding to the study area is generated, including: For any building unit within the research area, write the building height information corresponding to the building unit into the building plane information to generate the building three-dimensional information corresponding to each building unit within the research area.
[0010] In one implementation, the method further includes: Construct a multi-modal semantic segmentation sample set based on the multi-modal fusion data and building plane samples corresponding to a specified sub-region within the research area, and use the multi-modal semantic segmentation sample set to train the multi-modal semantic segmentation network; And, construct a multi-modal height inversion data set based on the digital surface model data, digital elevation model data, multi-modal fusion data, and building plane samples corresponding to a specified sub-region within the research area, and use the multi-modal height inversion data set to train the height inversion network.
[0011] In one implementation, constructing a multi-modal height inversion data set based on the digital surface model data, digital elevation model data, multi-modal fusion data, and building plane samples corresponding to a specified sub-region within the research area includes: Based on the digital surface model data and digital elevation model data corresponding to the specified sub-region, determine the relative height information corresponding to the specified sub-region; Using the building plane samples as a mask, respectively extract the multi-modal values and building height values corresponding to each building unit within the specified sub-region from the multi-modal fusion data and relative height information corresponding to the specified sub-region; Use the multi-modal values as the model input and the building height values as the labels to obtain the multi-modal height inversion data set.
[0012] In a second aspect, the present invention also provides a device for extracting building three-dimensional information based on multi-modal data, including: A data processing module, configured to obtain heterogeneous multi-modal data corresponding to the research area, and perform data fusion processing on the heterogeneous multi-modal data to obtain multi-modal fusion data corresponding to the research area, where the heterogeneous multi-modal data includes high-resolution multi-spectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data; A plane information prediction module, configured to predict the building plane information corresponding to the research area based on the multi-modal fusion data through a pre-trained multi-modal semantic segmentation network; A height information prediction module, configured to predict the building height information corresponding to the research area based on the multi-modal fusion data and the building plane information through a pre-trained height inversion network; A three-dimensional information generation module, configured to generate the building three-dimensional information corresponding to the research area based on the building height information and the building plane information.
[0013] In a third aspect, the present invention further provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of the first aspect.
[0014] In a fourth aspect, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method according to any one of the first aspect.
[0015] A method, apparatus, and device for extracting three-dimensional information of a building based on multi-modal data provided by an embodiment of the present invention first obtain heterogeneous multi-modal data corresponding to a research area, and perform data fusion processing on the heterogeneous multi-modal data to obtain multi-modal fusion data corresponding to the research area. The heterogeneous multi-modal data includes high-resolution multi-spectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data. Then, through a pre-trained multi-modal semantic segmentation network, the building plane information corresponding to the research area is predicted based on the multi-modal fusion data. Next, through a pre-trained height inversion network, the building height information corresponding to the research area is predicted based on the multi-modal fusion data and the building plane information. Finally, based on the building height information and the building plane information, the three-dimensional information of the building corresponding to the research area is generated. The above method performs data fusion on the heterogeneous multi-modal data, effectively improving the data validity and solving the problem of data redundancy. In addition, a multi-modal semantic segmentation model and a height inversion model applicable to the multi-modal fusion data are constructed to further improve the data utilization rate to obtain the building plane information and the building height information respectively, and the efficient three-dimensional information of the building can be realized even in the scenario of large-area detection.
[0016] Other features and advantages of the present invention will be described in the following description, and some of them will be obvious from the description or learned through the implementation of the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description, claims, and drawings.
[0017] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 Schematic flow chart of a method for extracting three-dimensional information of buildings based on multi-modal data provided by an embodiment of the present invention; Figure 2 Technical framework diagram of a method for extracting three-dimensional information of buildings based on multi-modal data provided by an embodiment of the present invention; Figure 3 Schematic flow chart of a method for processing heterogeneous multi-modal data provided by an embodiment of the present invention; Figure 4 Schematic flow chart of a method for multi-modal data fusion provided by an embodiment of the present invention; Figure 5 Schematic flow chart of a method for extracting planar features provided by an embodiment of the present invention; Figure 6 Schematic structural diagram of a multi-modal semantic segmentation network provided by an embodiment of the present invention; Figure 7 Schematic flow chart of a method for inverting height information provided by an embodiment of the present invention; Figure 8 Overall structural diagram of an inversion deep learning network constructed by using a Transformer provided by an embodiment of the present invention; Figure 9 Schematic structural diagram of a device for extracting three-dimensional information of buildings based on multi-modal data provided by an embodiment of the present invention; Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0021] At present, the existing technologies have problems such as slow speed, inability to monitor large areas, and data redundancy. Therefore, there is an urgent need for a method that can quickly and widely apply the three-dimensional information of building objects. With the development of the Internet of Things and multi-sensor technologies, the multi-source data fusion method is one of the most promising methods. To solve the difficulty of multi-source data fusion and improve the extraction performance at the same time, a method for extracting the three-dimensional information of buildings based on heterogeneous multi-modal data is designed. This method can stably and efficiently extract the three-dimensional information of large-scale buildings.
[0022] Based on this, the embodiments of the present invention provide a method, device, and equipment for extracting the three-dimensional information of buildings based on multi-modal data, which can significantly improve the problems existing in the existing technologies, such as slow speed, inability to detect large areas, and data redundancy, and effectively improve the problem of difficulty in multi-source data fusion existing in the existing technologies.
[0023] To facilitate the understanding of this embodiment, first, a method for extracting the three-dimensional information of buildings based on multi-modal data disclosed in the embodiments of the present invention will be introduced in detail. Refer to Figure 1 the flow schematic diagram of a method for extracting the three-dimensional information of buildings based on multi-modal data shown in the figure. This method mainly includes the following steps S102 to S108: Step S102, obtain heterogeneous multi-modal data corresponding to the research area, and perform data fusion processing on the heterogeneous multi-modal data to obtain multi-modal fusion data corresponding to the research area.
[0024] Among them, the heterogeneous multi-modal data includes high-resolution multi-spectral data, synthetic aperture radar data (Synthetic Aperture Radar, SAR), digital surface model data (Digital Surface Model, DSM), digital elevation model data (Digital Elevation Model, DEM), and land surface temperature data. The multi-modal fusion data is obtained based on the processed surface reflectance, SAR backscatter coefficient data, and land surface temperature data corresponding to the research area.
[0025] In one example, preprocess the high-resolution multi-spectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data respectively, perform geometric correction processing and resampling processing on each of the preprocessed data, and then the processed data corresponding to the research area can be obtained. On this basis, the multi-modal fusion data is determined through band combination processing, normalization processing, and principal component analysis processing.
[0026] Step S104, based on the multi-modal fusion data, predict the building plane information corresponding to the research area through a pre-trained multi-modal semantic segmentation network.
[0027] Among them, the input of the multi-modal semantic segmentation network is multi-modal fusion data, and the output is the building plane information corresponding to the research area. This building plane information describes the distribution of each building unit included in the research area. In one example, by inputting the multi-modal fusion data into the multi-modal semantic segmentation network, the building plane information output by the multi-modal semantic segmentation network can be obtained.
[0028] Step S106, based on the pre-trained building height inversion network, predict the building height information corresponding to the research area based on the multi-modal fusion data and the building plane information.
[0029] Among them, the input of the building height inversion network is the multi-modal value corresponding to each building unit in the research area, and the output is the building height information corresponding to each building unit in the research area. In one example, the building plane information can be used as a mask to extract the multi-modal fusion data corresponding to each building unit, and then determine the multi-modal value corresponding to each building unit. By inputting it into the building height inversion network, the building height information output by the building height inversion network can be obtained.
[0030] Step S108, generate the building three-dimensional information corresponding to the research area based on the building height information and the building plane information.
[0031] In one example, the building height information can be written into the building plane information to obtain the building three-dimensional information corresponding to each building unit in the research area. The building three-dimensional information is used to describe the plane coordinates of the building unit (i.e., the building plane information) and the building height information.
[0032] The method for extracting building three-dimensional information based on multi-modal data provided by the embodiments of the present invention performs data fusion on heterogeneous multi-modal data, effectively improving the data validity and solving the problem of data redundancy. In addition, by constructing a multi-modal semantic segmentation model and a building height inversion model applicable to multi-modal fusion data, the data utilization rate is further improved to respectively obtain the building plane information and the building height information. Even in the scenario of large-area detection, efficient building three-dimensional information can be realized.
[0033] Furthermore, the embodiments of the present invention also provide an implementation manner for training the multi-modal semantic segmentation network and the building height inversion network, including: (1) Construct a multi-modal semantic segmentation sample set based on the multi-modal fusion data and the building plane samples corresponding to the specified sub-region in the research area, and use the multi-modal semantic segmentation sample set to train the multi-modal semantic segmentation network.
[0034] In one example, using the multi-modal fusion data corresponding to the specified sub-region and the building plane samples, combined with sample enhancement techniques such as inversion and rotation, a multi-modal semantic segmentation sample set of size (s, m, m, 3) is constructed, where s is the number of samples. The multi-modal semantic segmentation sample set is trained, validated, and tested in a ratio of 6, 2, 2. In addition, key hyperparameters of the multi-modal semantic segmentation network, such as training batch size, learning rate, c1, and other key parameters, are set, and the optimal multi-modal semantic segmentation network is obtained through training.
[0035] (2) Construct a multi-modal height inversion data set based on the digital surface model data, digital elevation model data, multi-modal fusion data, and building plane samples corresponding to the specified sub-region within the research area, and use the multi-modal height inversion data set to train the height inversion network.
[0036] In one example, the process of constructing the multi-modal height inversion data set is as follows: (2.1) Based on the digital surface model data and digital elevation model data corresponding to the specified sub-region, determine the relative height information corresponding to the specified sub-region. The relative elevation data is obtained based on the processed digital surface model data and digital elevation model data corresponding to the specified sub-region within the research area.
[0037] Specifically, use the digital surface model data corresponding to the specified sub-region and the digital elevation model data , and calculate the relative height data H: ; (2.2) Using the building plane sample as a mask, extract the multi-modal value and building height value corresponding to each building unit within the specified sub-region from the multi-modal fusion data and relative height information corresponding to the specified sub-region. Specifically, according to the spatial information of the building plane sample, overlay the multi-modal fusion data and relative height information to obtain the multi-modal fusion data and relative height information within each building unit, and calculate the mean value of each band of the multi-modal fusion data and relative height information within each building unit to obtain the multi-modal value and building height value of each building unit.
[0038] (2.3) Use the multi-modal value as the model input and the building height value as the label to obtain the multi-modal height inversion data set. Specifically, process all building units one by one according to (2.2), and the multi-modal height inversion data set can be obtained, where the size of the multi-modal height inversion data set is (3, s), and the size of the height data is (1, s), and s is the number of samples.
[0039] The multi-modal height inversion dataset is trained, validated, and tested in a ratio of 6, 2, 2. Key hyperparameters of the inversion deep learning network are set, such as training batch, learning rate, d, and other key parameters, and the optimal model is obtained through training.
[0040] On this basis, an embodiment of the present invention provides a technical framework diagram of a method for extracting building three-dimensional information based on multi-modal data as shown in Figure 2 Figure 5. It includes: First, process data such as high-resolution multi-spectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data to generate heterogeneous multi-modal data to be fused; then use techniques such as standardization and principal component analysis to generate multi-modal fusion data, reducing the noise of heterogeneous multi-modal data and improving the usability of the data; use existing building plane samples and multi-modal fusion data to generate a multi-modal semantic segmentation sample set, and construct a semantic segmentation network suitable for multi-modal fusion data. Through model training and validation, the optimal multi-modal semantic segmentation network is obtained, and the building plane information in a larger range is obtained by predicting the multi-modal fusion data in a larger range; calculate the relative height of the digital surface model data and the digital elevation model data to generate relative height data, and superimpose the multi-modal fusion data and the building plane samples to generate a multi-modal height inversion dataset. At the same time, a height inversion network suitable for multi-modal data is constructed. Through model training and validation, the optimal height inversion model is obtained; use the predicted building plane information data, superimpose the multi-modal fusion data in a larger range, and predict through the height inversion model to obtain the building height information; finally, write the building height information into the building plane information to obtain the building three-dimensional information.
[0041] In Figure 2 On the basis of this, an embodiment of the present invention provides a specific implementation manner of a method for extracting building three-dimensional information based on multi-modal data.
[0042] (1) Heterogeneous multi-modal data processing: Preprocess, geometrically register, and resample high-resolution multi-spectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data to obtain processed surface reflectivity, SAR backscattering coefficient data, digital surface model data, digital elevation model data, and land surface temperature data.
[0043] In specific implementation, refer to Figure 3A schematic flowchart of heterologous multimodal data processing is shown, including: First, perform data preprocessing on the original high-resolution multispectral data and synthetic aperture radar (SAR) data. Among them, the high-resolution multispectral data undergoes preprocessing such as orthorectification, atmospheric correction, and fusion to generate surface reflectance data; the SAR data undergoes processing such as radiometric correction, filtering, and terrain correction to generate SAR backscatter coefficient data; perform filtering processing on the DSM to eliminate the influence of fine noise; Since the spatial references and spatial positions of heterologous multimodal data are not unified, taking the optical surface reflectance data as the reference, perform geometric registration on the SAR backscatter coefficient data, DSM, and surface temperature data. And since the optical data is orthorectified based on the digital elevation model (DEM) data, the geometric positions of the DEM data and the surface reflectance data are consistent. After geometric registration, the spatial positions and spatial references of all heterologous multimodal data are consistent; Then use resampling technology to process other data into the same spatial resolution as the surface reflectance data. Denote the surface reflectance data, SAR backscatter coefficient data, and surface temperature data at this time as SR, BC, and BT; Since the DSM data is usually data obtained by drones in a small area, while other data is basically obtained by satellites, and the acquisition range is much larger than the DSM range, it is necessary to perform an intersection process on the data to obtain the overlapping area of all data. Denote the surface reflectance data, SAR backscatter coefficient, DEM, DSM, and surface temperature data at this time as M1, M2, M3, M4, M5. After these data, the number of bands is N1, N2, N3, N4, N5, and the length and width are both (H, W).
[0044] It should be noted that during the training stage of the network, it is necessary to perform an intersection process on the data to obtain the overlapping area of all data, so as to construct a multimodal semantic segmentation sample set and a multimodal height inversion data set. mainly use the DSM to determine the building height value (as a label) in the multimodal height inversion data set; During the application stage of the network, there is no need to perform an intersection process on the data. Only need to obtain the surface reflectance data, SAR backscatter coefficient, and surface temperature data of the entire study area through preprocessing, geometric calibration processing, and resampling processing, which is convenient for subsequent determination of multimodal fusion data and serves as the input of the multimodal semantic segmentation network and height inversion network.
[0045] (2) Multimodal data fusion: Perform band combination processing, normalization processing, and principal component analysis processing on the processed surface reflectance, SAR backscatter coefficient data, and surface temperature data to obtain multimodal fusion data.
[0046] Due to the characteristics of heterogeneous multimodal data such as diversity, multi-dimensions, and easy redundancy, and at the same time, the radiation value ranges and physical meaning representations of optical, SAR, and temperature data are inconsistent, and various types of noise may appear between the data, so it is necessary to perform standardization processing, PCA (principal component analysis), etc. to achieve the purpose of multimodal data fusion and effective information extraction.
[0047] See Figure 4 A schematic flow diagram of a multimodal data fusion shown below, including: First, perform band combination on the preprocessed optical, SAR, and temperature data. Since semantic segmentation deep learning networks usually face natural pictures and the data needs to be quantized to 0 - 255, and multimodal data includes not only image data but also non-standardized data such as SAR and temperature with inconsistent value ranges, so standardization processing is performed on the combined data; then use PCA to extract the main effective information from the combined data, reduce the noise interference of different data sources, and generate multimodal fusion data. The specific process is as follows: First, combine the M1, M2, and M5 bands to generate the F1 multi-dimensional array with a size of (N, H, W), then transform the three-dimensional array into the F2 two-dimensional array with a size of (N, H*W), where N = N1 + N2 + N5, and then perform standardization processing to generate F3: ; ; ; where i is the i-th band of the mean value of the i-th band, the standard deviation of the i-th band, the standardized data of the i-th band.
[0048] Then use PCA to reduce the dimension of F3 as follows: ; ; where is the mean value of the i-th dimensional array, and C is the covariance of
[0049] Perform eigenvalue decomposition on the covariance C to obtain eigenvalues and eigenvectors: ; where is the eigenvector with a size of (H*W, 1), is the eigenvalue.
[0050] Select the eigenvectors v1, v2, and v3 corresponding to the largest three eigenvalues, and finally generate the PCA dimensionality reduction array P, with a size of (3, H*W): ; Perform an array size transformation on P to generate the fusion data F, with a size of (3, H, W).
[0051] (III) Plane feature extraction: Refer to Figure 5 the schematic flow diagram of a plane feature extraction method shown in the figure, including: First, construct a deep learning network suitable for multi-modal fusion data, combine existing building plane samples to construct a multi-modal semantic segmentation sample set, then perform model training, and use the verified multi-modal semantic segmentation network to predict the multi-modal fusion data, and finally generate large-scale building plane information. Specifically include: Multi-modal semantic segmentation network: Refer to Figure 6 the schematic structural diagram of a multi-modal semantic segmentation network shown in the figure, including a block layer, multiple Mix Transform encoding layers, a feature extraction layer, and a decoding layer. In the embodiments of the present invention, a Mix Transformer suitable for multi-modal data is used as the basic backbone to construct a building semantic segmentation network. The core part of this network is the Mix Transformer Block, which combines the advantages of a convolutional neural network (CNN) and a Transformer, can effectively extract hierarchical features of multi-modal data, and can effectively obtain building plane information.
[0052] Based on this, the process of using this network to predict building plane information is as follows: Through the block layer, the multi-modal fusion data is divided into multiple multi-modal fusion sub-data of a fixed size, and the multi-modal fusion sub-data is converted into one-dimensional feature vectors; through multiple Mix Transform encoding layers, hierarchical feature vectors corresponding to the one-dimensional feature vectors are effectively extracted; through the feature extraction layer, downsampling feature extraction processing and upsampling feature fusion processing are performed on the hierarchical feature vectors to obtain target feature vectors, which are used to gradually reduce the spatial resolution of the feature map, increase the abstraction level of the features, and fuse features at different levels to enhance the model's perception ability of global and local information; through the decoding layer, the feature map is restored to the original input resolution to decode the target feature vector for predicting the building plane information corresponding to the research area.
[0053] In specific implementation, refer to the structural information of the multi-modal semantic segmentation network shown in Table 1: Table 1 Structural information of the multi-modal semantic segmentation network
[0054] Based on this, the specific process is as follows: First, set the size of the multi-modal sample data to (m, m, 3) and input it into the chunking layer. The sample data is chunked into sub-data blocks of (4, 4). After each sub-database block undergoes linear projection, it becomes a one-dimensional feature vector, and the output data size is (m / 4, m / 4, c1). Then the data enters the encoding layer with Mix Transformer as the basic network. The encoding layer contains 3 Mix Transformer Blocks. Each Mix Transformer Block includes a normalization layer, a Multi-Head Self-Attention (MHSA) layer, a Residual Connection layer, a LayerNormalization layer, a Mixer layer, and a Residual Connection layer. The output data size is (m / 4, m / 4, c1). The data passing through the encoding layer flows into the feature extraction layer, which includes a downsampling feature extraction module and an upsampling feature fusion module. The output data size of this layer is (m / 4, m / 4, c1). Then the data flows into the decoding layer, which includes an upsampling module, a convolutional layer, and an activation function. The output data size is (m, m, 1). The loss function uses Cross-Entropy Loss.
[0055] The training process of the multi-modal semantic segmentation network can be referred to the foregoing (1), and this embodiment of the present invention will not elaborate on it. After the training of the multi-modal semantic segmentation network is completed, perform multi-modal data fusion on SR, BC, and BT to generate a larger range of multi-modal fusion data, and use the optimal multi-modal semantic segmentation network to predict the large-range multi-modal fusion data to obtain the building plane information D1.
[0056] (4) Height information inversion: Refer to Figure 7 As shown in the flow schematic diagram of a height information inversion, it includes: First, use M3 and M4 data to calculate the relative height data. By superimposing the relative height data and the multi-modal fusion data on the building plane samples, generate a multi-modal height inversion data set, construct a height inversion network suitable for multi-modal data, and perform model training and verification to obtain the optimal height inversion network. Then, by superimposing the building plane information and the multi-modal fusion data, generate the data to be inverted for large-range inversion of building height information. Then use the optimal height inversion data to predict the data to be inverted, and finally obtain the large-range building height information.
[0057] Height inversion network: General height inversion networks include random forests, support vector machines, recurrent neural networks, etc. However, the performance of these networks often fails to meet the requirements when adapting to multimodal data. Deep learning networks based on Transformer can be applied to complex multimodal data, such as Figure 8 As shown in , a general structure of an inversion deep learning network constructed using Transformer includes an input layer, an embedding layer, a positional encoding layer, a Transformer encoder, a global pooling layer, and an output layer.
[0058] The data to be inverted is sent to the embedding layer after the input layer, mapped to a high-dimensional space, and then the positional encoding layer is used to assign positional information to the features so that they can be effectively utilized by the Transformer encoder; the Transformer encoder learns the mutual relationships between data features through the multi-head self-attention mechanism and the feed-forward neural network, and extracts key information; finally, the global pooling layer integrates the encoded information, and the output layer completes the linear transformation to generate building height information.
[0059] In the specific implementation, refer to the structure information of the height inversion network shown in Table 2: Table 2 Structure information of the height inversion network
[0060] Based on this, the specific process is as follows: Let the training sample data be X(3, s), where s is the number of samples and is the number of fused data features; first, the embedding layer is used to map the training composed of 3 features to a high-dimensional space to better capture complex feature relationships, and the output data size becomes (1, s, d), where d is the hidden dimension; then the positional encoding layer adds positional information to the data, and the positional encoding helps the algorithm identify the positions of different features, and the output positional encoding data size is (1, s, d); the Transformer encoder is the core part of the network, which consists of multiple layers of Transformer encoders. Through the self-attention mechanism and the feed-forward neural network, it learns the mutual dependencies between features and deeply encodes the input feature sequence. The output encoded data retains important feature information, and the output data size is (1, s, d); the global pooling layer performs a pooling operation on the encoded feature sequence to extract the key features of each sample, and the output data size is (s, d); finally, these features are integrated by the output layer and passed through a linear transformation to obtain the final regression prediction value, and the output size is (1, s). This network uses the mean square error (MSE) as the loss function and the Adam optimizer to train the model.
[0061] The training process of the height inversion network can be referred to the aforementioned (2), and the embodiments of the present invention will not elaborate on this. After the training of the height inversion network is completed, the building height information can be predicted according to the following process: using the building plane information as a mask, extracting the multi-modal fusion data corresponding to each building unit in the research area from the multi-modal fusion data; for any building unit, performing mean processing by band on the multi-modal fusion data corresponding to the building unit to obtain the multi-modal value corresponding to the building unit; through the pre-trained height inversion network, based on the multi-modal value corresponding to each building unit, predicting the building height information corresponding to each building unit in the research area. In specific implementation, perform multi-modal data fusion on SR, BC, and BT to generate a larger range of multi-modal fusion data, and then overlay the building plane information D1 to obtain the mean value of the multi-modal values in each D1 building feature. Construct a large range of multi-modal data to be inverted based on the multi-modal mean, and use the optimal height inversion network to predict the large range of multi-modal fusion data to be inverted to obtain the large range of building height information D2.
[0062] (5) Stereo information acquisition: For any building unit in the research area, write the building height information corresponding to the building unit into the building plane information to generate the building stereo information corresponding to each building unit in the research area. In specific implementation, through the integration of D1 and D2, write the height information of D2 into the D1 data to obtain the building stereo information D.
[0063] Furthermore, the embodiments of the present invention also provide an application example for extracting building stereo information based on multi-modal data, including: a. Heterogeneous multi-modal data processing: The research area in the embodiments of the present invention is Hebi City, Henan Province. The heterogeneous data includes multi-spectral images, SAR data, temperature products, DSM, and DEM. Among them, the multi-spectral image is collected by the GF1 PMS sensor, and after preprocessing, it is surface reflectance data with a resolution of 2m and 4 bands; the SAR is collected by Sentinel1, and after preprocessing, the VV and VH backscattering coefficients are obtained with a resolution of 20m and 2 bands; the temperature product is obtained by the Landsat9 TIRS sensor with a resolution of 30m; the DSM is obtained by a drone with a resolution of 0.2m and 1 band; the DEM has a resolution of 12m and 1 band; except for the DEM data, the collection time of other data is May 2024. Taking the multi-spectral surface reflectance data and DEM as the benchmarks, perform registration on SAR, temperature, and DSM respectively, and then sample all the data to a resolution of 2m to generate SR, BC, and BT; then perform intersection extraction on all the data to obtain M1, M2, M3, M4, and M5, with the number of bands being 4, 2, 1, 1, and 1 respectively, and the size being (8671, 10829).
[0064] (b)Multi-modal data fusion: First, perform band combination on the preprocessed data of M1, M2, and M5, then perform standardization processing on the combined data, and then use PCA to extract effective information from the standardized data to generate multi-modal fusion data F with a size of (3, 8671, 10829).
[0065] (c)Plane information extraction: Set the sample size m in the multi-modal semantic segmentation network to 512 and c1 to 128. Use F and building plane samples to construct 2150 samples, among which 1290 are for training, 428 are for validation, and 428 are for testing. The training batch is 200, and the learning rate is 0.1. Finally, generate the optimal model and predict the multi-modal fusion data in a larger range to obtain the building plane data D1.
[0066] (d)Height information inversion: Set the training batch, learning rate, and the number of hidden network layers d in the height inversion network to 100, 0.1, and 64 respectively. The sample size m is 512 and c1 is 128. Use the multi-modal fusion data F, building plane samples, and building height data to construct 2150 inversion samples, among which 1290 are for training, 428 are for validation, and 428 are for testing. Through training and validation, finally generate the optimal inversion model and predict the multi-modal fusion data in a larger range to obtain the building height data D2.
[0067] (e)Stereo information acquisition: Integrate D1 and D2 to obtain the stereo information corresponding to each building unit in Hebi City, Henan Province.
[0068] In summary, the method for extracting building stereo information based on multi-modal data provided by the embodiments of the present invention uses techniques such as principal component analysis and standardization to improve the effectiveness of data, realizes the fusion of heterogeneous multi-modal data, and can further improve the data utilization efficiency by constructing a semantic segmentation network and an inversion network suitable for multi-modal fusion data, and respectively obtain the building plane and height information. This method is applicable to various heterogeneous data such as optical, SAR, temperature, DSM, and DEM, and has the advantages of strong robustness, high efficiency, and convenience for large-scale monitoring, playing a basic technical support role in urban monitoring and supervision.
[0069] Based on the foregoing embodiments, the embodiments of the present invention provide a device for extracting building stereo information based on multi-modal data. Refer to Figure 9 the structural schematic diagram of a device for extracting building stereo information based on multi-modal data shown. The device mainly includes the following parts: The data processing module 902 is configured to obtain heterogeneous multi-modal data corresponding to the research area, and perform data fusion processing on the heterogeneous multi-modal data to obtain multi-modal fusion data corresponding to the research area. The heterogeneous multi-modal data includes high-resolution multi-spectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and land surface temperature data. The plane information prediction module 904 is configured to predict the building plane information corresponding to the research area based on the multi-modal fusion data through a pre-trained multi-modal semantic segmentation network. The height information prediction module 906 is configured to predict the building height information corresponding to the research area based on the multi-modal fusion data and the building plane information through a pre-trained height inversion network. The three-dimensional information generation module 908 is configured to generate the building three-dimensional information corresponding to the research area based on the building height information and the building plane information.
[0070] The building three-dimensional information extraction device based on multi-modal data provided by the embodiments of the present invention performs data fusion on heterogeneous multi-modal data, effectively improving the data validity and solving the problem of data redundancy. In addition, by constructing a multi-modal semantic segmentation model and a height inversion model applicable to multi-modal fusion data, the data utilization rate is further improved to respectively obtain the building plane information and the building height information. Even in the scenario of large-area detection, efficient building three-dimensional information can be achieved.
[0071] For the device provided by the embodiments of the present invention, the implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For a brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding contents in the foregoing method embodiments.
[0072] The embodiments of the present invention provide an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and the computer program executes the method according to any one of the foregoing implementation manners when being run by the processor.
[0073] Figure 10 FIG. is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device 100 includes: a processor 10, a memory 11, a bus 12, and a communication interface 13. The processor 10, the communication interface 13, and the memory 11 are connected through the bus 12; the processor 10 is configured to execute an executable module stored in the memory 11, such as a computer program.
[0074] Among them, the memory 11 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is implemented through at least one communication interface 13 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0075] The bus 12 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of easy representation, Figure 10 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0076] Among them, the memory 11 is used to store a program. After receiving an execution instruction, the processor 10 executes the program. The method executed by the device defined by the flow process disclosed in any embodiment of the foregoing embodiments of the present invention can be applied to or implemented by the processor 10.
[0077] The processor 10 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 10 or by instructions in software form. The above-mentioned processor 10 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 11, and the processor 10 reads the information in the memory 11 and combines its hardware to complete the steps of the above method.
[0078] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the foregoing method embodiments, which will not be elaborated here.
[0079] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other various media that can store program code.
[0080] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for extracting three-dimensional information of a building based on multimodal data, characterized in that: include: Acquire heterogeneous multimodal data corresponding to the study area, and perform data fusion processing on the heterogeneous multimodal data to obtain multimodal fusion data corresponding to the study area, wherein the heterogeneous multimodal data includes high-resolution multispectral data, synthetic aperture radar data, digital surface model data, digital elevation model data, and surface temperature data; Predicting the building plan information corresponding to the study area based on the multimodal fusion data using a pre-trained multimodal semantic segmentation network; Predicting building height information corresponding to the study area based on the multimodal fusion data and the building plane information through a pre-trained height inversion network; Based on the building height information and the building plane information, three-dimensional building information corresponding to the study area is generated.
2. The method for extracting three-dimensional building information based on multimodal data according to claim 1, characterized in that: Performing data fusion processing on the heterogeneous multimodal data to obtain multimodal fusion data corresponding to the research area includes: Preprocessing, geometric registration and resampling the high-resolution multispectral data, the synthetic aperture radar data, the digital surface model data, the digital elevation model data and the land surface temperature data to obtain processed land surface reflectivity, SAR backscatter coefficient data, the digital surface model data, the digital elevation model data and the land surface temperature data; The processed surface reflectivity, the SAR backscatter coefficient data and the surface temperature data are subjected to band combination processing, standardization processing and principal component analysis processing to obtain multimodal fusion data.
3. The method for extracting three-dimensional building information based on multimodal data according to claim 1, characterized in that: The multimodal semantic segmentation network includes a block layer, multiple Mix Transform encoding layers, a feature extraction layer and a decoding layer; the multimodal semantic segmentation network obtained by pre-training predicts the building plane information corresponding to the study area based on the multimodal fusion data, including: By means of the block layer, the multimodal fusion data is divided into a plurality of multimodal fusion sub-data, and the multimodal fusion sub-data is converted into a one-dimensional feature vector; Extracting a hierarchical feature vector corresponding to the one-dimensional feature vector through a plurality of Mix Transform coding layers; Through the feature extraction layer, the hierarchical feature vector is subjected to down-sampling feature extraction processing and up-sampling feature fusion processing to obtain a target feature vector; The target feature vector is decoded through the decoding layer to predict the building plane information corresponding to the study area.
4. The method for extracting three-dimensional building information based on multimodal data according to claim 1, characterized in that: Predicting the building height information corresponding to the study area based on the multimodal fusion data and the building plane information through the pre-trained height inversion network, including: Using the building plane information as a mask, extracting the multimodal fusion data corresponding to each building unit in the study area from the multimodal fusion data; For any of the building units, the multimodal fusion data corresponding to the building unit is processed according to the band mean value to obtain the multimodal value corresponding to the building unit; The building height information corresponding to each building unit in the study area is predicted based on the multimodal value corresponding to each building unit through the pre-trained height inversion network.
5. The method for extracting three-dimensional building information based on multimodal data according to claim 1, characterized in that: Generating three-dimensional building information corresponding to the study area based on the building height information and the building plane information, including: For any building unit in the study area, the building height information corresponding to the building unit is written into the building plane information to generate the building three-dimensional information corresponding to each building unit in the study area.
6. The method for extracting three-dimensional building information based on multimodal data according to claim 1, characterized in that: The method further comprises: Constructing a multimodal semantic segmentation sample set based on the multimodal fusion data and building plane samples corresponding to the designated sub-area within the study area, and training the multimodal semantic segmentation network using the multimodal semantic segmentation sample set; Furthermore, a multimodal height inversion dataset is constructed based on the digital surface model data, the digital elevation model data, the multimodal fusion data and the building plane samples corresponding to the designated sub-area in the study area, and the height inversion network is trained using the multimodal height inversion dataset.
7. The method for extracting three-dimensional building information based on multimodal data according to claim 6, characterized in that: Constructing a multimodal height inversion dataset based on the digital surface model data, the digital elevation model data, the multimodal fusion data and the building plane sample corresponding to the designated sub-area in the study area, including: Determining relative height information corresponding to the designated sub-area based on the digital surface model data and the digital elevation model data corresponding to the designated sub-area; Using the building plane sample as a mask, respectively extracting the multimodal value and the building height value corresponding to each building unit in the designated sub-area from the multimodal fusion data and the relative height information corresponding to the designated sub-area; The multimodal values are used as model inputs, and the building height values are used as labels to obtain a multimodal height inversion dataset.
8. A device for extracting three-dimensional information of a building based on multimodal data, characterized in that: include: A data processing module, used for acquiring heterogeneous multimodal data corresponding to a study area, and performing data fusion processing on the heterogeneous multimodal data to obtain multimodal fusion data corresponding to the study area, wherein the heterogeneous multimodal data includes high-resolution multispectral data, synthetic aperture radar data, digital surface model data, digital elevation model data and surface temperature data; A plane information prediction module, used to predict the plane information of buildings corresponding to the study area based on the multimodal fusion data by using a pre-trained multimodal semantic segmentation network; A height information prediction module, used to predict the building height information corresponding to the study area based on the multimodal fusion data and the building plane information through a pre-trained height inversion network; The three-dimensional information generating module is used to generate the three-dimensional information of the buildings corresponding to the research area based on the building height information and the building plane information.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Three-dimensional building fine geometric reconstruction method integrating airborne and vehicle-mounted three-dimensional laser point clouds and streetscape images
CN111815776A
Plateau lake extraction method, device and equipment based on multi-source data and medium
CN117115666A
Semantic segmentation method and device based on multi-modal feature fusion and medium
CN118657940A
Real estate surveying and mapping method and system based on machine vision
CN118967554A
Oblique photography and three-dimensional laser point cloud fused building model construction method
CN119152142A
Cited By
Building height collaborative estimation method based on intelligent agent and multi-source data fusion
CN121053528A