Residential building three-dimensional diagram enhancement processing method
Through the collaborative work of cross-modal feature encoding and decoder, combined with convolutional neural network and Transformer, the deep fusion of image and point cloud data was successfully achieved, and high-quality three-dimensional images were generated, solving the problem of point cloud data sparsity and lack of depth information in images, and improving the accuracy and consistency of three-dimensional image reconstruction.
Patent Information
- Application Number
- CN202510743460.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
In the prior art, when processing three-dimensional images, the sparsity and irregular shape of point cloud data lead to complex processing and analysis, and the lack of depth information in traditional image data leads to insufficient accuracy in object recognition and modeling.
A cross-modal feature encoder is used to combine convolutional neural networks and Transformer to extract image features, combine random point sampling and spherical neighborhood search to extract point cloud features, and fuse image and point cloud information through a multimodal feature bridge module, enhance feature interaction using a cross-modal attention mechanism, and reconstruct three-dimensional images using deconvolution and dynamic graph convolution neural networks.
The deep fusion of image and point cloud data is achieved, and more comprehensive and accurate three-dimensional images are generated, improving the spatial consistency and image quality of reconstruction.
Smart Images

Figure CN120259102A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional map enhancement processing methods. More specifically, it particularly relates to a three-dimensional map enhancement processing method for residential buildings. Background Art
[0002] With the continuous development of three-dimensional vision technology, the processing and analysis of three-dimensional data have been widely applied in multiple fields, especially in fields such as architecture, urban planning, and autonomous driving. The acquisition of three-dimensional data usually relies on technologies such as lidar, structured light, RGB-D sensors, stereo vision systems, and deep learning models. These data provide richer spatial information than traditional two-dimensional images and can better express the geometric shape and spatial structure of objects.
[0003] In the field of three-dimensional image enhancement and reconstruction, point cloud data, as a commonly used three-dimensional data representation method, provides detailed information about the object surface and spatial position. However, the sparsity and irregularity of point cloud data make its processing and analysis relatively complex, especially in three-dimensional enhancement and reconstruction tasks. Traditional image data performs well in object detection and recognition tasks, but when dealing with three-dimensional spatial structures, due to the lack of depth information, it has limitations in accurate modeling and object recognition.
[0004] To achieve high-quality three-dimensional image enhancement, a method based on multimodal learning has been proposed. This method combines image data and three-dimensional point cloud data. Through the deep fusion of the two types of modal information, it makes up for the problem of insufficient expression ability of single-modal information and generates three-dimensional images with rich details and significantly improved overall effects. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a three-dimensional map enhancement processing method for residential buildings to solve the above problems.
[0006] A three-dimensional map enhancement processing method for residential buildings includes the following steps: S1. Obtain residential building images and three-dimensional point cloud data, and preprocess the residential building images and three-dimensional point cloud data respectively; S2. Construct a cross-modal feature encoder, which is used to extract the features of residential building images and three-dimensional point cloud data; S3. Construct a multimodal feature bridging module, which is used to fuse the features of residential building images and three-dimensional point cloud data; S4. Construct a cross-modal feature decoder, which is used to reconstruct the features fused by the multimodal feature bridging module and generate an enhanced three-dimensional map; S5. Construct a 3D map enhancement model, which consists of an input, a cross-modal feature encoder, a multi-modal feature bridging module, a cross-modal feature decoder, and an output.
[0007] Preferably, in step S2, for the cross-modal feature encoder, for the residential building image, input the residential building image into the 3D map enhancement model, , , and are the height, width, and number of channels of the residential building image respectively. Divide the residential building image into non-overlapping image patches, and use linear projection to embed each image patch into the feature vector space. , is the linear projection, is the position embedding, is the modality embedding. Use a feature extractor combining ResNet50 and Transformer to extract the features of the residential building image and obtain the residential building image features , containing image tokens, and the feature dimension of each token is . For the point cloud data, input the point cloud data into the 3D map enhancement model, , is the 3D real number space. Use random point sampling to downsample the point cloud data and select representative points. , . For each representative point , use spherical neighborhood search to determine the neighborhood of , and use a modular network based on Transformer to extract the features of the neighborhood to obtain the central feature of each neighborhood . For each , use linear projection to embed it into the feature vector space. , and obtain the point cloud features , containing point cloud tokens, and the feature dimension of each token is . In step S3, for the multi-modal feature bridging module, use the linear transformation matrices and to map the residential building image features and the point cloud features to the shared embedding space, where , , is the input feature dimension, is the dimension of the embedding space, and obtain the residential building image features and the point cloud features , satisfying , , calculate the enhanced representation of the point cloud.
[0008] Preferably, use the point cloud features as the query , the residential building image features as the key and the value , calculate the attention weights of the point cloud features , , use the value of the residential building image features and the weights of the point cloud features , calculate the enhanced representation of the point cloud, , calculate the enhanced representation of the residential building image, using the residential building image features as the query , the point cloud features as the key and the value , calculate the attention weights of the residential building image features , , use the value of the point cloud features and the weights of the residential building image features , calculate the enhanced representation of the residential building image, , use residual connections to combine the attention-enhanced features with the original features respectively, , , obtain the residential building image features and the point cloud features .
[0009] Preferably, in step S4, for the cross-modal feature decoder, for the residential building image features generated in step S3 , use deconvolution and upsampling operations to restore the residential building image features to a higher-dimensional image space, and then use a multi-layer perceptron to restore to the original image space to obtain the decoded residential building image features , for the point cloud features generated in step S3 , use back-projection operations to restore the geometric structure of the point cloud, and use a dynamic graph convolutional neural network to reconstruct the structure of the point cloud to obtain the decoded point cloud features , splice and fuse the decoded residential building image features and the point cloud features, , obtain the fused features , the fused features are used for 3D image reconstruction through 3D convolution, , generate the enhanced 3D image of the residential building .
[0010] Compared with the prior art, the present invention has the following beneficial effects: The technical solution provided by the present invention proposes a 3D map enhancement model. In the cross-modal feature encoder, a method combining a convolutional neural network and a Transformer is used to extract the features of the residential building image, and a method combining random point sampling and spherical neighborhood search with a Transformer-based modular network is used to extract the features of the 3D point cloud data; in the multi-modal feature bridging module, a cross-modal attention mechanism is used to fully fuse the residential building image and point cloud information, extract complementary features, and enhance the consistency of the reconstructed 3D space; in the cross-modal feature decoder, deconvolution, upsampling, and a multi-layer perceptron are used to restore the residential building image features to the original image space, an inverse projection operation and a dynamic graph convolutional neural network are used to reconstruct the structure of the point cloud, and a 3D convolutional operation is used to generate an enhanced 3D image of the residential building.
[0011] Advantages of multi-modal fusion: Through the collaborative work of the cross-modal feature encoder, the multi-modal feature bridging module, and the cross-modal feature decoder, the present invention successfully realizes the deep fusion of the residential building image and the 3D point cloud data. This fusion method makes full use of the texture and semantic information of the image and the geometric structure information of the point cloud, compensates for the deficiencies of single-modal data in expressing the 3D structure of the residential building, and enables the generated 3D map to present the true appearance of the residential building more comprehensively and accurately.
[0012] Effectiveness of feature extraction: In the cross-modal feature encoder, a method combining a convolutional neural network and a Transformer is used to extract the features of the residential building image, which can not only capture the local texture and semantic information of the image, but also obtain the global context relationship through the self-attention mechanism, enhancing the feature expression ability. For the 3D point cloud data, a method combining random point sampling and spherical neighborhood search with a Transformer-based modular network is used for feature extraction, which can effectively capture the local geometric features of the point cloud and provide a high-quality feature basis for subsequent fusion and reconstruction.
[0013] Enhancement of spatial consistency: In the fusion process, the cross-modal attention mechanism of the multi-modal feature bridging module enhances the interaction between the image and the point cloud by calculating the attention weights between different modal features, making the fused features have better consistency when reconstructing the 3D space. This helps to improve the accuracy and stability of the 3D map and reduce errors and distortions in the reconstruction process.
[0014] High-quality image reconstruction: The cross-modal feature decoder processes the image features of the residential building through operations such as deconvolution, upsampling, and multi-layer perceptron, which can effectively restore the details and original spatial structure of the image, ensuring an accurate match between the decoded image features and the original image space. At the same time, for the point cloud features, back-projection and dynamic graph convolutional neural network are used for reconstruction, which can accurately restore the geometric shape and spatial positioning of the point cloud, further improving the quality of the 3D map. Description of the Drawings
[0015] Figure 1 is the flowchart of the 3D map enhancement processing of the residential building provided by the present invention.
[0016] Figure 2 is the structural diagram of the cross-modal feature encoder module provided by the present invention.
[0017] Figure 3 is the structural diagram of the cross-modal feature decoder module provided by the present invention. Detailed Embodiment
[0018] The following further describes in detail the embodiments of the present invention in conjunction with the drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0019] Please refer to Figures 1 - 3 , the present invention provides a method for enhancing the 3D map of a residential building, aiming to propose a 3D map enhancement model. In the cross-modal feature encoder, a method combining a convolutional neural network and a Transformer is used to extract the features of the residential building image, and a method combining random point sampling and ball neighborhood search with a Transformer-based modular network is used to extract the features of the 3D point cloud data; in the multi-modal feature bridging module, a cross-modal attention mechanism is used to fully fuse the residential building image and point cloud information, extract complementary features, and enhance the reconstruction of 3D spatial consistency; in the cross-modal feature decoder, deconvolution, upsampling, and multi-layer perceptron are used to restore the residential building image features to the original image space, back-projection operation and dynamic graph convolutional neural network are used to reconstruct the structure of the point cloud, and 3D convolutional operation is used to generate the enhanced 3D image of the residential building.
[0020] Please refer to Figure 1 shown, a method for enhancing the 3D map of a residential building in the embodiment of the present application.
[0021] S1. Obtain the residential building image and 3D point cloud data, and preprocess the residential building image and 3D point cloud data respectively.
[0022] Further, in step S1, for the preprocessing of the residential building image, the Gaussian filtering algorithm is used to denoise the image. After normalizing the image, the image is scaled to For the preprocessing of the residential building image data set, in terms of size, the method of radiometric transformation is adopted to enhance it; for the preprocessing of 3D point cloud data, the statistical outlier removal algorithm is used to remove the noisy point cloud data, and the voxel grid downsampling method is adopted to reduce the number of point clouds. At the same time, sufficient structural information is retained, surface reconstruction is performed on the sparse point cloud to fill the vacant areas between the point clouds, and for the point cloud after surface reconstruction, a smoothing algorithm is used to remove irregular or abnormal points to improve the quality of the point cloud.
[0023] S2. Construct a cross-modal feature encoder, which is used to extract the features of the residential building image and 3D point cloud data.
[0024] Furthermore, in step S2, for the cross-modal feature encoder, its structure is as Figure 2 shown. For the residential building image, the residential building image is input into the 3D map enhancement model, , , and are respectively the height, width and number of channels of the residential building image. The residential building image is divided into non-overlapping image patches, and each image patch is embedded into the feature vector space using linear projection, , is the linear projection, is the positional embedding, is the modality embedding. A feature extractor combining ResNet50 and Transformer is used to extract the features of the residential building image, and the residential building image feature is obtained, which contains image tokens, and the feature dimension of each token is ; for the point cloud data, the point cloud data is input into the 3D map enhancement model, , is the 3D real number space. Random point sampling is used to downsample the point cloud data, and representative points are selected, , , for each representative point , the neighborhood of is determined using spherical neighborhood search, and a modular network based on Transformer is used to extract the features of the neighborhood, and the central feature of each neighborhood is obtained. For each , it is embedded into the feature vector space using linear projection, , and the point cloud feature is obtained, which contains point cloud tokens, and the feature dimension of each token is .
[0025] S3. Construct a multimodal feature bridging module, which is used to fuse the features of the residential building image and the 3D point cloud data.
[0026] Further, in step S3, for the multimodal feature bridging module, use the linear transformation matrix and to map the residential building image features and the point cloud features to the shared embedding space, where 、 , is the input feature dimension, is the dimension of the embedding space, and the residential building image features and the point cloud features are obtained, satisfying , ,calculate the enhanced representation of the point cloud, and use the point cloud feature as the query ,the residential building image feature as the key and the value to calculate the attention weight , ,use the value of the residential building image feature and the weight of the point cloud feature to calculate the enhanced point cloud representation, ; calculate the enhanced representation of the residential building image, use the residential building image feature as the query ,the point cloud feature as the key and the value to calculate the attention weight , ,use the value of the point cloud feature and the weight of the residential building image feature to calculate the enhanced residential building image representation, ,use the residual connection to combine the attention-enhanced features with the original features respectively, , ,and obtain the residential building image feature and the point cloud feature .
[0027] S4. Construct a cross-modal feature decoder, which is used to reconstruct the features fused by the multimodal feature bridging module and generate an enhanced 3D map.
[0028] Further, in step S4, for the cross-modal feature decoder, its structure is as Figure 3 shown. For the residential building image features generated in step S3 Restore the residential building image features to a higher-dimensional image space using deconvolution and upsampling operations, and then restore them to the original image space using a multi-layer perceptron to obtain the decoded residential building image features ; For the point cloud features generated in step S3 , restore the geometric structure of the point cloud using back-projection operations, and reconstruct the structure of the point cloud using a dynamic graph convolutional neural network to obtain the decoded point cloud features ; Concatenate and fuse the decoded residential building image features and point cloud features to obtain the fused features ; The fused features are used to reconstruct a three-dimensional image through 3D convolution to generate an enhanced three-dimensional image of the residential building . .
[0029] S5. Construct a three-dimensional graph enhancement model, which consists of an input, a cross-modal feature encoder, a multi-modal feature bridging module, a cross-modal feature decoder, and an output
[0030] Further, in step S5, for the three-dimensional graph enhancement model, it is written based on the Pytorch framework, using the Adam optimizer, the mean squared error loss function, a learning rate of 0.001, and a training batch size of 300
[0031] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical applications, and to enable those of ordinary skill in the art to understand the invention and design various embodiments with various modifications suitable for specific purposes
Claims
1. A method for enhancing the three-dimensional drawing of a residential building, characterized in that, Including the following steps: S1. Obtain the residential building image and 3D point cloud data, and perform preprocessing on the residential building image and 3D point cloud data respectively; S2. Construct a cross-modal feature encoder, which is used to extract the features of the residential building image and 3D point cloud data; S3. Construct a multi-modal feature bridging module, which is used to fuse the features of the residential building image and 3D point cloud data; S4. Construct a cross-modal feature decoder, which is used to reconstruct the features fused by the multi-modal feature bridging module and generate an enhanced 3D map; S5. Construct a 3D map enhancement model, which consists of an input, a cross-modal feature encoder, a multi-modal feature bridging module, a cross-modal feature decoder and an output.
2. The three-dimensional map enhancement processing method for a residential building according to claim 1, wherein, In step S2, for the cross-modal feature encoder, for the residential building image, input the residential building image to the 3D map enhancement model, , , and are the height, width, and number of channels of the residential building image respectively. Divide the residential building image into non-overlapping image patches, and use linear projection to embed each image patch into the feature vector space, , is the linear projection, is the position embedding, is the modality embedding. Use a feature extractor that combines ResNet50 and Transformer to extract the features of the residential building image and obtain the residential building image features , which contains image tokens, and the feature dimension of each token is .
3. The three-dimensional map enhancement processing method for a residential building according to claim 2, wherein For point cloud data, input the point cloud data to the 3D map enhancement model, , which is a three-dimensional real space. Downsample the point cloud data using random point sampling and select representative points, , , for each representative point , use spherical neighborhood search to determine its neighborhood, and use a Transformer-based modular network to extract features from the neighborhood to obtain the central feature of each neighborhood , for each embed it into the feature vector space using linear projection, , to obtain the point cloud feature , which contains point cloud tokens, and the feature dimension of each token is .
4. The three-dimensional map enhancement processing method for a residential building according to claim 1, characterized in that, In step S3, for the multi-modal feature bridging module, a linear transformation matrix and are used to map the residential building image features and point cloud features into a shared embedding space, where 、 , is the input feature dimension, is the dimension of the embedding space, and the residential building image features and the point cloud features are obtained, satisfying , , and the enhanced representation of the point cloud is calculated.
5. The three-dimensional map enhancement processing method for a residential building according to claim 4, wherein, Using point cloud features as queries and residential building image features as keys and values to calculate the attention weights of the point cloud features , and using the values of the residential building image features and the weights of the point cloud features to calculate the enhanced representation of the point cloud .
6. The three-dimensional drawing enhancement processing method of a residential building according to claim 5, wherein Calculate an enhanced representation of the residential building image using the residential building image features As a query , point cloud features As keys And values , calculate the attention weights of the residential building image features , , using the values of the point cloud features And the weights of the residential building image features , calculate the enhanced representation of the residential building image .
7. The three-dimensional map enhancement processing method for a residential building according to claim 6, characterized in that, Use residual connections to combine the attention-enhanced features with the original features respectively, , , obtaining the image features of the residential building and the point cloud features .
8. The three-dimensional map enhancement processing method for a residential building according to claim 7, wherein In step S4, for the cross-modal feature decoder, for the residential building image features generated in step S3 , the deconvolution and upsampling operations are used to restore the residential building image features to a higher-dimensional image space, and then a multi-layer perceptron is used to restore them to the original image space, obtaining the decoded residential building image features .
9. The three-dimensional map enhancement processing method for a residential building according to claim 8, wherein, For the point cloud features generated in step S3 , use the back-projection operation to restore the geometric structure of the point cloud, and use the dynamic graph convolutional neural network to reconstruct the structure of the point cloud to obtain the decoded point cloud features .
10. The three-dimensional map enhancement processing method for a residential building according to claim 9, wherein, The decoded image features and point cloud features of the residential building are spliced and fused, to obtain the fused features The fused features Perform 3D image reconstruction through 3D convolution, to generate the enhanced 3D image of the residential building .
Citation Information
Patent Citations
Dual-modal target detection method and system based on cross-modal attention mechanism fusion
CN117422971A
Three-dimensional anomaly detection method based on image-point cloud data double-branch hybrid model
CN118628808A
Multi-modal three-dimensional target detection method and device for automatic driving
CN118674916A
Unmanned aerial vehicle 3D target detection multi-modal fusion method based on Transform
CN118837875A
Object point cloud reconstruction method based on guided cross-modal robot
CN119478220A