Building model generation method and device, electronic equipment and medium

By combining building segmentation models and surface reconstruction algorithms, the problem of low accuracy in extracting building point clouds in large-scale urban scenes is solved, and efficient and high-precision reconstruction of building 3D models is achieved.

CN121904259APending Publication Date: 2026-04-21SICHUAN JIANSHAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN JIANSHAN TECH CO LTD
Filing Date
2024-10-21
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of building point cloud extraction in large-scale urban scenes is low, which affects the efficiency and accuracy of reconstructing building models.

Method used

By acquiring orthophotos and point cloud data, image segmentation is performed using a building segmentation model to generate a building segmentation mask, which is then matched with the point cloud data. Finally, a surface reconstruction algorithm is used to classify, filter, and reconstruct the point cloud data to generate a 3D model of the building.

Benefits of technology

It improves the accuracy of building point cloud data extraction and the efficiency and precision of building 3D model reconstruction, ensuring the accuracy of building boundary information and the integrity of overall point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904259A_ABST
    Figure CN121904259A_ABST
Patent Text Reader

Abstract

According to the building model generation method and device, the electronic equipment and the medium provided by the invention, image segmentation is performed on the orthoimage based on the building segmentation model to obtain the building segmentation mask; matching the building segmentation mask with the point cloud data to obtain fusion data of the building mask and the point cloud; determining overall point cloud data of the building based on the fused data; and carrying out point cloud classification, point cloud screening and point cloud reconstruction on the overall point cloud data of the building based on a surface reconstruction algorithm to generate a building three-dimensional model. According to the method and the device, the problem that the three-dimensional point cloud data is difficult to show accurate boundary information is avoided, the extraction precision of the building point cloud data is improved, large-scale point cloud data is screened in the building surface reconstruction process, the processing amount of the point cloud data in the reconstruction process is reduced, the building model reconstruction efficiency is improved, and therefore, the method and the device are suitable for popularization and application. By means of the method, the efficiency and precision of building three-dimensional model reconstruction are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer data processing technology, and in particular to methods, apparatus, electronic devices and media for generating building models. Background Technology

[0002] The emergence of concepts such as smart cities, digital twins of cities, and digital earth has accelerated the pace of urban development. Buildings, roads, vegetation, and other geographical features are crucial components of urban planning and construction, and their representation in virtual scenarios is particularly important. Buildings, as the core of the city, require reconstructed models for the creation of virtual environments.

[0003] Currently, when reconstructing buildings in urban scenes, point cloud data of the entire urban scene is typically collected. However, urban scenes contain numerous objects, such as buildings, roads, and vegetation. To reconstruct buildings in urban scenes, it is necessary to extract building point cloud data from the overall urban scene point cloud data, and then reconstruct the buildings based on the extracted building point cloud data. Existing technologies typically use rule-based classifiers, decision trees, and random forests to extract building point clouds. However, for large-scale urban scenes, the scenes are complex and the point cloud data is massive. Using existing building point cloud extraction methods to extract building point clouds from large-scale scenes can easily lead to low accuracy of the extracted building point clouds, thus affecting the efficiency and accuracy of the reconstructed building models.

[0004] Improving the efficiency and accuracy of reconstructed building models is an urgent problem to be solved in related technologies. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a building model generation method, apparatus, electronic device and medium to improve the accuracy of the reconstructed building model.

[0006] In a first aspect, embodiments of this application provide a method for generating a building model, the method comprising:

[0007] Acquire orthophotos of the target city and point cloud data of the target city;

[0008] The orthophoto is segmented based on a building segmentation model to obtain the building segmentation mask for the target city.

[0009] The building segmentation mask is matched with the point cloud data of the target city to obtain the fused data of the building mask and point cloud in the target city;

[0010] The overall point cloud data of buildings in the target city is determined based on the fusion data of building masks and point clouds in the target city;

[0011] Based on the surface reconstruction algorithm, the overall point cloud data of buildings in the target city is classified, filtered, and reconstructed to generate a 3D model of the buildings in the target city.

[0012] In one embodiment, matching the building segmentation mask with the point cloud data of the target city to obtain fused data of the building mask and point cloud in the target city includes:

[0013] Multiple key points are selected from the building segmentation mask, and the position coordinates of each key point are determined.

[0014] The point in the point cloud data of the target city that corresponds to the position coordinates of the target mask key point is determined as the point cloud key point that matches the target mask key point; the target mask key point is any one of multiple mask key points;

[0015] The fused data of building masks and point clouds in the target city is determined based on the point cloud key points that match each mask key point.

[0016] In one embodiment, determining the overall point cloud data of buildings in the target city based on the fusion data of building masks and point clouds in the target city includes:

[0017] The building boundary contours in the target city are determined based on the fusion data of building masks and point clouds in the target city.

[0018] Starting from the target point cloud data, a target ray corresponding to the target point cloud data is formed in a specified direction, wherein the target point cloud data is any point cloud data in the point cloud data of the target city;

[0019] Whether the target point cloud data is within the building boundary contour is determined based on the number of intersections between the target ray and the building boundary contour.

[0020] All point cloud data within the boundary contour of the building are identified as the overall point cloud data of the buildings in the target city.

[0021] In one embodiment, determining all point cloud data within the building boundary contour as the overall point cloud data of buildings in the target city includes:

[0022] Any point cloud in the fused data of the building mask and point cloud is determined as the target point cloud of the building outline;

[0023] Point cloud data in the point cloud data of the target city whose distance to the target point cloud of the building outline is within a preset distance threshold are determined as candidate point clouds corresponding to the target point cloud of the building outline;

[0024] The candidate point clouds corresponding to the fused data of all the building masks and point clouds, as well as all the point cloud data within the boundary contours of the buildings, are determined as the overall point cloud data of the buildings in the target city.

[0025] In one embodiment, the network structure of the surface reconstruction algorithm includes a feature classification unit, a feature filtering unit, and a feature reconstruction unit. The step of performing point cloud classification, point cloud filtering, and point cloud reconstruction on the overall point cloud data of buildings in the target city based on the surface reconstruction algorithm to generate a 3D model of the buildings in the target city includes:

[0026] Based on the feature classification unit, the overall point cloud data of buildings in the target city is classified to obtain classified point cloud data, which includes point cloud of building surface convexity, point cloud of building surface corner, and point cloud of building surface inflection point.

[0027] Based on the feature filtering unit, key feature point cloud filtering is performed on the point cloud data after target category classification to obtain key point cloud data of target category in the building. The point cloud data after target category classification is any one of the following: building surface convex point cloud, building surface corner point cloud, and building surface inflection point point cloud.

[0028] Based on the feature reconstruction unit, the point clouds of all types of buildings are reconstructed to generate a 3D model of the buildings in the target city.

[0029] In one embodiment, the reconstruction of the point cloud of all categories of buildings based on the feature reconstruction unit to generate a 3D model of the buildings in the target city includes:

[0030] The point cloud level of the building selection point cloud for each category is determined based on the feature encoding unit.

[0031] Based on the level of the point cloud, a nearest neighbor threshold is set for each category of building point cloud.

[0032] The building selection point cloud of each category is encoded based on the nearest neighbor threshold corresponding to the building selection point cloud of each category to obtain the corresponding encoded building features;

[0033] Based on the feature reconstruction unit, feature reconstruction is performed on all encoded building features to generate a 3D model of the buildings in the target city.

[0034] In one embodiment, the step of performing feature reconstruction on all encoded building features based on the feature reconstruction unit to generate a 3D model of the buildings in the target city includes:

[0035] The feature representation capability of the corresponding encoded building features is enhanced to obtain the corresponding enhanced building features;

[0036] Based on the feature reconstruction unit, feature reconstruction is performed on all enhanced building features to generate a three-dimensional model of the buildings in the target city.

[0037] Secondly, embodiments of this application also provide a building model generation apparatus, the apparatus comprising:

[0038] The acquisition module is used to acquire orthophotos of the target city and point cloud data of the target city;

[0039] The image segmentation module is used to segment the orthophoto based on the building segmentation model to obtain the building segmentation mask of the target city.

[0040] The matching module is used to match the building segmentation mask with the point cloud data of the target city to obtain the fused data of the building mask and point cloud in the target city;

[0041] The point cloud determination module is used to determine the overall point cloud data of buildings in the target city based on the fusion data of building masks and point clouds in the target city;

[0042] The model reconstruction module is used to classify, filter, and reconstruct the overall point cloud data of buildings in the target city based on the surface reconstruction algorithm, and generate a three-dimensional model of the buildings in the target city.

[0043] Thirdly, embodiments of this application also provide an electronic device, including: a memory and a processor, wherein the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, it performs the building model generation method as described in the first aspect above.

[0044] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer program instructions that, when executed by a processor, perform the building model generation method as described in the first aspect above.

[0045] In the above implementation process, the orthophoto is segmented by the building segmentation model to obtain the building segmentation mask. The building segmentation mask is obtained by processing the orthophoto with coding units of different convolution kernel sizes to obtain coding information of different precisions, and then merging the coding information of different precisions. Thus, the obtained building segmentation mask can accurately represent the building outlines in the target city. Furthermore, the building segmentation mask, including the building outline, is matched with the point cloud data of the target city to obtain fused data of the building mask and point cloud in the target city. Based on the fused data of the building mask and point cloud in the target city, the overall point cloud data of the buildings in the target city is determined. This effectively matches the building segmentation mask, which includes two-dimensional building outline information, with three-dimensional point cloud data, which includes accurate spatial structure information. This avoids the problem that three-dimensional point cloud data is difficult to represent accurate boundary information, thereby effectively improving the accuracy of extracting point cloud data of buildings in the target city. Then, the extracted point cloud data is classified, filtered, and reconstructed according to the surface reconstruction algorithm to obtain the three-dimensional model of the buildings in the target city. This reduces the amount of point cloud data processing. By constructing the three-dimensional model of the buildings with accurate building point cloud data, the efficiency and accuracy of the reconstructed three-dimensional building model are effectively improved. Attached Figure Description

[0046] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a schematic diagram of a building model generation method provided in an embodiment of this application;

[0048] Figure 2 This is a schematic diagram of a building segmentation mask provided in an embodiment of this application;

[0049] Figure 3 This is a schematic diagram of a light projection method provided in an embodiment of this application;

[0050] Figure 4 This is a schematic diagram of the network structure of a surface reconstruction algorithm provided in an embodiment of this application;

[0051] Figure 5 This is a schematic diagram of the network structure of another surface reconstruction algorithm provided in this application embodiment;

[0052] Figure 6This is a schematic diagram of the structure of a building segmentation model provided in an embodiment of this application;

[0053] Figure 7 This is a schematic diagram of the structure of a building model generation device provided in an embodiment of this application;

[0054] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0055] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application.

[0057] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" and "a variety" mean two or more, unless otherwise explicitly defined.

[0058] The emergence of concepts such as smart cities, digital twins of cities, and digital earth has accelerated the pace of urban development. Buildings, roads, vegetation, and other geographical features are crucial components of urban planning and construction, and their representation in virtual scenarios is particularly important. Buildings, as the core of the city, require reconstructed models for the creation of virtual environments.

[0059] Currently, when reconstructing buildings in urban scenes, point cloud data of the entire urban scene is typically collected. However, urban scenes contain numerous objects, such as buildings, roads, and vegetation. To reconstruct buildings in urban scenes, it is necessary to extract building point cloud data from the overall urban scene point cloud data, and then reconstruct the buildings based on the extracted building point cloud data. Existing technologies typically use rule-based classifiers, decision trees, and random forests to extract building point clouds. However, for large-scale urban scenes, the scenes are complex and the point cloud data is massive. Using existing building point cloud extraction methods to extract building point clouds from large-scale scenes can easily lead to low accuracy of the extracted building point clouds, thus affecting the efficiency and accuracy of the reconstructed building models.

[0060] Improving the efficiency and accuracy of reconstructing 3D building models is an urgent problem to be solved in related technologies.

[0061] Please see Figure 1 , Figure 1 This is a schematic flowchart of a building model generation method provided in an embodiment of this application. The building model generation method may include the following steps, such as steps S101-S105:

[0062] Step S101: Obtain the orthophoto of the target city and the point cloud data of the target city.

[0063] Step S102: Perform image segmentation on the orthophoto based on the building segmentation model to obtain the building segmentation mask of the target city.

[0064] The network structure of the building segmentation model includes a first encoding unit and a second encoding unit. The kernel of the feature convolutional layer in the first encoding unit is smaller than the kernel of the feature convolutional layer in the second encoding unit.

[0065] Step S103: Match the building segmentation mask with the point cloud data of the target city to obtain the fused data of the building mask and point cloud in the target city.

[0066] Step S104: Determine the overall point cloud data of buildings in the target city based on the fusion data of building masks and point clouds in the target city.

[0067] Step S105: Based on the surface reconstruction algorithm, perform point cloud classification, point cloud filtering and point cloud reconstruction on the overall point cloud data of buildings in the target city to generate a 3D model of the buildings in the target city.

[0068] For example, orthophotos are remote sensing images with orthogonal projection properties. Orthophotos of a target city can be acquired by drones, and these orthophotos are typically two-dimensional images that include urban landscapes such as buildings, roads, and vegetation.

[0069] Point cloud data can be generated by LiDAR measuring the distance between an object and a sensor by emitting laser pulses and receiving reflected signals, calculating the time required for the laser pulse to travel from emission to return, and converting this time into distance. Point cloud data for a target city can be obtained by a drone equipped with LiDAR scanning objects within the city.

[0070] For example, the building segmentation model is obtained by training an initial segmentation model. Specifically, before using the building segmentation model to segment buildings in orthophotos, remote sensing data can be collected and training, validation, and test sets can be created. The resolution of remote sensing images is generally 8192*8192 pixels. In the early stage, the images need to be preprocessed to adjust the image size to 1024*1024 pixels. Alternatively, publicly available building segmentation datasets such as WHU Building Dataset, Buildings2Vec, and ChesapeakRSC can be used to train and validate the initial segmentation model, thereby obtaining the building segmentation model. The model structure of the initial building segmentation model can be obtained by improving the DeepLabV3+ network.

[0071] Furthermore, a building segmentation model is used to segment the orthophoto image, obtaining a building segmentation mask for the target city. This building segmentation mask includes the contour region of each building. Figure 2 This is a schematic diagram of a building segmentation mask provided in an embodiment of this application, such as... Figure 2 As shown, Figure 2 Each white area in the image corresponds to a segmentation mask for a building.

[0072] Specifically, the network structure of the building segmentation model includes a first encoding unit and a second encoding unit. The kernel of the feature convolutional layer in the first encoding unit is smaller than that in the feature convolutional layer in the second encoding unit. By extracting features from orthophotos using this building segmentation model, orthophotos can be processed by encoding units with different kernel sizes to obtain encoding information of different precision. Then, the building segmentation mask in the orthophoto is determined based on the encoding information of different precision, thereby improving the accuracy of building segmentation mask extraction.

[0073] Then, the 2D building segmentation mask is matched with the corresponding 3D point cloud data to obtain fused data of the building mask and point cloud in the target city. Specifically, key points in the 2D building segmentation mask can be identified, which are called mask key points. Further, based on the position coordinate information of the mask key points, the point cloud corresponding to the position coordinate information of the mask key points is determined in the 3D point cloud data, which are called point cloud key points, thus obtaining fused data of the building mask and point cloud in the target city.

[0074] Furthermore, the overall point cloud data of buildings in the target city can be determined based on the fusion data of building masks and point clouds in the target city. Specifically, the overall point cloud data of buildings in the target city can be determined using the ray casting method.

[0075] Furthermore, based on the surface reconstruction algorithm, the determined overall point cloud of buildings is classified, filtered, and reconstructed to generate a 3D model of the buildings in the target city. Specifically, the surface reconstruction algorithm can be an improved version of the DeepDT deep learning-based Delaunay triangulation surface reconstruction network. The algorithm sequentially classifies, filters, and reconstructs the overall point cloud of buildings to ultimately generate a 3D model of the buildings in the target city.

[0076] Because 3D point cloud data struggles to accurately represent boundary information, the extracted building point cloud data often suffers from low accuracy in boundary information. Furthermore, the reconstruction of building models using massive amounts of point cloud data leads to low efficiency. In this embodiment, an orthophoto is segmented using a building segmentation model to obtain a building segmentation mask. This mask is obtained by processing the orthophoto with coding units of different convolutional kernel sizes to obtain coding information of varying precision, and then merging these different precision coding information. This results in a building segmentation mask that accurately represents the building outlines in the target city. Furthermore, the building segmentation mask, including the building outline, is matched with the point cloud data of the target city to obtain fused data of the building mask and point cloud in the target city. This effectively matches the building segmentation mask, which includes two-dimensional building outline information, with three-dimensional point cloud data, which includes accurate spatial structure information, avoiding the problem that three-dimensional point cloud data is difficult to represent accurate boundary information. Based on the fused data of the building mask and point cloud in the target city, the overall point cloud data of the buildings in the target city is determined, thereby effectively improving the accuracy of the extraction of building point cloud data in the target city. Then, the extracted point cloud data is classified, filtered, and reconstructed according to the surface reconstruction algorithm to obtain the three-dimensional model of the buildings in the target city. By constructing the three-dimensional model of the buildings using accurate building point cloud data, the efficiency and accuracy of the three-dimensional model reconstruction of the buildings are effectively improved.

[0077] In one embodiment, matching building segmentation masks with point cloud data of the target city to obtain fused data of building masks and point clouds in the target city may include the following steps:

[0078] Step 1: Select multiple mask key points from the building segmentation mask and determine the position coordinates of each mask key point.

[0079] Step 2: Determine the points in the point cloud data of the target city that correspond to the location coordinates of the target mask key points as the point cloud key points that match the target mask key points.

[0080] Among them, the target mask key point is any one of multiple mask key points.

[0081] Step 3: Determine the fusion data of building masks and point clouds in the target city based on the point cloud key points that match each mask key point.

[0082] For example, at least one of the following algorithms can be used to extract feature points and feature descriptors of buildings in the building segmentation mask, thereby obtaining multiple mask key points in the building segmentation mask and determining the position coordinates of each mask key point in the building segmentation mask.

[0083] Any one of the multiple mask keypoints is used as the target mask keypoint, and the points in the point cloud data of the target city that correspond to the position coordinates of the target mask keypoint are determined as the point cloud keypoints that match the target mask keypoint. Specifically, if the building segmentation mask is a two-dimensional image, then the position coordinates of the target mask keypoint are two-dimensional coordinates (x, y), and the point cloud data of the target city is three-dimensional data (x', y', z'). Therefore, the point clouds in the point cloud data of the target city whose position coordinates (x', y') have the same value as (x, y) are determined as the point cloud keypoints that match the target mask keypoint, thus ensuring a one-to-one correspondence between the target mask keypoints in the building segmentation mask and the point cloud keypoints in the point cloud data of the target city.

[0084] This method is used to identify point cloud key points that match all mask key points, thereby obtaining fused data of building masks and point clouds in the target city.

[0085] In the above implementation process, multiple mask key points are selected from the building segmentation mask, and the position coordinates of each mask key point are determined. Then, based on the position coordinates of the mask key points, point cloud key points corresponding to the position coordinates of each mask key point are determined from the point cloud data of the target city. This improves the accuracy of determining the fusion data of building masks and point clouds in the target city and realizes the determination of the boundary position of building point clouds.

[0086] In one embodiment, determining the overall point cloud data of buildings in the target city based on the fusion data of building masks and point clouds in the target city may include the following steps:

[0087] Step 1: Determine the building boundary contours in the target city based on the fusion data of building masks and point clouds in the target city.

[0088] Step 2: Starting from the target point cloud data, form a target ray corresponding to the target point cloud data in the specified direction.

[0089] The target point cloud data is any point cloud data in the point cloud data of the target city.

[0090] Step 3: Determine whether the target point cloud data is within the building boundary contour based on the number of intersections between the target ray and the building boundary contour.

[0091] Step 4: Determine all point cloud data within the building boundary outline as the overall point cloud data of buildings in the target city.

[0092] For example, ray casting can be used to determine the overall point cloud data of buildings in a target city. Figure 3 This is a schematic diagram of a ray casting method provided in an embodiment of this application. The boundary contours of buildings in the target city are determined based on the fusion data of building masks and point clouds in the target city. Figure 3 The octagon shown represents the boundary outline of buildings in the target city.

[0093] Furthermore, any point cloud data point in the point cloud data of the target city is identified as the target point cloud data. Figure 3 Any point in the target point cloud can be used as the target point cloud data. A ray is emitted from the target point cloud data in a specified direction, that is, the target ray corresponding to the target point cloud data, and the specified direction is to the right.

[0094] It should be noted that the embodiment of this application uses the specified direction to the right as an example for illustration. In actual application, the specified direction can also be to the left, that is, the specified direction can be adaptively adjusted according to the actual situation.

[0095] The number of intersections between the target ray and the building boundary contour determines whether the target point cloud data is within the building boundary contour. If the number of intersections is odd, the target point cloud data is within the building boundary contour; if the number of intersections is even, the target point cloud data is not within the building boundary contour. Furthermore, this method is used to determine whether all point cloud data in the target city's point cloud data are within the building boundary contour.

[0096] Furthermore, all point cloud data within the building boundary contours are identified as the overall point cloud data of buildings in the target city.

[0097] In the above implementation process, the point cloud data within the building boundary contour is determined by the number of intersections between the ray corresponding to each point cloud data in the target city and the building boundary contour, thereby improving the accuracy of determining the overall point cloud data of buildings in the target city.

[0098] In one embodiment, determining all point cloud data within the building boundary contour as the overall point cloud data of buildings in the target city may include the following steps:

[0099] Step 1: Determine any point cloud from the fused data of the building mask and point cloud as the target point cloud of the building outline.

[0100] Step 2: Select the point cloud data of the target city whose distance to the target point cloud of the building outline is within a preset distance threshold as the candidate point cloud corresponding to the target point cloud of the building outline.

[0101] Step 3: Determine the candidate point clouds corresponding to the fused data of all building masks and point clouds, as well as all point cloud data within the building boundary contours, as the overall point cloud data of buildings in the target city.

[0102] For example, any point cloud in the fused data of building mask and point cloud is determined as the target point cloud of building outline, and a preset distance threshold is set. The preset distance threshold can be 3cm, 5cm, 10cm or other values, which are not limited here.

[0103] Furthermore, point cloud data in the target city whose distance to the building outline target point cloud is within a preset distance threshold are identified as candidate point clouds corresponding to the building outline target point cloud. In addition, this method is used to determine the candidate point clouds corresponding to the fused data of all building masks and point clouds.

[0104] The candidate point clouds corresponding to the fused data of all building masks and point clouds are combined with all point cloud data within the building boundary contours to determine the overall point cloud data of buildings in the target city.

[0105] In the above implementation process, point cloud data in the target city whose distance to the building outline target point cloud is within a preset distance threshold is determined as candidate point clouds corresponding to the building outline target point cloud. Thus, point clouds around the fused data of the building mask and point cloud within the preset distance threshold are used as candidate point clouds. All candidate point clouds corresponding to the fused data of the building mask and point cloud, as well as all point cloud data within the building boundary outline, are determined together as the overall point cloud data of the buildings in the target city. This avoids the loss of the overall point cloud data of the buildings in the target city due to errors in the building boundary position, ensures the integrity of the overall point cloud data of the buildings in the target city, and further improves the accuracy of the 3D model reconstruction of the buildings in the target city.

[0106] In one embodiment, the network structure of the surface reconstruction algorithm includes a feature classification unit, a feature filtering unit, and a feature reconstruction unit. Based on the surface reconstruction algorithm, point cloud classification, point cloud filtering, and point cloud reconstruction are performed on the overall point cloud data of buildings in the target city to generate a 3D model of the buildings in the target city. This may include the following steps:

[0107] Step 1: Classify the overall point cloud data of buildings in the target city based on feature classification units to obtain classified point cloud data.

[0108] The classified point cloud data includes point clouds of building surface convexity and concaveness, point clouds of building surface corners, and point clouds of building surface inflection points.

[0109] Step 2: Based on the feature filtering unit, perform key feature point cloud filtering on the point cloud data after classifying the target categories to obtain key point cloud data of the target categories in the building.

[0110] The point cloud data after target category classification is any one of the following: point cloud of building surface convexity, point cloud of building surface corner, and point cloud of building surface inflection point.

[0111] Step 3: Based on the feature reconstruction unit, the point cloud of all types of buildings is reconstructed to generate a 3D model of the buildings in the target city.

[0112] For example, Figure 4 This is a schematic diagram of the network structure of a surface reconstruction algorithm provided in an embodiment of this application, such as... Figure 4 As shown, the network structure of the surface reconstruction algorithm includes a feature classification unit, a feature selection unit, and a feature reconstruction unit. The surface reconstruction algorithm can be obtained by improving the feature extraction part of the DeepDT network and training it. Specifically, the feature extraction part of the DeepDT network is divided into a feature classification unit and a feature selection unit. The feature classification unit classifies the building point cloud data to obtain classified point cloud data, which includes point clouds of building surface convexity and concaveness, point clouds of building surface edges and corners, and point clouds of building surface inflection points. Specifically, the feature classification unit can be the classification network in the PointNet++ network.

[0113] Furthermore, point clouds of any category among the building surface convexity / concaveness point clouds, building surface corner point clouds, and building surface inflection point point clouds are identified as point cloud data after target category classification. A feature filtering unit then performs key feature point cloud filtering on the target category-classified point cloud data to obtain key point cloud data for the building. This method is used to filter key point clouds for each category of classified point cloud data, reducing the number of building point clouds and achieving lightweighting of each category's building point cloud. This results in the selected point cloud for each category, ensuring that the lightweight point cloud data only includes key point information of the building surface convexity / concaveness, corners, inflection points, etc. Specifically, the building point cloud filtering method can include at least one of point cloud filtering, point cloud sampling, point cloud compression, and feature extraction and balancing, without limitation.

[0114] When training the surface reconstruction model, publicly available datasets can be used for training and validation, such as the German Vaihingen dataset, the DTU dataset, and the DublinCity dataset. Alternatively, a custom point cloud dataset can be created. Oblique images captured by drone aerial photography are input into Capture Content software to generate point cloud data. In areas with missing information, LiDAR is used for targeted scanning, and the generated point cloud data fills in the missing information. Finally, the surface reconstruction model is obtained.

[0115] Then, the feature reconstruction unit of the DeepDT network is used to reconstruct the point cloud of all types of buildings to generate a 3D model of the buildings in the target city.

[0116] In the above implementation process, the building point cloud is classified by the feature classification unit to obtain building point cloud data of different categories. The key point cloud of each category of buildings is filtered by the feature filtering unit, thereby reducing the number of building point clouds and realizing the lightweighting of building point clouds of each category. Then, the key point cloud of all categories of buildings is reconstructed by the feature reconstruction unit to generate the 3D model of the target city's buildings, which effectively improves the generation efficiency of the 3D model of buildings.

[0117] In one embodiment, reconstructing the point clouds of all categories of buildings based on feature reconstruction units to generate a 3D model of buildings in the target city may include the following steps:

[0118] Step 1: Determine the point cloud level for the point cloud of each building category.

[0119] Step 2: Set the nearest neighbor threshold for each type of building point cloud based on the level of the point cloud.

[0120] Step 3: Encode the corresponding building selection point cloud based on the nearest neighbor threshold of each building selection point cloud to obtain the corresponding encoded building features.

[0121] Step 4: Based on the feature reconstruction unit, perform feature reconstruction on all encoded building features to generate a 3D model of the buildings in the target city.

[0122] For example, the building selection point cloud includes three categories. The point cloud level of each category can be determined according to the building selection point cloud category. Specifically, the building selection point cloud can be divided into three levels. For example, the point cloud of the building surface undulation is divided into level 1, the point cloud of the building surface corner is divided into level 2, and the point cloud of the building surface inflection point is divided into level 3.

[0123] Furthermore, based on the level of the point cloud, a nearest neighbor threshold is set for the point cloud filtering of each building category. Specifically, a nearest neighbor threshold can be set for each building category by setting a lower threshold for higher levels. For example, the nearest neighbor threshold for the bump point cloud of a building surface at level 1 is 'a', the nearest neighbor threshold for the corner point cloud of a building surface at level 2 is 'b', and the nearest neighbor threshold for the inflection point cloud of a building surface at level 3 is 'c', where a>b>c.

[0124] Figure 5 This is a schematic diagram of the network structure of another surface reconstruction algorithm provided in this application embodiment, such as... Figure 5 As shown, the network structure of the surface reconstruction algorithm can also include a feature encoding unit. This unit can use nearest neighbor thresholds to encode the point clouds of corresponding buildings. Specifically, nearest neighbor threshold 'a' is used to geometrically encode the concave and convex point cloud data of the building surface, obtaining the encoded building features of the concave and convex parts; nearest neighbor threshold 'b' is used to geometrically encode the point cloud of the building surface corners, obtaining the encoded building features of the corner parts; and nearest neighbor threshold 'c' is used to geometrically encode the point cloud of the building surface inflection points, obtaining the encoded building features of the inflection points. By encoding the point cloud features of corresponding categories using different nearest neighbor thresholds, the geometric features of the point cloud are enriched. Specifically, the encoding process can involve creating a tangent plane perpendicular to the normal of each point in the k nearest neighbors and calculating the signed distance from the current point to these k planes. Based on this, the point cloud normal is decomposed into the relative normals of the tangent planes of neighboring points as input to provide richer local geometric information about the surface.

[0125] The feature reconstruction unit reconstructs the features of all encoded buildings to generate a 3D model of the buildings in the target city. For example... Figure 5 As shown, the feature reconstruction unit may include a Delaunay triangulation layer, a graph network layer, and a surface reconstruction layer.

[0126] Specifically, the Delaunay triangulation layer divides all the encoded building features into tetrahedral meshes using the Delaunay triangulation algorithm, thereby capturing the geometric and topological information in the point cloud more effectively through the tetrahedral network structure.

[0127] Furthermore, the obtained tetrahedral network structure is transformed into tetrahedral nodes expressed by a dual graph through a graph network layer, where each tetrahedral node becomes a node in the graph. The tetrahedral nodes are then classified according to the graph network in the graph network layer to obtain the classification result. This classification result helps to extract a clear surface mesh model from complex point cloud data.

[0128] Furthermore, the classification results are used to generate a 3D model of the target city's buildings through a surface reconstruction layer, thereby realizing the transformation from disordered point cloud data to an ordered surface mesh.

[0129] In the above implementation process, a corresponding nearest neighbor threshold is set for each category of building screening point cloud according to the level of the point cloud. The corresponding building screening point cloud is then encoded using the nearest neighbor threshold to obtain the encoded building features. This results in the encoded building features having richer local geometric information, which in turn improves the accuracy of the target city's 3D building model generated by further reconstructing the encoded building features using the feature reconstruction unit.

[0130] In one embodiment, performing feature reconstruction on all encoded building features based on the feature reconstruction unit to generate a 3D model of buildings in the target city may include the following steps:

[0131] Step 1: Enhance the feature representation capability of the corresponding encoded building features to obtain the corresponding enhanced building features;

[0132] Step 2: Based on the feature reconstruction unit, perform feature reconstruction on all enhanced building features to generate a 3D model of the buildings in the target city.

[0133] For example, the feature representation capabilities of the encoded building features of the concave and convex parts, the encoded building features of the angular parts, and the encoded building features of the inflection point parts are enhanced, thereby improving the accuracy of building point cloud feature extraction and thus improving the accuracy of building 3D model reconstruction.

[0134] Furthermore, feature reconstruction units are used to reconstruct all enhanced building features to generate a 3D model of the target city's buildings.

[0135] In the above implementation process, the corresponding encoded building features are encoded according to the nearest neighbor threshold to obtain the corresponding encoded building features. The feature representation capability of each category of building point cloud is enhanced, which improves the accuracy of building point cloud feature extraction and facilitates further improvement of the accuracy of building model reconstruction.

[0136] In one embodiment, the building segmentation model further includes a decoding module, which performs image segmentation on the orthophoto based on the building segmentation model to obtain a building segmentation mask for the target city. This may include the following steps:

[0137] Step 1: Input the orthophoto into the first coding unit to obtain the first coding result.

[0138] Step 2: Input the orthophoto into the second coding unit to obtain the second coding result.

[0139] Step 3: Merge the first encoding result and the second encoding result to obtain the merged feature information.

[0140] Step 4: Decode the merged feature information in the decoding module to obtain the building segmentation mask in the orthophoto.

[0141] For example, when inputting an orthophoto into a building segmentation model, the orthophoto can be input into a first encoding unit to obtain a first encoding result. The first encoding result is used to describe the building segmentation features of a first precision in the orthophoto. The orthophoto can be input into a second encoding unit to obtain a second encoding result. The first encoding result can be used to describe the building segmentation features of a second precision in the orthophoto. Further, the first encoding result and the second encoding result are merged to obtain merged feature information. Specifically, the Merge module can be used to merge the first encoding result and the second encoding result. The merged feature information is used to describe the building segmentation features in the orthophoto after the first encoding result and the second encoding result are merged.

[0142] The building segmentation model may also include a decoding module, which can then input the merged feature information into the decoding module for decoding, thereby outputting the feature information describing the building segmentation in the form of an image, i.e., obtaining the building segmentation mask in the orthophoto.

[0143] In the above implementation process, orthophotos are input into the first encoding unit and the second encoding unit respectively, resulting in first and second encoding results describing building segmentation features of different precisions. The first and second encoding results are then merged to obtain merged feature information, which includes building segmentation features of two different precisions. Further, the merged feature information is input into a value decoding module for decoding, thereby causing the building segmentation model to output a building segmentation mask.

[0144] In one embodiment, the second encoding unit includes a second feature extraction network and a second feature processing network; inputting the orthophoto into the second encoding unit to obtain the second encoding result may include the following steps:

[0145] Step a: Input the orthophoto into the second feature extraction network to obtain the second feature information.

[0146] Step b: Input the second feature information into the second feature processing network to obtain the second encoding result.

[0147] For example, the second encoding unit may include a second feature extraction network and a second feature processing network.

[0148] Specifically, the second feature extraction network can be a deep convolutional neural network (DCNN). The orthophoto is input into the second feature extraction network for feature extraction, thereby obtaining the second feature information. The second feature processing network can be an atrous spatial pyramid pooling module (ASPP). The second feature information is input into the second feature processing network for atrous convolution, thereby obtaining the second encoding result.

[0149] Specifically, the ASPP module is a key component for improving the receptive field and multi-scale feature extraction capabilities of Convolutional Neural Networks (CNNs) in semantic segmentation tasks. It primarily captures multi-scale contextual information in images by using atrous convolutions with varying dilation rates. The main function of the ASPP module is to expand the receptive field, extract multi-scale features, and integrate global contextual information through atrous convolutions and global average pooling. These characteristics enable ASPP to more accurately segment targets of different scales in semantic segmentation tasks, improving the model's segmentation accuracy.

[0150] In the above implementation process, the orthophoto is input into the second feature extraction network for feature extraction to obtain the second feature information. Furthermore, the second feature information is input into the second feature processing network for dilated convolution to obtain the second encoding result, thus realizing the feature extraction and feature processing of buildings in the orthophoto.

[0151] In one embodiment, the second feature processing network includes a second standard convolutional layer, four second feature convolutional layers with different convolution rates, a second global pooling layer, and a second channel convolutional layer.

[0152] Step b: Input the second feature information into the second feature processing network to obtain the second encoding result, which may include the following steps:

[0153] Step b1: Input the second feature information into a second standard convolutional layer, four second feature convolutional layers with different convolution rates, and a second global pooling layer to obtain the corresponding second output results.

[0154] Step b2: Perform feature fusion on all second output results to obtain the second fused features.

[0155] Step b3: Input the second fused feature into a second-channel convolutional layer for channel dimensionality reduction to obtain the second encoding result.

[0156] For example, the second feature processing network may include a second standard convolutional layer, four second feature convolutional layers with different convolution rates, a second global pooling layer, and a second channel convolutional layer. The second standard convolutional layer may be a 1*1 dilated convolution. The kernel sizes of the four second feature convolutional layers with different convolution rates may be the same or different. In this embodiment, taking the kernel size of the four second feature convolutional layers with different convolution rates as an example, for instance, the four second feature convolutional layers with different convolution rates may all be dilated convolutions, the kernel size may all be 5*5, and the convolution rates may be rate3, rate6, rate12, and rate18, respectively. The second channel convolutional layer may be a 1*1 standard convolution.

[0157] Specifically, the second extracted features are input into a 1*1 second standard convolutional layer with dilated convolution, a 5*5 kernel, and convolution rates of rate3, rate6, rate12, and rate18, respectively. This is followed by a second feature convolutional layer with dilated convolution and a second global pooling layer with dilated convolution, yielding corresponding second output results. Further, all second output results are fused to obtain second fused features. This can be achieved using the Merge module. The second fused features are then input into a 1*1 second channel convolutional layer for channel dimensionality reduction, reducing the number of channels to increase the image size, thus obtaining the second encoding result.

[0158] In the above implementation process, the second extracted features are convolved and pooled through a second standard convolutional layer, four second feature convolutional layers with different convolution rates, and a second global pooling layer, thereby realizing the feature processing of the second extracted features. All the processed second output results are merged, and then the merged second fused features are channel-reduced through a second channel convolutional layer, thereby realizing the feature processing of the second feature information and finally obtaining the second encoding result with second precision.

[0159] In one embodiment, the first encoding unit includes a first feature extraction network and a first feature processing network; inputting an orthophoto into the first encoding unit to obtain a first encoding result may include the following steps:

[0160] Step A: Input the orthophoto into the first feature extraction network to obtain the first feature information.

[0161] Step B: Process the first feature information based on the first feature processing network to obtain the first encoding result.

[0162] For example, the first encoding unit may include a first feature extraction network and a first feature processing network.

[0163] Specifically, the first feature extraction network can be a DCNN network. The orthophoto is input into the first feature extraction network for feature extraction, thereby obtaining the first feature information. The first feature processing network can be an ASPP with dilated convolutions. The first feature information is input into the first feature processing network for dilated convolutions, thereby obtaining the first encoding result. The kernel size of the first feature processing network is smaller than the kernel size of the second feature processing network, thus making the accuracy of the first feature information greater than the accuracy of the second feature information.

[0164] In the above implementation process, the first feature information is obtained by inputting the orthophoto into the first feature extraction network for feature extraction. Furthermore, the first feature information is input into the first feature processing network for dilated convolution to obtain the first encoding result, thereby realizing feature extraction and feature processing of building outline information in the orthophoto.

[0165] In one embodiment, the first feature processing network includes a first standard convolutional layer, four first feature convolutional layers with different convolution rates, a first global pooling layer, and a first channel convolutional layer.

[0166] The second coding unit includes a second feature processing network, which includes four second feature convolutional layers with different convolution rates; the kernel size of each of the four first feature convolutional layers with different convolution rates is smaller than the kernel size of the four second feature convolutional layers with different convolution rates.

[0167] Step B: Process the first feature information based on the first feature processing network to obtain the first encoding result, which may include the following steps:

[0168] Step B1: Input the first feature information into a first standard convolutional layer, four first feature convolutional layers with different convolution rates, and a first global pooling layer to obtain the corresponding first output results.

[0169] Step B2: Perform feature fusion on all the first output results to obtain the first fused feature.

[0170] Step B3: Input the first fused feature into a first-channel convolutional layer for channel dimensionality reduction to obtain the first encoding result.

[0171] Then, the first encoding result and the second encoding result are merged to obtain merged feature information, which may include: merging the first encoding result, the second encoding result and the first feature information to obtain merged feature information.

[0172] For example, Figure 6 This is a structural schematic diagram of a building segmentation model provided in an embodiment of this application, such as... Figure 6 The network structure of the building segmentation model shown includes a first coding unit and a second coding unit. The first coding unit includes a first feature extraction network and a first feature processing network, and the second coding unit may include a second feature extraction network and a second feature processing network.

[0173] It should be noted that the first feature extraction network and the second feature extraction network can be a single DCNN network or two DCNN networks with the same structure. In this embodiment, in order to clearly represent the data flow, the first feature extraction network and the second feature extraction network in the building segmentation model of this embodiment are set as two DCNN networks with the same structure, but this application is not limited to this.

[0174] The first feature processing network may include a first standard convolutional layer, four first feature convolutional layers with different convolution rates, a first global pooling layer, and a first channel convolutional layer. The first standard convolutional layer may be a 1*1 dilated convolution. The kernel sizes of the four first feature convolutional layers with different convolution rates may be the same or different. In this embodiment, taking the kernel size of the four first feature convolutional layers with different convolution rates as an example, for example, the four first feature convolutional layers with different convolution rates may all be dilated convolutional layers, the kernel size may all be 3*3, and the convolution rates may be rate3, rate6, rate12, and rate18, respectively. The first channel convolutional layer may also be a 1*1 standard convolution.

[0175] The kernel sizes of the four first-feature convolutional layers with different convolutional rates are all smaller than the kernel sizes of the four second-feature convolutional layers with different convolutional rates. For example, if the kernel size of the four first-feature convolutional layers with different convolutional rates in the first-feature processing network is 3*3, the kernel size of the four second-feature convolutional layers with different convolutional rates in the second-feature processing network can all be 5*5; if the kernel size of the four first-feature convolutional layers with different convolutional rates in the first-feature processing network is 5*5, the kernel size of the four second-feature convolutional layers with different convolutional rates in the second-feature processing network can all be 7*7, without any restrictions. In deep learning, smaller convolutional kernels (such as 3*3) generally have fewer parameters and less computation than larger convolutional kernels (such as 5*5 or larger). This helps to speed up model training while reducing the demand for memory and computing resources. However, a small convolution kernel can limit the expressive power of the model to some extent. Therefore, when processing orthophotos using a building segmentation model that includes two sizes of convolution kernels, not only can the building features obtained by the small convolution kernel be preserved, but the orthophotos can also be processed by the large convolution kernel, which improves the expressive power of the building segmentation model and further increases the accuracy of building segmentation in orthophotos.

[0176] Specifically, the first feature information is input into a 1*1 first standard convolutional layer with dilated convolution, a first feature convolutional layer with dilated convolution and a first global pooling layer with dilated convolution, all with a kernel size of 3*3 and convolution rates of rate3, rate6, rate12 and rate18, respectively, to obtain the corresponding first output results.

[0177] Furthermore, all the first output results are fused to obtain the first fused feature. As an example, the Merge module can be used to fuse all the first output results to obtain the first fused feature. Then, the first fused feature is input into a 1*1 first channel convolutional layer for channel dimensionality reduction, which reduces the number of channels to increase the image size, thereby obtaining the first encoding result.

[0178] Furthermore, in the decoding module, the first encoding result, the second encoding result, and the first feature information are merged to obtain merged feature information.

[0179] In the above implementation process, the first feature information is convolved and processed through a first standard convolutional layer, four first feature convolutional layers with different convolutional rates, and a first global pooling layer, thereby realizing the feature processing of the first feature information. Then, all the first output results after feature processing are fused to obtain the first fused feature. The first fused feature is then passed through a first channel convolutional layer for channel dimensionality reduction to obtain the first encoding result. Finally, the first feature information, the first encoding result, and the second encoding result are merged in the decoding module to obtain the merged feature information. This allows the decoding module to output a building segmentation mask based on the merged feature information. As a result, the output building segmentation mask includes not only the feature information obtained after processing by two encoding units with different convolutional kernel sizes, but also the first feature information, effectively improving the accuracy of the output building segmentation mask.

[0180] In one embodiment, the decoding module further includes a channel dimensionality reduction convolutional layer, a first upsampling unit, and a second upsampling unit.

[0181] The first encoding result, the second encoding result, and the first feature information are merged to obtain the merged feature information, which may include the following steps:

[0182] Step S1: Input the first encoding result into the first upsampling unit for upsampling to obtain the first upsampling information.

[0183] Step S2: Input the second encoding result into the second upsampling unit for upsampling to obtain the second upsampling information.

[0184] Step S3: Input the first feature information into the channel dimensionality reduction convolutional layer for channel dimensionality reduction to obtain the first feature information after dimensionality reduction.

[0185] Step S4: Merge the first upsampled information, the second upsampled information, and the dimensionality-reduced first feature information to obtain merged feature information.

[0186] For example, such as Figure 6 As shown, the decoding module may further include a channel-reduction convolutional layer, a first upsampling unit, and a second upsampling unit. The channel-reduction convolutional layer can be a standard 1x1 convolutional layer. Both the first and second upsampling units can be 4x upsampling units. The first encoding result is input into the first upsampling unit for upsampling to obtain first upsampling information; the second encoding result is input into the second upsampling unit for upsampling to obtain second upsampling information, thereby adjusting the number of channels in the first and second encoding information to be consistent. Furthermore, the first feature information is input into the channel-reduction convolutional layer for channel-reduction to obtain the dimensionality-reduced first feature information.

[0187] It should be noted that the first upsampling unit and the second upsampling unit in this embodiment can be the same upsampling unit. In this embodiment, in order to improve the running efficiency of the model, two identical first upsampling units and second upsampling units are set in the building segmentation model to upsample the output results of the corresponding channel convolutional layers.

[0188] Finally, the first upsampled information, the second upsampled information, and the dimensionality-reduced first feature information are merged to obtain the merged feature information, thus increasing the feature dimension of the merged feature information. Specifically, the concat module can be used to merge features to obtain the merged feature information.

[0189] In the above implementation process, the first encoding result and the second encoding result are upsampled respectively to adjust the channels of the first encoding result and the second encoding result to be consistent. Furthermore, the first feature information is reduced in dimensionality through a channel dimensionality reduction convolutional layer to avoid the influence of too many channels of the first feature information on the first encoding result and the second encoding result. Then, the first upsampled information, the second upsampled information, and the dimensionality-reduced first feature information are merged to increase the resolution and feature dimension of the merged feature information.

[0190] In one embodiment, the decoding module includes a channel-reversal convolutional layer and a third upsampling unit. Decoding the merged feature information in the decoding module to obtain the building segmentation mask in the orthophoto image may include the following steps:

[0191] Step 1: Input the merged feature information into the channel-reduction convolutional layer to perform channel reduction, and obtain the merged feature information after convolution.

[0192] Step 2: Input the merged feature information after convolution into the third upsampling unit for upsampling to obtain the building segmentation mask in the orthophoto.

[0193] For example, such as Figure 6 The decoding module in the building segmentation model shown also includes a channel decomposition convolutional layer and a third upsampling unit. The channel decomposition convolutional layer can be a 3x3 convolutional layer used to restore the number of channels for merging feature information, and the third upsampling unit can be a 4x upsampling unit.

[0194] In the decoding module, the channel restoration convolution layer is used to restore the input channel of the merged feature information to obtain the merged feature information after convolution.

[0195] Finally, the merged feature information after convolution is input into the third upsampling unit for upsampling to obtain the building segmentation mask in the orthophoto, so that the size of the building segmentation mask is consistent with the size of the orthophoto.

[0196] In the above implementation process, the channel restoration of the merged feature information is realized through the channel restoration convolutional layer of the decoding module. The merged feature information after convolution is further upsampled to obtain a building segmentation mask with the same size as the orthophoto image.

[0197] Figure 7 This is a schematic diagram of the structure of a building model generation device provided in an embodiment of this application, as shown below. Figure 7 The building model generation device shown includes:

[0198] The acquisition module 701 is used to acquire orthophotos of the target city and point cloud data of the target city.

[0199] Image segmentation module 702 is used to segment orthophotos based on a building segmentation model to obtain building segmentation masks for the target city.

[0200] The matching module 703 is used to match the building segmentation mask with the point cloud data of the target city to obtain the fused data of the building mask and point cloud in the target city;

[0201] The point cloud determination module 704 is used to determine the overall point cloud data of buildings in the target city based on the fusion data of building masks and point clouds in the target city;

[0202] The model reconstruction module 705 is used to classify, filter, and reconstruct the overall point cloud data of buildings in the target city based on the surface reconstruction algorithm, and generate a 3D model of the buildings in the target city.

[0203] In one embodiment, the matching module 703 is specifically used for:

[0204] Multiple mask key points are selected from the building segmentation mask, and the position coordinates of each mask key point are determined;

[0205] The point in the point cloud data of the target city that corresponds to the location coordinates of the target mask key point is determined as the point cloud key point that matches the target mask key point; the target mask key point is any one of multiple mask key points;

[0206] The fusion data of building masks and point clouds in the target city is determined based on the point cloud key points that match each mask key point.

[0207] In one embodiment, the point cloud determination module 704 is specifically used for:

[0208] The building boundary contours in the target city are determined based on the fusion data of building masks and point clouds in the target city.

[0209] Starting from the target point cloud data, a target ray corresponding to the target point cloud data is formed in a specified direction. The target point cloud data is any point cloud data in the point cloud data of the target city.

[0210] The number of intersections between the target ray and the building boundary contour determines whether the target point cloud data is within the building boundary contour;

[0211] All point cloud data within the building boundary contours are identified as the overall point cloud data of buildings in the target city.

[0212] In one embodiment, the point cloud determination module 704 is specifically used for:

[0213] Determine any point cloud in the fused data of building mask and point cloud as the target point cloud of the building outline;

[0214] Point cloud data in the target city whose distance to the target point cloud of building outlines is within a preset distance threshold are identified as candidate point clouds corresponding to the target point cloud of building outlines.

[0215] The candidate point clouds corresponding to the fused data of all building masks and point clouds, as well as all point cloud data within the building boundary contours, are determined as the overall point cloud data of buildings in the target city.

[0216] In one embodiment, the network structure of the surface reconstruction algorithm includes a feature classification unit, a feature selection unit, and a feature reconstruction unit. The model reconstruction module 705 is specifically used for:

[0217] Based on the feature classification unit, the overall point cloud data of buildings in the target city is classified to obtain classified point cloud data, which includes point cloud of building surface convexity, point cloud of building surface corner, and point cloud of building surface inflection point.

[0218] Based on the feature filtering unit, key feature point cloud filtering is performed on the point cloud data after target category classification to obtain key point cloud data of target category in the building. The point cloud data after target category classification is any one of the following: building surface convex point cloud, building surface corner point cloud, and building surface inflection point point cloud.

[0219] Based on feature reconstruction units, point clouds of buildings of all categories are reconstructed to generate 3D models of buildings in the target city.

[0220] In one embodiment, the model reconstruction module 705 is specifically used for:

[0221] Determine the point cloud level for each category of building selection point cloud;

[0222] The threshold for filtering the nearest neighbors of each building category is set based on the level of the point cloud.

[0223] The building selection point cloud of each category is encoded based on the nearest neighbor threshold corresponding to the building selection point cloud of each category to obtain the corresponding encoded building features;

[0224] Based on the feature reconstruction unit, feature reconstruction is performed on all encoded building features to generate a 3D model of the buildings in the target city.

[0225] In one embodiment, the model reconstruction module 705 is specifically used for:

[0226] The feature representation capability of the corresponding encoded building features is enhanced to obtain the corresponding enhanced building features;

[0227] Based on the feature reconstruction unit, feature reconstruction is performed on all enhanced building features to generate a 3D model of the buildings in the target city.

[0228] Please refer to Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. An electronic device 800 provided in this application includes a processor 801 and a memory 802. These components are interconnected and communicate with each other via a communication bus 803 and / or other forms of connection mechanisms (not shown). The memory 802 stores a computer program executable by the processor 801. When executed by the processor 801, the computer program performs the building model generation method described in the first aspect above.

[0229] This application also provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor 801, perform the building model generation method described in the first aspect above.

[0230] The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0231] It should be understood that the disclosed apparatus / systems and methods can also be implemented in other ways, as provided in the embodiments of this application. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0232] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0233] The above description is only an optional implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application.

Claims

1. A method for generating building models, characterized in that, The method includes: Acquire orthophotos of the target city and point cloud data of the target city; The orthophoto is segmented based on a building segmentation model to obtain the building segmentation mask for the target city. The building segmentation mask is matched with the point cloud data of the target city to obtain the fused data of the building mask and point cloud in the target city; The overall point cloud data of buildings in the target city is determined based on the fusion data of building masks and point clouds in the target city; Based on the surface reconstruction algorithm, the overall point cloud data of buildings in the target city is classified, filtered, and reconstructed to generate a 3D model of the buildings in the target city.

2. The method according to claim 1, characterized in that, The step of matching the building segmentation mask with the point cloud data of the target city to obtain fused data of the building mask and point cloud in the target city includes: Multiple key points are selected from the building segmentation mask, and the position coordinates of each key point are determined. The point in the point cloud data of the target city that corresponds to the position coordinates of the target mask key point is determined as the point cloud key point that matches the target mask key point; the target mask key point is any one of multiple mask key points; The fused data of building masks and point clouds in the target city is determined based on the point cloud key points that match each mask key point.

3. The method according to claim 1, characterized in that, The determination of the overall point cloud data of buildings in the target city based on the fusion data of building masks and point clouds in the target city includes: The building boundary contours in the target city are determined based on the fusion data of building masks and point clouds in the target city. Starting from the target point cloud data, a target ray corresponding to the target point cloud data is formed in a specified direction, wherein the target point cloud data is any point cloud data in the point cloud data of the target city; Whether the target point cloud data is within the building boundary contour is determined based on the number of intersections between the target ray and the building boundary contour. All point cloud data within the boundary contour of the building are identified as the overall point cloud data of the buildings in the target city.

4. The method according to claim 3, characterized in that, The step of determining all point cloud data within the boundary contour of the building as the overall point cloud data of the buildings in the target city includes: Any point cloud in the fused data of the building mask and point cloud is determined as the target point cloud of the building outline; Point cloud data in the point cloud data of the target city whose distance to the target point cloud of the building outline is within a preset distance threshold are determined as candidate point clouds corresponding to the target point cloud of the building outline; The candidate point clouds corresponding to the fused data of all the building masks and point clouds, as well as all the point cloud data within the boundary contours of the buildings, are determined as the overall point cloud data of the buildings in the target city.

5. The method according to claim 1, characterized in that, The network structure of the surface reconstruction algorithm includes a feature classification unit, a feature filtering unit, and a feature reconstruction unit. The process of classifying, filtering, and reconstructing the overall point cloud data of buildings in the target city based on the surface reconstruction algorithm to generate a 3D model of the buildings in the target city includes: Based on the feature classification unit, the overall point cloud data of buildings in the target city is classified to obtain classified point cloud data, which includes point cloud of building surface convexity, point cloud of building surface corner, and point cloud of building surface inflection point. Based on the feature filtering unit, key feature point cloud filtering is performed on the point cloud data after target category classification to obtain key point cloud data of target category in the building. The point cloud data after target category classification is any one of the following: building surface convex point cloud, building surface corner point cloud, and building surface inflection point point cloud. Based on the feature reconstruction unit, the point clouds of all types of buildings are reconstructed to generate a 3D model of the buildings in the target city.

6. The method according to claim 5, characterized in that, The process of reconstructing the point clouds of all categories of buildings based on the feature reconstruction unit to generate a 3D model of the buildings in the target city includes: Determine the point cloud level for each category of building selection point cloud; Based on the level of the point cloud, a nearest neighbor threshold is set for each category of building point cloud. The building selection point cloud of each category is encoded based on the nearest neighbor threshold corresponding to the building selection point cloud of each category to obtain the corresponding encoded building features; Based on the feature reconstruction unit, feature reconstruction is performed on all encoded building features to generate a 3D model of the buildings in the target city.

7. The method according to claim 6, characterized in that, The step of reconstructing features from all encoded building features using the feature reconstruction unit to generate a 3D model of the buildings in the target city includes: The feature representation capability of the corresponding encoded building features is enhanced to obtain the corresponding enhanced building features; Based on the feature reconstruction unit, feature reconstruction is performed on all enhanced building features to generate a three-dimensional model of the buildings in the target city.

8. A building model generation device, characterized in that, The device includes: The acquisition module is used to acquire orthophotos of the target city and point cloud data of the target city; The image segmentation module is used to segment the orthophoto based on the building segmentation model to obtain the building segmentation mask of the target city. The matching module is used to match the building segmentation mask with the point cloud data of the target city to obtain the fused data of the building mask and point cloud in the target city; The point cloud determination module is used to determine the overall point cloud data of buildings in the target city based on the fusion data of building masks and point clouds in the target city; The model reconstruction module is used to classify, filter, and reconstruct the overall point cloud data of buildings in the target city based on the surface reconstruction algorithm, and generate a three-dimensional model of the buildings in the target city.

9. An electronic device, characterized in that, The electronic device includes: Memory; processor; The memory stores a computer program executable by the processor, which, when executed by the processor, performs the building model generation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, perform the building model generation method according to any one of claims 1-7.