Monocular remote sensing image building roof model construction method based on point-line-plane prediction
By using a monocular remote sensing image method based on point, line, and surface prediction, combined with feature fusion and topology correction, the problem of insufficient detail and topological consistency of 3D models in complex building structures in existing technologies is solved, and high-precision LoD-2 model construction is achieved.
Patent Information
- Application Number
- CN202510964808.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-07-14
AI Technical Summary
Existing methods for building 3D architectural models rely on a single data source, which results in insufficient detail when dealing with complex architectural structures. This leads to defects in geometric accuracy and topological consistency, making it difficult to generate detailed LoD-2 models.
A method for constructing building roof models based on monocular remote sensing images using point, line, and surface prediction is adopted. By acquiring monocular remote sensing images and initial depth estimation maps, point, line, and surface segmentation maps are extracted using a feature fusion model, and topological errors are corrected by combining a topology correction library, and finally a 3D model is constructed.
It improves the geometric accuracy, detail representation, and topological consistency of 3D models, generating high-quality LoD-2 models suitable for generating 3D models of rooftops for complex building structures.
Smart Images

Figure CN120852664B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of real three-dimensional technology, more particularly to a monocular remote sensing image building roof model construction method based on point-line-surface prediction. BACKGROUND
[0002] With the rapid advancement of smart city construction and urbanization in China, the economic and social development needs spatial information to move from two-dimensional to three-dimensional, and people urgently need to express more realistic geographic real scene space through high-precision and fine-grained modeling methods. Buildings, as the main ground category in urban scenes, the construction of large-scale three-dimensional real scene models is the data basis for urban land resource management and various spatial analysis applications.
[0003] Traditional three-dimensional building model construction methods usually rely on a single data source, such as laser point cloud. Although these methods can provide building geometry information to some extent, they have limitations in dealing with complex building structures. For example, a single data source may not provide enough details to construct a detailed LoD-2 model, resulting in defects in model geometry accuracy and topological consistency.
[0004] In view of the above problems, the present application is proposed. SUMMARY
[0005] The present application is proposed in view of the above problems. According to one aspect of the present application, a monocular remote sensing image building roof model construction method based on point-line-surface prediction is provided, comprising:
[0006] Obtaining a monocular remote sensing image containing a target building roof and an initial depth estimation map;
[0007] Respectively extracting image features of the monocular remote sensing image and depth features of the initial depth estimation map;
[0008] Inputting the image features and the depth features into a feature fusion model to obtain a point-line-surface segmentation map, the point-line-surface segmentation map including roof key points, ridge lines, roof contour lines and a two-dimensional segmentation map of the target building roof;
[0009] Using a topological correction library to correct point, line and surface topological errors in the point-line-surface segmentation map to obtain a corrected image;
[0010] Assigning depth values to points on the corrected image using the initial depth estimation map to obtain a three-dimensional point cloud;
[0011] Based on the three-dimensional point cloud and the building area vector contour of the target building roof, constructing a three-dimensional model of the target building roof.
[0012] Exemplarily, the topology correction library comprises a plurality of correction entries respectively corresponding to different point-line-surface error modes, each of the correction entries comprising an error subgraph and a correct subgraph; the correction of the point-line-surface topology errors in the point-line-surface segmentation graph by using the topology correction rules in the topology correction library comprises:
[0013] converting the point-line-surface segmentation graph into a topology graph;
[0014] correcting the topology graph by using the correct subgraph of at least part of the correction entries in the plurality of correction entries based on the matching degree of the error subgraph of each of the correction entries in the plurality of correction entries with the topology graph;
[0015] wherein the at least part of the correction entries are the correction entries whose error subgraphs match the topology graph to a preset degree.
[0016] Exemplarily, the correction of the topology graph by using the correct subgraph of at least part of the correction entries in the plurality of correction entries based on the matching degree of the error subgraph of each of the correction entries in the plurality of correction entries with the topology graph comprises:
[0017] repeating the following steps until the matching degree of the error subgraph of any of the correction entries in the plurality of correction entries with the topology graph is lower than a preset degree:
[0018] determining the matching degree of the error subgraph of each of the correction entries in the plurality of correction entries with the topology graph;
[0019] when there is at least one correction entry in the plurality of correction entries whose error subgraph matches the topology graph to a preset degree, replacing the part of the topology graph matching the corresponding error subgraph by using the correct subgraph in the correction entry matching the topology graph to the highest degree.
[0020] Exemplarily, the correction of the topology graph by using the correct subgraph of at least part of the correction entries in the plurality of correction entries based on the matching degree of the error subgraph of each of the correction entries in the plurality of correction entries with the topology graph comprises:
[0021] determining the matching degree of the error subgraph of each of the correction entries in the plurality of correction entries with the topology graph;
[0022] when there is at least one correction entry in the plurality of correction entries whose error subgraph matches the topology graph to a preset degree, replacing the part of the topology graph matching the corresponding error subgraph by using the correct subgraph in each of the at least one correction entry.
[0023] Exemplarily, the determining the matching degree of each of the plurality of correction items with the topology graph comprises:
[0024] For the error subgraph of any one of the plurality of correction items, a graph pattern matching algorithm is used to determine a local graph similar to the error subgraph in the topology graph and a similarity between the local graph and the error subgraph;
[0025] The matching degree is represented by the similarity.
[0026] Exemplarily, the feature fusion model comprises a first feature conversion module, a second feature conversion module and a third feature conversion module; and the inputting the image feature and the depth feature into the feature fusion model to obtain a point-line-surface segmentation graph comprises:
[0027] The depth feature is input into the first feature conversion module;
[0028] The image feature is input into the second feature conversion module;
[0029] The output of the first feature conversion module and the output of the second feature conversion module are spliced in the depth dimension and then input into the third feature conversion module;
[0030] The first feature conversion module and the second feature conversion module share weights; and the point-line-surface segmentation graph is obtained based on the output decoding of the third feature conversion module.
[0031] Exemplarily, before the depth feature is input into the first feature conversion module, the method further comprises:
[0032] The key vector and the value vector of the first feature conversion module and the second feature conversion module are exchanged.
[0033] Exemplarily, the constructing a three-dimensional model of the target building roof based on the three-dimensional point cloud and the building area vector contour of the target building roof comprises:
[0034] The points in the three-dimensional point cloud are processed by using a RANSAC algorithm to obtain an equation of each roof plane of the target building roof;
[0035] Based on the equation of each roof plane of the target building roof, a roof polygon is generated;
[0036] The building area vector contour of the target building roof is lifted to a facade polygon;
[0037] The roof polygon and the facade polygon are grown and collided at a constant speed using a planar assembly dynamics algorithm to obtain a polyhedral cell.
[0038] The portion of the polyhedral cell located outside the building area vector outline of the target building's roof is discarded to obtain a three-dimensional model of the target building's roof.
[0039] According to another aspect of the present invention, an electronic device is provided, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the method as described above.
[0040] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores a computer program / instructions that, when executed by a processor, implement the method described above.
[0041] In the above technical solution, by combining monocular remote sensing imagery, initial depth estimation maps, and building area vector contours, the advantages of multi-source data can be fully utilized, which helps improve the geometric accuracy and detail of the final 3D model. The feature fusion model automatically extracts point, line, and surface segmentation maps based on image and depth features, enabling the automatic extraction of key points on the roof, ridge lines, and roof outlines, which helps improve model generation efficiency and reduce time and labor costs. Simultaneously, by utilizing a topology correction library for topology correction, topological errors between points, lines, and surfaces can be automatically corrected, improving the topological consistency of the 3D model and ensuring the accuracy of the final 3D model. In summary, applying this method to the generation of 3D models of building roofs yields high-quality LoD-2 models with high geometric accuracy, precision, and detail. This method is particularly suitable for generating 3D models of roofs of complex building structures, meeting the needs of high-precision modeling.
[0042] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0043] The above and other objects, features, and advantages of the present invention will become more apparent from the more detailed description of the embodiments of the invention in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same parts or steps.
[0044] Figure 1A schematic flowchart illustrating a method for constructing building roof models from monocular remote sensing images based on point, line, and area prediction according to an embodiment of the present invention is shown.
[0045] Figure 2 A schematic diagram of a feature fusion module according to an embodiment of the present invention is shown;
[0046] Figure 3 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the present invention more apparent, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of the present invention, and not all of the embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described herein, all other embodiments obtained by those skilled in the art without inventive effort should fall within the protection scope of the present invention.
[0048] In the field of real-world 3D imaging, establishing accurate 3D building models is a key technical challenge. However, existing 3D building modeling techniques have shortcomings in data fusion, topology error correction, and model detail processing. As mentioned above, model building methods relying on a single data source cannot provide sufficient detail to construct a detailed LoD-2 model, resulting in defects in geometric accuracy and topological consistency. Some related technologies employ deep learning-based methods to assist in 3D model construction. While this approach can improve modeling efficiency, it is prone to topological errors. For example, the model may incorrectly connect two non-adjacent roof planes or omit certain important ridge lines, leading to insufficient robustness. In short, a method for constructing detailed and accurate LoD-2 models is currently lacking. Therefore, this invention provides a method for constructing building roof models from monocular remote sensing imagery based on point, line, and surface prediction. This method effectively improves model accuracy by fusing multi-source data and utilizing a topology correction library for topology correction, thereby contributing to the generation of detailed LoD-2 models.
[0049] According to one aspect of the present invention, a method for constructing building roof models from monocular remote sensing images based on point, line, and surface prediction is provided. Figure 1 A schematic flowchart illustrating a method for constructing building roof models from monocular remote sensing imagery based on point, line, and area prediction according to an embodiment of the present invention is shown. Figure 1 As shown, the method may include the following steps: S110, S120, S130, S140, S150, and S160.
[0050] In step S110, a monocular remote sensing image containing the roof of the target building and an initial depth estimation map are acquired.
[0051] The initial depth estimation map can be obtained using the Depth-Anything depth estimation model. In some implementations of this embodiment, monocular remote sensing images and pre-trained weights can be input into Depth-Anything to obtain the initial depth estimation map. The pre-trained weights are obtained by training on a large amount of labeled depth data, enabling the model to extract rich feature information from the image and predict the depth value of each pixel based on this feature information; details are omitted here.
[0052] In step S120, the image features of the monocular remote sensing image and the depth features of the initial depth estimation map are extracted respectively.
[0053] Image features from monocular remote sensing images and depth features from initial depth estimation maps can be obtained using any existing or future neural network for feature extraction. In some implementation schemes of this embodiment, ResNet can be used as the backbone network. That is, the monocular remote sensing image and the initial depth estimation map are respectively input into ResNet, and processed using a series of convolutional layers and residual modules in ResNet to obtain image features and depth features. Both features are output in the form of feature maps.
[0054] In step S130, image features and depth features are input into the feature fusion model to obtain a point-line-plane segmentation map, which includes the roof key points, ridge line, roof outline, and two-dimensional segmentation map of the target building's roof.
[0055] The feature fusion model can predict and generate a point-line-polygon segmentation map from image and depth features. This segmentation map identifies key roof points such as vertices and intersections, as well as the ridge line, roof outline, and a two-dimensional segmentation of the roof. This segmentation map provides rich detail for the subsequent generation of the 3D model.
[0056] In step S140, the topology errors of points, lines, and surfaces in the point-line-surface segmentation map are corrected using the topology correction library to obtain a corrected image.
[0057] In this embodiment, multiple topology correction rules (which can be correction entries as described below) can be pre-stored in a topology correction library. Each topology correction rule can correspond to an error mode. Thus, the error modes corresponding to the multiple topology correction rules can cover various topology errors that may occur between points and lines, lines and surfaces, and points and surfaces, such as missing points, missing lines, missing surfaces, incorrect connections between points and lines, discontinuities between lines and surfaces, and incorrect intersections between surfaces.
[0058] Topology correction rules can store potential errors (such as erroneous subgraphs below) and their corrected versions (such as correct subgraphs below). When an error exists in the point-line-plane segmentation diagram, the erroneous part can be replaced using the corrected version from the topology correction rules. This effectively avoids topology errors such as connecting two non-adjacent roof planes or omitting important ridge lines.
[0059] In step S150, depth values are assigned to points on the corrected image using the initial depth estimation map to obtain a three-dimensional point cloud.
[0060] It is understood that the initial depth estimation map stores the depth value of each pixel. In this embodiment, the depth values of keypoints, points on segmentation lines, and points on segmentation planes in the corrected image can all be extracted from the initial depth estimation map. Since the initial depth estimation map is essentially a two-dimensional array, in some implementations, the corresponding pixel coordinates in the map can be directly found based on the vector coordinates. The nearest neighbor interpolation method is used to assign depth values to each point to provide geometric details for subsequent 3D modeling.
[0061] In step S160, a 3D model of the target building's roof is constructed based on the 3D point cloud and the building region vector contour of the target building's roof. This 3D model is a LoD-2 model.
[0062] After obtaining the 3D point cloud, a 3D model can be constructed based on the point cloud. For example, the RANSAC algorithm can be combined with the building region vector contour of the target building's roof to construct a 3D model of the target building's roof.
[0063] The aforementioned technical solution, by combining monocular remote sensing imagery, initial depth estimation maps, and building area vector contours, fully leverages the advantages of multi-source data, which helps improve the geometric accuracy and detail of the final 3D model. Through a feature fusion model, point, line, and surface segmentation maps are automatically extracted based on image and depth features, enabling the automatic extraction of key points on the roof, ridge lines, and roof outlines. This improves model generation efficiency and reduces time and labor costs. Simultaneously, topological correction using a topology correction library automatically corrects topological errors between points, lines, and surfaces, improving the topological consistency of the 3D model and ensuring its accuracy. In summary, applying this method to the generation of 3D models of building roofs yields high-quality LoD-2 models with high geometric accuracy, precision, and detail. This method is particularly suitable for generating 3D models of roofs of complex building structures, meeting the requirements for high-precision modeling.
[0064] For example, the topology correction library includes multiple correction entries corresponding to different point, line, and surface error patterns. Each correction entry includes an erroneous subgraph and a correct subgraph. The topology correction rules in the topology correction library are used to correct the topology errors of points, lines, and surfaces in the point, line, and surface segmentation map, including: converting the point, line, and surface segmentation map into a topology map; and correcting the topology map using the correct subgraphs of at least some of the correction entries based on the degree of matching between the erroneous subgraphs of each correction entry and the topology map. Wherein, at least some of the correction entries are correction entries whose erroneous subgraphs match the topology map to a preset degree.
[0065] In this example, the point-line-polygon segmentation map can be vectorized first to construct a complete topology map, which facilitates matching with erroneous subgraphs in subsequent steps.
[0066] After obtaining the topology graph, it can be matched with erroneous subgraphs in the topology correction library. If a matching erroneous subgraph is found (i.e., the matching degree reaches a preset level), the correct subgraph corresponding to the erroneous subgraph can be used to replace the matching part of the topology graph, thereby correcting the topological errors. For example, for incorrect connections between points and lines, replacing the correct subgraph can adjust the position of the points or redefine the connection relationship of the lines; for discontinuities between lines and surfaces, replacing the correct subgraph can repair the boundary relationship between lines and surfaces, making it continuous and geometrically consistent; for incorrect intersections between surfaces, replacing the correct subgraph can adjust the boundary of the surfaces or redivide the regions of the surfaces, ensuring the topological consistency and accuracy of the model.
[0067] In this embodiment, the degree of matching can be represented by similarity or by the reciprocal of graph distance. Similarly, the preset degree can be represented by a similarity threshold or by the reciprocal of a distance threshold, which will not be elaborated further.
[0068] The above technical solution can quickly and accurately find the location of topological errors in the topology graph by matching the erroneous subgraph of each correction entry with the topology graph. By replacing these locations one by one with the corresponding correct subgraph, the topological errors in the topology graph can be quickly corrected. This helps to provide an accurate basis for the construction of the 3D model in the subsequent process, so that the 3D model can reach an accurate state in terms of geometric structure.
[0069] For example, based on the degree of matching between the erroneous subgraph of each of the multiple correction entries and the topology graph, the topology graph is corrected using the correct subgraphs of at least some of the multiple correction entries, including: repeatedly performing the following steps until the degree of matching between the erroneous subgraph of any of the multiple correction entries and the topology graph is lower than a preset level: determining the degree of matching between the erroneous subgraph of each of the multiple correction entries and the topology graph; when the degree of matching between the erroneous subgraph of at least one of the multiple correction entries and the topology graph reaches a preset level, replacing the part of the topology graph that matches the corresponding erroneous subgraph with the correct subgraph of the correction entry with the highest degree of matching with the topology graph.
[0070] In this example, only the portion of the topology graph that matches the erroneous subgraph is corrected at a time. After each correction, the matching degree between the erroneous subgraph of each correction entry and the topology graph is reassessed, and the portion of the topology graph that matches the erroneous subgraph with the highest degree of matching is corrected again, until the matching degree between the topology graph and all erroneous subgraphs in the topology correction library is lower than the preset degree. This method of correcting only one error in the topology graph at a time avoids mutual interference between the corrections of different errors, ensuring correction efficiency and effectiveness, and thus guaranteeing the topological consistency of the topology graph.
[0071] For example, based on the degree of matching between the erroneous subgraph of each of the multiple correction entries and the topology graph, the topology graph is corrected using the correct subgraphs of at least some of the multiple correction entries, including: determining the degree of matching between the erroneous subgraph of each of the multiple correction entries and the topology graph; when the degree of matching between the erroneous subgraph of at least one of the multiple correction entries and the topology graph reaches a preset level, the part of the topology graph that matches the corresponding erroneous subgraph is replaced by the correct subgraph of each of the at least one correction entry.
[0072] In this example, all erroneous locations in the topology graph can be corrected simultaneously, which helps improve correction efficiency and, consequently, model building efficiency.
[0073] For example, determining the degree of matching between the erroneous subgraph of each of the multiple correction entries and the topology graph includes: for the erroneous subgraph of any of the multiple correction entries, using a graph pattern matching algorithm to determine a local graph in the topology graph that is similar to the erroneous subgraph and the similarity between the local graph and the erroneous subgraph; wherein the degree of matching is represented by similarity.
[0074] Graph pattern matching algorithms include, but are not limited to, subgraph isomorphism algorithms and similarity-based pattern matching algorithms. In a specific implementation, the graph pattern matching algorithm can be a graph distance algorithm. In this embodiment, graph distance is used as a similarity metric. Graph distance is the minimum number of operations required to transform one graph into another. These operations include adding or deleting points, lines, or polygons, or changing the connection relationships between elements. In this embodiment, basic editing operations can be defined: adding points, deleting points, adding lines, deleting lines, adding polygons, deleting polygons, and changing connection relationships (changing the connection relationships between points and lines, lines and polygons, and points and polygons). For each error pattern in the input topology graph and the topology pattern library, the graph distance between them is calculated. Each time a basic editing operation is performed, the graph distance is incremented by 1 until the edited input topology graph and each error pattern in the topology pattern library are completely identical. The smaller the graph distance, the higher the similarity.
[0075] In the above scheme, the graph pattern matching algorithm can be used to accurately determine the local image that matches the erroneous subgraph in the topological graph, which can provide a more accurate basis for topological correction in subsequent steps.
[0076] For example, the feature fusion model includes a first feature conversion module, a second feature conversion module, and a third feature conversion module; inputting image features and depth features into the feature fusion model to obtain a point-line-plane segmentation map includes: inputting depth features into the first feature conversion module; inputting image features into the second feature conversion module; concatenating the outputs of the first feature conversion module and the second feature conversion module in the depth dimension and then inputting them into the third feature conversion module; wherein, the first feature conversion module and the second feature conversion module share weights; the point-line-plane segmentation map is obtained by decoding the output of the third feature conversion module.
[0077] Figure 2 A schematic diagram of a feature fusion module according to an embodiment of the present invention is shown. Figure 2 In this architecture, all feature transformation modules utilize the Transformer class. For example... Figure 2As shown, depth features and image features are input into two independent Transformers, which share weights, forming a two-layer interactive attention model. These two independent Transformers can enhance the spatial information of features within a modality. The enhanced features are concatenated along the depth dimension and then passed through the next-level Transformer to achieve information interaction between modalities. This allows for precise, fine-grained information fusion from a global perspective.
[0078] In this embodiment, the input to the third feature conversion module is obtained by concatenating the outputs of the first and second feature conversion modules in the depth dimension. For example, the shape of the output feature map can be represented as H×W×C, where H is the height, W is the width, and C is the depth (number of channels). Concatenating the two feature maps output by the first and second feature conversion modules in the depth dimension means adding their channel numbers to obtain a new feature map with the shape H×W×(C1+C2), where C1 and C2 are the channel numbers of feature maps F1 and F2 output by the first and second feature conversion modules, respectively. After processing and output by the third feature conversion module, this concatenated feature map can be decoded by the decoder to obtain a point-line-surface segmentation map.
[0079] It is understandable that, in order to improve output accuracy, the feature fusion model can be trained using a triple consisting of sample monocular remote sensing images, sample initial depth estimation maps, and sample point-line-area segmentation maps, which will not be elaborated upon here.
[0080] The above scheme effectively reduces model parameters, lowers computational complexity, and improves computational efficiency by sharing weights between the first and second feature conversion modules. It also allows for better learning of the characteristics of different modalities within the same region, enhancing the model's constraint. By processing depth and image features separately using the first and second feature conversion modules, spatial information enhancement can be applied to both features, improving their respective detail representation. Finally, by processing the features obtained after concatenation in the depth dimension using the third feature conversion module, information interaction between modalities can be achieved, enabling precise, fine-grained information fusion from a global perspective.
[0081] For example, before inputting deep features into the first feature conversion module, the method further includes: swapping the key vectors and value vectors of the first feature conversion module and the second feature conversion module.
[0082] In this example, the query vector, key vector, and value vector can be generated first using the first feature transformation module and the second feature transformation module, respectively. For example, they can be denoted as Q1, K1, V1, and Q2, K2, V2, respectively. Then, the key vector and value vector can be swapped. That is, the first feature transformation module uses Q1, K2, V2, and the second feature transformation module uses Q2, K1, V1.
[0083] In the technical solution of this invention, the inventors considered the rich correlation information between different modalities. To generate a high-quality 3D model, it is necessary to effectively fuse this modal information. The inventors found that in the original attention mechanism, directly using the same K and V vectors for the same modality may cause the model to focus on local features within the modality rather than cross-modal correlations. Therefore, the technical solution in this example, by exchanging the key vector and value vector, facilitates two modules to better learn the characteristics between different modalities in the same region. Specifically, exchanging K and V forces the model to explicitly establish cross-modal matching relationships when calculating attention, compelling the model to learn the features of the two modalities in the same region. Therefore, the two modules not only focus on their own modal information but also obtain richer cross-information through each other's key and value vectors, achieving information complementarity and enhancement. This helps improve the enhancement effect of spatial information during feature processing, thereby helping to improve the accuracy of the final model.
[0084] For example, a 3D model of the roof of a target building is constructed based on a 3D point cloud and the building area vector contour of the roof of the target building. This includes: processing the points in the 3D point cloud using the RANSAC algorithm to obtain the equation of each roof plane of the target building; generating roof polygons based on the equations of each roof plane of the target building; elevating the building area vector contour of the roof of the target building into a facade polygon; using a planar assembly dynamics algorithm to make the roof polygon and the facade polygon grow and collide at a constant speed to obtain polyhedral cells; and discarding the parts of the polyhedral cells located outside the building area vector contour of the roof of the target building to obtain the 3D model of the roof of the target building.
[0085] In this example, the RANSAC algorithm is first used to fit the equations for each roof plane. Then, a line-segmented image is used to extract the roof polygons. Simultaneously, existing planar assembly dynamics algorithms allow the roof and facade polygons to grow at a constant rate. During growth, collisions occur between polygons. When a collision occurs, the polygon stops growing and the 3D space is divided according to the collision boundary, forming polyhedral cells (a polyhedral cell is a closed region in 3D space composed of multiple planar polygon faces. These faces are interconnected, forming the boundary of a polyhedron, dividing the 3D space into internal and external parts).
[0086] In the polyhedral cells divided by polygon growth and collision, some cells are located outside the building. These external cells are unnecessary for constructing the 3D model of the building, so in this embodiment, these points are filtered out and discarded. In some implementations of this invention, the following method can be used to discard the portion of the polyhedral cells located outside the building area vector outline of the target building's roof: For each cell, select a point inside it (e.g., the centroid) and determine whether the point is located within the 3D boundary of the building. Specifically, first determine whether the point is located within the 2D outline of the building (i.e., the building area vector outline). If the point is within the 2D outline, then determine whether its height is within the height range of the building (the coordinates and height of the point are known, i.e., (x, y, z) are known, find the corresponding (x, y, z') in the initial depth estimation map, and compare the values of z and z'). If the height exceeds the building's range, then the point is located outside the building, so the cell is determined to be outside the building. After this step, the remaining cells are the set of polyhedral cells representing the building entity.
[0087] The above method is simple to operate and can quickly and accurately reconstruct the LoD-2 model of the building roof based on the 3D point cloud and the building area vector contour of the target building roof.
[0088] According to another aspect of the present invention, an electronic device is also provided. Figure 3 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. Figure 3 As shown, the electronic device 300 includes a processor 310 and a memory 320. The memory 320 stores a computer program, which the processor 310 executes to implement the method described above.
[0089] According to another aspect of the present invention, a computer-readable storage medium is also provided. The storage medium stores a computer program / instructions that, when executed by a processor, implement the method described above. The storage medium may, for example, include a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0090] Those skilled in the art will readily understand the implementation structure, working principle, and beneficial effects of electronic devices and computer-readable storage media by reading the above methods. For the sake of brevity, further details will not be elaborated here.
[0091] Although exemplary embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above exemplary embodiments are merely illustrative and are not intended to limit the scope of the invention. Various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of the invention. All such changes and modifications are intended to be included within the scope of the invention as claimed in the appended claims.
[0092] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0093] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed.
[0094] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0095] Similarly, it should be understood that, in order to streamline the invention and aid in understanding one or more of the various aspects of the invention, features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the description of exemplary embodiments of the invention. However, this approach should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, its inventive point lies in solving the corresponding technical problem with fewer features than all of those in a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.
[0096] Those skilled in the art will understand that, apart from the mutual exclusion of features, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or elements of any method or apparatus so disclosed may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0097] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.
[0098] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules in the electronic device according to embodiments of the present invention. The present invention can also be implemented as an apparatus program (e.g., a computer program and computer program product) for performing some or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0099] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0100] The above description is merely a specific embodiment of the present invention or an explanation of that embodiment. The scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for constructing a building roof model of a monocular remote sensing image based on point-line-plane prediction, characterized in that, The method comprises: acquiring a monocular remote sensing image containing a target building roof and an initial depth estimation map; extracting image features of the monocular remote sensing image and depth features of the initial depth estimation map, respectively; inputting the image features and the depth features into a feature fusion model to obtain a point-line-surface segmentation map, wherein the point-line-surface segmentation map comprises roof key points, ridge lines, roof contour lines and a two-dimensional segmentation map of the target building roof; correcting point-line-surface topological errors in the point-line-surface segmentation map by using a topological correction library to obtain a corrected image; assigning depth values to points on the corrected image by using the initial depth estimation map to obtain a three-dimensional point cloud; constructing a three-dimensional model of the target building roof based on the three-dimensional point cloud and a building area vector contour of the target building roof; the construction of the three-dimensional model of the target building roof based on the three-dimensional point cloud and the building area vector contour of the target building roof comprises: processing points in the three-dimensional point cloud by using a RANSAC algorithm to obtain an equation of each roof plane of the target building roof; generating roof polygons based on the equation of each roof plane of the target building roof; promoting the building area vector contour of the target building roof to an elevation polygon; growing the roof polygons and the elevation polygon at a constant speed by using a plane assembly dynamics algorithm to obtain a polyhedral cell; discarding parts of the polyhedral cell located outside the building area vector contour of the target building roof to obtain a three-dimensional model of the target building roof.
2. The method of claim 1, wherein, The topological correction library comprises a plurality of correction entries corresponding to different point-line-surface error modes, and each correction entry comprises an error subgraph and a correct subgraph. The correction of point-line-surface topological errors in the point-line-surface segmentation map by using topological correction rules in the topological correction library comprises: converting the point-line-surface segmentation map into a topological graph; correcting the topological graph by using correct subgraphs of at least some of the correction entries based on the matching degree of the error subgraph of each correction entry in the plurality of correction entries and the topological graph; wherein the at least some correction entries are correction entries whose error subgraphs match the topological graph to a preset degree.
3. The method of claim 2, wherein, The correction of the topological graph by using correct subgraphs of at least some of the correction entries based on the matching degree of the error subgraph of each correction entry in the plurality of correction entries and the topological graph comprises: repeating the following steps until the matching degree of the error subgraph of any correction entry in the plurality of correction entries and the topological graph is lower than a preset degree: determining the matching degree of the error subgraph of each correction entry in the plurality of correction entries and the topological graph; when the matching degree of the error subgraph of at least one correction entry in the plurality of correction entries and the topological graph reaches a preset degree, replacing the part of the topological graph that matches the corresponding error subgraph with the correct subgraph in the correction entry that matches the topological graph to the highest degree.
4. The method of claim 2, wherein, The correcting the topology graph by using the correct sub-graphs of at least some of the plurality of correction entries based on the matching degrees of the error sub-graphs of each of the plurality of correction entries and the topology graph comprises: determining the matching degrees of the error sub-graphs of each of the plurality of correction entries and the topology graph; when the matching degrees of the error sub-graphs of at least one of the plurality of correction entries and the topology graph reach a preset degree, replacing the part of the topology graph that matches the corresponding error sub-graph by using the correct sub-graphs of each of the at least one correction entry.
5. The method according to claim 3 or 4, characterized in that, The determining the matching degrees of the error sub-graphs of each of the plurality of correction entries and the topology graph comprises: for the error sub-graph of any one of the plurality of correction entries, determining a local graph in the topology graph that is similar to the error sub-graph and a similarity between the local graph and the error sub-graph by using a graph pattern matching algorithm; wherein the matching degree is represented by the similarity.
6. The method of claim 2, wherein, The feature fusion model comprises a first feature conversion module, a second feature conversion module, and a third feature conversion module; The inputting the image feature and the depth feature into a feature fusion model to obtain a point-line-surface segmentation map comprises: inputting the depth feature into the first feature conversion module; inputting the image feature into the second feature conversion module; concatenating the output of the first feature conversion module and the output of the second feature conversion module in the depth dimension and inputting the concatenated output into the third feature conversion module; wherein the first feature conversion module and the second feature conversion module share weights; and the point-line-surface segmentation map is obtained by decoding the output of the third feature conversion module.
7. The method of claim 6, wherein, Before inputting the depth feature into the first feature conversion module, the method further comprises: swapping the key vector and the value vector of the first feature conversion module and the second feature conversion module.
8. An electronic device, comprising: The device comprises a processor and a memory, and the memory stores a computer program, and the processor is configured to execute the computer program to implement the method of any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The device stores a computer program / instruction, and the computer program / instruction is executed by a processor to implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Building structured model reconstruction method based on laser point cloud and image
CN118279516A
Building three-dimensional reconstruction method, device and equipment based on single remote sensing image
CN118470250A