Systems and methods of predicting three dimensional reconstructions of a building

EP4616372A4Pending Publication Date: 2026-04-22HOVER INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
HOVER INC
Filing Date
2023-11-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Current computer vision and machine learning models are limited in reconstructing three-dimensional representations of buildings from varying perspectives and incomplete data, as they rely on 'bottom-up' approaches that require explicit observations and are prone to noise and occlusions, lacking a top-down holistic understanding of object geometry.

Method used

The method employs a top-down approach using latent codes from a trained distribution of building objects, where latent codes are sampled and updated based on loss computation to infer semantic geometry, combining implicit neural fields and graph neural networks for generating structured 3D representations, even from incomplete views.

Benefits of technology

This approach enables the generation of plausible and complete 3D representations by iteratively minimizing loss between observed data and predicted models, maximizing the probability of outputs aligned with learned priors, resulting in more accurate and continuous predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The present disclosure describes systems, methods, and techniques for predicting plausible, semantic, and structure three-dimensional (3D) representations of a building from one or more images of the building. One or more two-dimensional (2D) images of the same building from different camera perspectives are input and used to predict corresponding 2D semantic representations of the building. Latent codes are iteratively sampled from a learned latent space representing a distribution of building structures and used to infer successive semantic representations based on losses between the predicted representations and the inferred representations until a convergence between the predicted representations and the inferred representations is detected. A resulting latent code is then decoded into a semantic geometry for the building.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS OF PREDICTING THREE DIMENSIONAL RECONSTRUCTIONS OF A BUILDINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 424,482, filed on November 10, 2022, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD

[0002] The present disclosure relates to systems, methods, and techniques for predicting three dimensional reconstructions of a building.BACKGROUND

[0003] Humans possess much implicit knowledge of the world without specifically observing every detail and can make assumptions and educated guesses for unobserved data. For example, humans can “know” what a scene may depict, or what an object may look like, from a different angle without explicitly viewing it in its entirety.

[0004] Computer vision pipelines, and computer-based reconstruction techniques are not so equipped, and without higher order learning must observe data to reproduce it, and even then, imperfect sensors lead to noisy observations. Even in high fidelity sensing there exists ambient interferences such as line of sight occlusions. Machine learning models are similarly limited, for example they are constrained to the limits of their domain; a network trained to identify a particular geometry is ill equipped to identify other forms of geometry.

[0005] In other words, computers can only reconstruct what they observe or predict within a range of what they have been exposed to during training. These may be thought of as “bottom-up” reconstruction techniques. In such bottom-up techniques, specific inputs are observed to form or fit a representation. What is needed is a top-down approach for reconstructing objects with complete parameters defining an object even in the absence of explicit views or varying domain species.BRIEF SUMMARY

[0006] Aspects and features of the present disclosure relate to reconstructing plausible, semantic, and structured 3D representations of a building. One aspect of the present disclosure relates to a method for predicting a semantic geometry of a house. The method may include receiving at least one image depicting the house. The at least one image may observe or depict the house from a first pose. The method may include generating semantic information of the house from the at least one image. The method may include sampling a first latent code for decoding into a representation. The first latent code may be one of a plurality of latent codes from a trained distribution of building objects. The method may include inferring a first representation from the first latent code based on the first pose. The method may include computing a loss between the semantic information and the first inferred representation. The method may include selecting a second latent code from the plurality of latent codes based on the loss; selecting may be choosing a new latent code or providing an update or modification to the latent code’ s position in latent space. The method may include inferring a second representation from the second latent code based on the first pose.

[0007] In some embodiments, each latent code of the plurality of latent codes is an encoded value of one or more parameters for a structural description of building objects. In some embodiments, the encoded value is learned from an implicit neural field. In some embodiments, the one or more parameters includes a roof ridge length. In some embodiments, inferring the first representation is performed by decoding the first latent code with an implicit neural field. Additionally, or alternatively, inferring the first representation is performed by decoding the first latent code with a graph neural network. In some embodiments, inferring the second representation is performed by decoding the second latent code with an implicit neural field. Additionally, or alternatively, inferring the second representation is performed by decoding the second latent code with a generative neural network.

[0008] In some embodiment, computing a loss between the semantic information and the first inferred representation further comprises projecting the first semantic representation into the semantic information of the at least one image. In some embodiments, selecting the second latent code is based on a gradient of the loss with respect to the first inferred representation.

[0009] The method may further comprise computing a second loss between the semantic information and the second inferred representation, selecting a third latent code from the plurality of latent codes based on the second loss, and inferring a third representation from the latent code based on the pose. In some embodiments, computing the second loss between the semantic information and the second inferred representation further comprises projecting the second inferred representation into the semantic information of the at least one image. In some embodiments, selecting the third latent code is based on a gradient of the loss with respect to the second inferred representation. In some embodiments, the semantic information comprises a silhouette of the house from the first pose. In some embodiments, the first or second inferred representation comprises a silhouette of the house from the first pose. In some embodiments, the semantic information comprises at least one architectural feature of the house. In some embodiments, the at least one architectural feature is a linear feature. In some embodiments, the at least one architectural feature is a roof ridge line. In some embodiments, the first or second inferred representation comprises at least one architectural feature of the house.

[0010] The method may further comprise designating a final inferred representation with a substantially similar geometry to the semantic information. In some embodiments, the substantially similar geometry is a minimum loss divergence between the first and second inferred representations. In some embodiments, the substantially similar geometry is a minimum Euclidean distance between the first and second inferred representations. In some embodiments, the substantially similar geometry is a minimum loss convergence between the second and third inferred representations. In some embodiments, the substantially similar geometry is a minimum Euclidean distance between the second and third inferred representations. In some embodiments, the method may further comprise selecting the final inferred representation as a structured geometry of the house.

[0011] Another aspect of the present disclosure relates to a method of predicting a semantic geometry of a house. The method may include receiving at least one image depicting the house. The at least one image may observe the house from a first pose. The method may further include generating first semantic information of the house from the at least one image. The method may further include sampling a first latent code for decoding into a representation. The first latent code may be one of a plurality of latent codes from a trained distribution of building objects. The method may further include inferring a first and second representation from the first latent code based on the first pose. The first inferredrepresentation is a same classification as the first semantic information and the second inferred semantic representation may be a structure wireframe comprising segmented geometries. The method may further include computing a loss between the semantic information and the first inferred representation or second inferred representation. The method may further include selecting a second latent code from the plurality of latent codes based on the loss. The method may further include inferring a third and fourth representation from the second latent code based on the first pose. The third inferred representation may be the same classification as the first semantic information. The fourth inferred semantic representation may be a structured wireframe comprising segmented geometries.

[0012] Another aspect of the present disclosure relates to a method of predicting a semantic geometry of a house. The method may include receiving at least one image depicting the house. The at least one image may observe the house from a first pose. The method may further include generating semantic information of the house from the at least one image. The method may further include sampling a latent code for decoding into a representation. The latent code may be one of a plurality of latent codes from a trained distribution of building objects. The method may further include inferring at least one representation from the latent code based on the first pose. The method may further include computing a loss between the semantic information and the at least one representation. The method may further include iteratively selecting a successive latent code from the plurality of latent codes based on the loss. The method may further include iteratively inferring at least one successive representation from the successive latent code based on the first pose.

[0013] In some embodiments, each latent code of the plurality of latent codes is an encoded value of one or more parameters for a structural description of building objects. In some embodiments, the encoded value is learned from an implicit neural field. In some embodiments, the one or more parameters is a roof ridge length. In some embodiments, inferring the at least one representation is performed by decoding the latent code with an implicit neural field. In some embodiments, inferring the at least one representation is performed by decoding the first latent code with a graph neural network. In some embodiments, inferring the at least one successive representation is performed by decoding the successive latent code with an implicit neural field. In some embodiments, inferring the at least one successive representation is performed by decoding the successive latent code with a generative neural network.

[0014] In some embodiments, computing the loss between the semantic information and the at least one inferred representation further comprises projecting the at least one inferred representation into the semantic information of the at least one image. In some embodiments, selecting a successive latent code is based on a gradient of the loss with respect to the at least one inferred representation. In some embodiments, the semantic information comprises a silhouette of the house from the first pose. In some embodiments, the at least one inferred representation or successive representation comprises a silhouette of the house from the first pose. In some embodiments, the semantic information comprises at least one architectural feature of the house. In some embodiments, the at least one architectural feature is a linear feature. In some embodiments, the at least one architectural feature is a roof ridge line. In some embodiments, the at least one inferred representation or successive representation comprises at least one architectural feature of the house.

[0015] The method may further comprise designating a final inferred representation with a substantially similar geometry to the semantic information. In some embodiments, the substantially similar geometry is a minimum loss divergence between iterative successive inferred representations. In some embodiments, the substantially similar geometry is a minimum Euclidean distance between iterative successive inferred representations. The method may further comprise selecting the final inferred representation as a structure geometry of the house.

[0016] Another aspect of the present disclosure relates to a non-transient computer- readable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method for predicting a semantic geometry of a house. The method may include receiving at least one image depicting the house. The at least one image may observe the house from a first pose. The method may include generating semantic information of the house from the at least one image. The method may include sampling a first latent code for decoding into a representation. The first latent code may be one of a plurality of latent codes from a trained distribution of building objects. The method may include inferring a first representation from the first latent code based on the first pose. The method may include computing a loss between the semantic information and the first inferred representation. The method may include selecting a second latent code from the plurality of latent codes based on the loss; selecting may be choosing a new latent code or providing an update or modification to the latent code’ s position in latentspace. The method may include inferring a second representation from the second latent code based on the first pose.

[0017] Yet another aspect of the present disclosure relates to a system configured for predicting a semantic geometry of a house. The system may include one or more hardware processors configured by machine-readable instructions. The processor(s) may be configured to receive at least one image depicting the house. The at least one image may observe the house from a first pose. The processor(s) may be configured to generate semantic information of the house from the at least one image. The processor(s) may be configured to sample a first latent code for decoding into a representation. The first latent code may be one of a plurality of latent codes from a trained distribution of building objects. The processor(s) may be configured to infer a first representation from the first latent code based on the first pose. The processor(s) may be configured to compute a loss between the semantic information and the first inferred representation. The processor(s) may be configured to select a second latent code from the plurality of latent codes based on the loss; selecting may be choosing a new latent code or providing an update or modification to the latent code’ s position in latent space. The processor(s) may be configured to infer a second representation from the second latent code based on the first pose.

[0018] Another aspect of the present disclosure relates to a computer-implemented method. The computer implemented method may include receiving a digital image. The digital image may comprise a set of pixels that depict a building observed from a first perspective. The computer implemented method may further include transforming the set of pixels that depict the building into a first semantic representation. The computer implemented method may further include sampling a latent code that corresponds to a first point in a learned latent space. Points in the learned latent space may represent a distribution of building structures according to encoded parameters. The computer implemented method may further include decoding the latent code into a second semantic representation according to the first perspective. The computer implemented method may further include determining a loss between the first semantic representation and the second semantic representation. The computer implemented method may further include updating the latent code to correspond to a second point in the learned latent space based on the loss. The computer implemented method may further include decoding the updated latent code into a third semantic representation according to the first perspective. The computer implemented method may further include detecting a convergence between the first semantic representation and thethird semantic representation. The computer implemented method may further include decoding the updated latent code into a semantic geometry for the building. The semantic geometry may represent classifications for one or more structural features of the building.

[0019] In some embodiments, the first semantic representation comprises a two- dimensional segmentation map. In some embodiments, the first semantic representation comprises a silhouette of the building. In some embodiments, decoding the latent code into a second semantic representation further comprises providing an implicit neural representation of the latent code. The implicit neural representation may have a density volume.

[0020] Embodiments of such a computer implemented method may further include providing a silhouette of the implicit neural representation. In some embodiments, determining a loss between the first semantic representation and the second semantic representation comprises detecting a difference between a projection of the first semantic representation and the second semantic representation. In some embodiments, updating the latent code comprises selecting a new latent code from the learned latent space. In some embodiments, updating the latent code comprises modifying the sampled latent code. In some embodiments, updating the latent code based on the loss comprises applying a gradient direction derived from the loss to the sampled latent code.

[0021] In some embodiments, updating the latent code based on the loss comprises traversing the latent space according to a distribution fit to the learned latent space. In some embodiments, decoding the updated latent code into a third semantic representation further comprises providing an updated implicit neural representation of the updated latent code, wherein the implicit neural representation has a density volume.

[0022] The computer implemented method may further include providing an updated silhouette of the updated implicit neural representation. In some embodiments, detecting a convergence between the first semantic representation and the third semantic representation comprises detecting a substantial similarity between a projection of the first semantic representation and the third semantic representation. In some embodiments, detecting a convergence between the first semantic representation and the third semantic representation comprises detecting a subsequent updated latent code does not produce a lower loss between a projection of the first semantic representation and the third semantic representation. As used herein, “substantial similarity” may mean meeting predefined threshold criteria or convergence between representations, such as pixel area overlap of at least 90% betweensilhouettes of representations, angular alignment of features within 5 degrees, or reprojection of key points or nodes within 5 pixels according to a 1024 resolution projection.

[0023] In some embodiments, decoding the updated latent code into the semantic geometry comprises executing a graph neural network on the updated latent code. In some embodiments, the semantic geometry comprises a graph structure. In some embodiments, the graph structure comprises a plurality of edges that correspond to edges on a roof of the building and a plurality of nodes that correspond to intersections between the edges on the roof of the building. In some embodiments, decoding the updated latent code into the semantic geometry for the building comprises classifying each edge of the plurality of edges as one of a plurality of edge types.

[0024] Another aspect of the present disclosure relates to a computer-implemented method. The computer-implemented method may comprise receiving a digital image. The digital image may comprise a set of pixels that depict a building observed from a first perspective. The computer-implemented method may further comprise transforming the set of pixels that depict the building into a two-dimensional (2D) segmentation map. The computer- implemented method may further comprise sampling a latent code from a learned latent space comprising a plurality of latent codes. Any one latent code may be sampled as a point from the learned latent space. The computer-implemented method may further comprise updating the latent code through one or more iterations until a convergence is detected between the 2D segmentation map and an inferred segmentation map decoded from the latent code. Each iteration may comprise decoding the latent code into the inferred segmentation map for a hypothetical building structure according the first perspective. Each iteration may further comprise determining a loss between the 2D segmentation map and the inferred segmentation map. The loss may represent differences between the 2D segmentation map and the inferred segmentation map when projected onto each other. Each iteration may further comprise modifying the latent code based on the loss. The computer-implemented method may further comprise decoding the modified latent code into a semantic geometry for the building. The semantic geometry may represent classifications for one or more structural features of the building.

[0025] Another aspect of the present disclosure relates to a computer-implemented method for learning a latent code for parameters of a building geometry. The computer-implemented method may comprise providing a multi-dimensional representation of a building structure.The multi-dimensional representation may have a parameterized geometry and image data associated with the building structure. The computer-implemented method may further comprise generating a first semantic representation of the image data associated with the building structure. The computer-implemented method may further comprise generating a vector of randomized values associated with the parameterized geometry for each building structure. The computer-implemented method may further comprise inputting the vector into a first network to output a second semantic representation. The computer-implemented method may further comprise determining a loss between the first semantic representation and the second semantic representation. The computer-implemented method may further comprise iteratively optimizing, based on the loss, the random values associated with the parameterized geometry of the vector and one or more weights of the first network until an output of the iteratively optimized first network executed upon an iteratively optimized vector converges with the first semantic representation.

[0026] The computer-implemented method may further comprise training a second network by providing a second network the iteratively optimized vector, generating a third semantic representation, determining a loss between the parameterized geometry of the multidimensional representation of the building structure and the third semantic representation, and iteratively optimizing, based on the loss, one or more weights of the second network until an output of the iteratively optimized second network executed upon the iteratively optimized vector converges with the parameterized geometry of the multi-dimensional representation of the building structure.

[0027] The computer-implemented method may further comprise generating a fourth semantic representation based on the image data associated with the building structure. The computer-implemented method may further comprise inputting the iteratively optimized vector into the first network to output an updated second semantic representation and the second network to output an updated third semantic representation. The computer- implemented method may further comprise determining an updated density loss between the first semantic representation and the updated second semantic representation. The computer- implemented method may further comprise determining an updated semantic loss between the updated third semantic representation and the fourth semantic representation. The computer-implemented method may further comprise iteratively re-optimizing, based on the updated density loss and the updated semantic loss, one or more values associated with the parameterized geometry of the iteratively optimized vector and one or more weights of thefirst network and one or more weights of the second network until an output of the iteratively re-optimized first network and iteratively re-optimized second network executed upon an iteratively re-optimized vector converges with the first semantic representation and fourth semantic representation respectively.

[0028] In some embodiments, generating the first semantic representation comprises generating a silhouette of the building structure based on the image data. In some embodiments, generating the vector of randomized values comprises generating a 512- dimensional vector. In some embodiments, the first network is an implicit neural representation decoder network, and the second semantic representation is an implicit neural representation having a volume density. In some embodiments, the second network is a graph neural network, and the third semantic representation is a semantic spatial graph of geometry inferred from the provided vector. In some embodiments, the fourth semantic representation is a two-dimensional semantic segmentation of the image data.

[0029] The computer-implemented method may further comprise learning a latent code for a plurality of multi-dimensional building structures based on the iteratively optimized vector and arranging a latent space comprising the plurality of known latent codes associated with a plurality of multi-dimensional building structures. The computer-implemented method may further comprise fitting a probability distribution to the plurality of known latent codes, wherein the probability distribution comprises additional points in the latent space corresponding to latent codes for hypothetical building structures among the known latent codes associated with the plurality of multi-dimensional building structures.

[0030] In some embodiments, the multi-dimensional representation is a three-dimensional geometrical model. In some embodiments, the multi-dimensional representation is a two- dimensional floorplan.

[0031] Another aspect of the present disclosure relates to a system. The system may comprise a database. The database may comprise a plurality of multi-dimensional building models, each multi-dimensional building model having a parameterized geometry and image data associated with the building structure. The system may further comprise an encoder network capable of generating encoded values for the parameterized geometry into a latent code. The latent code may be a multidimensional vector associated with a respective multidimensional model and the encoder network if further capable of iteratively optimizing the encoded values based on losses provided by one or more decoder networks. The system mayfurther comprise an image segmentation network configured to provide one or more semantic outputs based on the image data. The system may further comprise a first decoder network capable of decoding a provided latent code and generating a first semantic output for comparison with the one or more semantic outputs of the image segmentation network. The system may further comprise a second decoder network capable of decoding a provided latent code and generating a second semantic output for comparison with the parameterized geometry or one or more outputs of the image segmentation network.

[0032] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.

[0033] Desired representations for inference include structured geometry of an inferred, such as semantically meaningful lines and nodes of a building model. Though certain networks like a graph neural network are capable of making semantic predictions based on evidence (such as predicting a semantically meaningful three-dimensional wireframe from captured two-dimensional images), reconstruction of certain geometries is prone to failure because they cannot be effectively compared to ground truth images. For example, when capturing data with wide baseline changes, such as ground level imagery capturing a building by circling about the building’s perimeter, triangulation of semantically meaningful data is inconsistent. A first image may capture a roof apex, from which a semantic point may be inferred, but a second image may not observe that same point leading to difficulty in resolving a prediction based on both images when both images view different aspects of a common object. Given this variable interpretation of an object’s geometry due to the variable perspectives, the output space is not continuous, and differences cannot be learned using basic techniques such as gradient based optimization. Additionally, when a semantic output is generated, it may suffer from occlusion artifacts: evidence is encoded into the prediction because it was observed in the image merely because it was in the foreground of the image, or output such as a wireframe does not self-occlude certain geometries that should not bevisible from a given perspective (a wireframe is “see through”). This makes it difficult for a computer vision pipeline to resolve differences in the output and the provided evidence.

[0034] An implicit neural representation (INR) may complement a graph neural network to generate a semantically continuous representation for reconstruction-related predictions, while generating a mesh that will occlude certain geometries produced by the graph neural network, or producing a probability that geometries produced by the graph neural network would not be visible, based on the viewing perspective. It may thus be leveraged by a computer vision pipeline to evaluate the suitability of the prediction. When a neural rendering such as an INR matches a segmentation of presented evidence (for example, when a silhouette of an INR of a building mesh aligns to a silhouette of the building as from a corresponding 2D image), the same data that produced the INR can be used to produce a semantic wireframe, such as through a graph neural network, to produce structured geometry for the building in the 2D image.BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The specification references the following appended figures, in which use of like reference numerals in different figures is intended to illustrate like or analogous components.

[0036] FIG. 1 is a block diagram illustrating an example of a computing environment, according to certain aspects of the present disclosure.

[0037] FIG. 2 is a block diagram illustrating another example of a computing environment, according to certain aspects of the present disclosure.

[0038] FIG. 3 is a diagram illustrating an example of a process flow for training a system to generate semantic geometries for a building based on 2D images of the building, according to certain aspects of the present disclosure.

[0039] FIG. 4 illustrates exemplary comparisons between known silhouettes for a building and silhouettes inferred from a latent code, according to certain aspects of the present disclosure.

[0040] FIG. 5 illustrates exemplary inferred silhouettes decoded from iterations of a latent code, according to certain aspects of the present disclosure.

[0041] FIG. 6 illustrates exemplary renderings of a probability distribution fit to a learned latent space, according to certain aspects of the present disclosure.

[0042] FIG. 7 is a diagram illustrating an example of a process flow for generating semantic geometries for a building based on 2D images of the building, according to certain aspects of the present disclosure.

[0043] FIG. 8 illustrates exemplary data outputs from a GNN decoder for a given latent code, according to certain aspects of the present disclosure.

[0044] FIG. 9 illustrates an exemplary rendering of a preliminary semantic geometry from the outputs of a GNN decoder, according to certain aspects of the present disclosure.

[0045] FIG. 10 illustrates exemplary comparisons between iteratively inferred silhouettes and semantic geometries generated from a latent code and a predicted silhouette generated from an image of a building, according to certain aspects of the present disclosure.

[0046] FIG. 11 illustrates exemplary movement within the latent space between iterations of an INF decoder, according to certain aspects of the present disclosure.

[0047] FIG. 12 is a flowchart illustrating an example of a process for generating a semantic geometry for a building from a digital image of the building, according to certain aspects of the present disclosure.

[0048] FIG. 13 is a flowchart illustrating an example of a process for superimposing a semantic geometry of a house’ s roof over an image of the house for display on a mobile device, according to certain aspects of the present disclosure.

[0049] FIG. 14 is a flowchart illustrating an example of a process for superimposing measurements for a house’s roof over an image of the house for display on a mobile device, according to certain aspects of the present disclosure.DETAILED DESCRIPTION

[0050] Certain aspects and features of the present disclosure relate to reconstructing plausible, semantic, and structured 3D representations of a building. 3D reconstruction lies at the heart of 3D vision and enables myriad applications ranging from autonomous driving to medicine. However, many 3D reconstruction techniques do not go beyond recovering low- level 3D information (coordinates, colors, normals, pointwise semantics, etc.), whereas applications like indoor scanning, reverse engineering or shape parsing require the recoveryof semantics and parsimonious structure (e.g., parametric forms, hierarchies, or higher-order relationships) from data.

[0051] Many 3D reconstruction techniques rely on a bottoms-up approach. For example, such techniques may capture images of an object and generate a geometry for the object based on observed data, such as through structure from motion (SfM). This is an unstructured geometry, as the resultant reconstruction conveys no semantic information of the scene, it is simply a representation of the spatial data. The captured images may be processed by segmentation models to classify particular pixels in an image, and these classifications may be applied separately to the unstructured geometry to create structurally meaningful data.

[0052] In such a pipeline, noise and training data biases can create uncertain outcomes. For example, observations of certain data may lead to a reasonable reconstruction or prediction of a building’s geometry. However, errors in sensing or interpretation may lead to reconstructions that are valid for the observed data, but nonetheless are incorrect or even nonsensical. Completeness may also suffer. For example, structure from motion outputs produced from observed data using a Colmap algorithm may result in missing regions in the output where the structure should be present but was not observed.

[0053] Further, these techniques lack the benefits associated with a top-down approach that enable more accurate predictions. In a top-down approach, a parameterized geometry may be optimized for specific inputs. In some embodiments enabled by the disclosure herein, a parameterized geometry is expressed as one or more latent priors. A latent prior is an encoded value conforming to certain distributions of parameters that define a geometry of an object. A latent prior may be represented as within a probability distribution trained on a plurality of learned latent codes. As a result, a geometry generated from a latent prior will always produce a plausible geometrical representation for an encoded object (such as a building), as latent codes that produce unlikely geometries would fall outside the trained distribution.

[0054] In many cases, there is a lack of a top-down or holistic, structural, and semantic priors due in part to a lack of accurately labeled data. In some embodiments, a corpus of geometrical ground truth representations, or models of buildings, is generated and used in training to learn the values and distributions for the latent prior. Each geometrical ground truth representation in the corpus may comprise a plurality of parameters describing anunderlying object’s geometries. In some embodiments, and as specifically referred to herein, the underlying object is the exterior of a house.

[0055] In some embodiments, for a respective house of the corpus, a series of parameters are identified, such as roof ridge length, roof facet width, and eave information. A multidimensional vector is generated to represent the identified parameters. This multidimensional vector may be a 32-dimensional vector, a 64-dimensional vector, a 128- dimensional vector, a 256-dimensional vector, a 512-dimensional vector, or a 1024- dimensional vector but other numbers of dimensions can be used, including numbers of dimensions that are not a multiple or power of two. In some embodiments, the values within the vector are produced at random at the beginning of training and provided to the respective decoder networks for providing the desired representations. For example, a structure decoder such as graph neural network to produce semantic geometry predictions, or an INR decoder to produce a density field. In some embodiments, each multidimensional vector is a latent code for training samples. Outputs of the network are compared to ground truth of the house the multidimensional vector was based on, and then both the weights of the respective networks and the values of the multidimensional vector training sample are adjusted. The adjusted multidimensional vector is then decoded by the respective networks and comparisons to ground truth is made and iterative adjustments continue in this fashion until the network(s) converge on their respective samples. Training may be done with batches of multidimensional vector latent codes as training samples, or individual multidimensional vector latent code training samples. In some embodiments, training various networks is sequential; an INR decoder is trained first with the latent variability adjustments, and then a graph neural network has supervised training with the learned latent codes.

[0056] With a learned set of latent codes representing the corpus of property data as a plurality of multidimensional vectors, additional latent codes may be identified, such as through interpolation or extrapolation of the known or learned latent codes, or by varying values of the multidimensional vectors of the known or learned latent codes. Interpolation or extrapolation may be by fitting a distribution to the known or learned latent codes, such as Mixture of Gaussians. Varying the values, in some embodiments, is performed by applying a gaussian to particular values within any one multidimensional vector. These additional latent codes represent hypothetical structures only. Taken together, the additional latent codes with the known or learned latent codes, an entire latent space of latent codes is provided. In some embodiments, a distribution fit to the latent space, either the known or learned latent codes oradditional latent codes, maps similarities between latent codes. Similarities between latent codes may be expressed as clusters within a distribution space.

[0057] The latent space may be structured, therefore, in a way that links each encoded value, implicitly linking each decoded geometry, such that each point in the latent space represents a latent code for a building geometry, with peaks or clusters indicating locations of latent codes with higher likelihood of conforming to particular distributions of parameters identified in the corpus of ground truth representations.

[0058] In a test scenario, an image of a particular house is presented. The image may be segmented, such as using the segmentation network described in co-owned US Pat. Pub. 2021 / 0243362, hereby incorporated by reference in its entirety. A latent code may be selected from the plurality of latent codes, and a decoder network produces a representation of the geometry associated with that latent code according to the pose of the image. The decoder network may be trained to predict a representation of the latent code according to a particular segmentation protocol. For example, the image may be segmented to produce a silhouette of the house and the decoder network, such as an implicit neural field, may decode the latent code to produce a representation from which semantics like a silhouette for that latent code from a similar pose can be rendered. In some embodiments, other semantics like a normalized object coordinate space, edge detection or other semantics are rendered. In some embodiments, a graph neural network trained on the semantic information of the corpus of ground truth representations is used to predict a semantic wireframe representation of the particular house from the same latent code with the same pose data.

[0059] In some embodiments, a loss between the predicted representation and the semantic information of the image is calculated. The loss is a difference in the geometry of the predicted representation and the semantic information of the image, such as determined from reprojecting the geometry or silhouette of one into the other.

[0060] In some embodiments, the selected latent code is updated based on the loss. In some embodiments, updating may be selecting a next latent code or otherwise updating the selected latent code’ s position in the latent space. The updated latent code is used to produce a decoded representation that better fits the segmented information of the image than the previous latent code. In some embodiments, the updated latent code is based on a gradient descent of the loss, such as where the value of one or more dimensions adjusts an incremental amount. In some embodiments, this updating or iteration of latent codes is repeated until adecoded representation of an instant latent code matches an applicable representation of the segmented image. In some embodiments, matching the segmented image is substantial alignment of the respective geometries, either in a Euclidean sense in that a successive latent code does not improve the geometric fit on the whole despite a better fit among select points and features, or a successive latent code has a loss divergence substantially similar to the previous latent code selection.

[0061] With a final or matching latent code selected, the geometry of the resultant decoded latent code can be further decoded as desired, such as by a graph neural network to produce a predicted segmented geometry, or structured geometry, for the house in the image.

[0062] By performing the above, the problem of incomplete or inaccurate data in a reconstruction pipeline leading to incomplete or implausible reconstructions is solved by iteratively minimizing a loss between observed data and outputs of a predictive model while jointly maximizing the probability assigned to the outputs by a learned prior. Each prediction is a plausible house geometry and therefore will be complete in its output for facilitating a successive prediction; the network will not get “lost” trying to optimize a prediction because it is always predicting a likely incrementally better fit. Further, by updating latent codes at each iteration based on the loss from an inferred density representation, as opposed to a loss between a predicted wireframe, the resulting predictions at each iteration may follow a more continuous and predictable path within the latent space, with higher fidelity changes from one prediction to the next.

[0063] These illustrative examples are given to introduce the reader to the general subject matter discussed here and are not intended to limit the scope of the disclosed concepts. The following text describes various additional and optional features and examples with reference to the drawings in which like numerals indicate like elements, and directional descriptions are used to describe the illustrative embodiments but, like the illustrative embodiments, should not be used to limit the present disclosure. The elements included in the illustrations herein may not be drawn to scale.

[0064] FIG. 1 is a block diagram illustrating an example of a computing environment 100, according to certain aspects of the present disclosure. Computing environment 100 may include user device 110 and server 120. User device 110 may be any portable (e.g., mobile devices, such as smartphones, tablets, laptops, application specific integrated circuits (ASICs), or the like) or non-portable computing device (e.g., desktop computer, electronickiosk, or the like). User device 110 may be connected to gateway 140 (e.g., a Wi-Fi access point), which provides access to network 130. Network 130 may be any public network (e.g., Internet), private network (e.g., Intranet), or cloud network (e.g., a private or public virtual cloud). User device 110 may communicate with server 120 through network 130.

[0065] A native or web application may be executing on user device 110. The native or web application may be configured to perform various functions relating to analyzing an image of a physical structure, such as a house. As an illustrative example, the native or web application may be configured to perform a function that captures one or more 2D images of house 150 within the camera’s field of view 160. The function may also transmit the 2D image(s) to server 120 for analysis. Server 120 may analyze the 2D image(s) to automatically detect or compute a predicted segmented geometry, or structured geometry, for house 150. For example, server 120 may execute one or more computer vision techniques, such as feature detection and feature map determination, edge detection, or semantic image segmentation, or produce silhouettes of house 150 from the same poses or perspectives as the images provided to server 120. Server 120 may then execute a decoder network on a latent code to produce inferred semantics similar to as performed on the images. For example, server 120 may produce a silhouette of house 150 from an image and then extract a silhouette from an inferred density representation from a latent code according to the same poses or perspectives as the images provided to server 120. Server 120 may then compare the silhouettes predicted from the latent code with the silhouettes produced from the images to determine a loss between the two sets of silhouettes. Based on the amount of loss, server 120 may then iteratively update the latent code, execute the decoder, and determine a new loss until a resulting loss meets one or more predefined threshold criteria. Server 120 may then decode the final latent code, such as by executing a graph neural network on the final latent code, to produce a predicted segmented, structured, or semantic geometry for house 150.

[0066] Once the semantic geometry for house 150 has been produced, server 120 may transmit the semantic geometry for house 150 to the native or web application executing on user device 110. In response to receiving the semantic geometry for house 150, the native or web application may display a final image 170, which presents an overlay of the semantic geometry on top of one or more of the 2D images of house 150. The present disclosure is not limited to performing the function on server 120. The function can be entirely performed on user device 110 without the need for or use of server 120. Additionally, the present disclosure is not limited to the use of a native or web application executing on user device 110. Anyexecutable code (whether or not the code is a native or web application) can be configured to perform at least a part of the function for algorithmically determining the dimensions of house 150.

[0067] FIG. 2 is a block diagram illustrating components of server 120, according to certain aspects of the present disclosure. In some implementations, server 120 may include a plurality of services such as: image processor 211; segmentation engine 212; latent space module 213; Implicit Neural Field (INF) decoder 214; probability engine 215; loss module 216; and Graph Neural Network (GNN) decoder 217. One or more of the services of server 120 may be implemented as standalone and / or integrated software applications stored in a memory of server 120. Server 120 may also include processor 200 and one or more databases, such as house model database 218.

[0068] Processor 200 may be a processor or processing apparatus (e.g., a server) configured to execute executable code that performs the various functions provided by the plurality of services of server 120, according to implementations described herein. The executable code may be stored in a memory or non-transitory storage device (not shown) associated with the processor 200. Processor 200 may be used to train and / or execute INF decoder 214, GNN decoder 217, and other machine learning models described herein, such as region proposal models, 2D graph models, 3D graph models, or distance determination models, according to certain implementations described herein.

[0069] House model database 218 may be configured to include a data structure that stores one or more existing 3D models of physical structures. Non-limiting examples of a 3D model of a physical structure include a CAD model, a 3D shape including a cuboid base with an angled roof, a pseudo-voxelized volumetric representation, mesh geometric representation, a graphical representation, a 3D point cloud, or any other suitable 3D model of a virtual or physical structure. The 3D models of physical structures may be generated by a professional or may be automatically generated (e.g., a 3D point cloud may be generated from a 3D camera). In some cases, house model database 218 may be used to store 3D models generated using techniques described herein.

[0070] Additionally, or alternatively, house model database 218 may store 2D images of physical structures. The 2D images may be captured by professionals or users of the native or web application, or may be generated automatically by a computer (e.g., a virtual image). Referring to the example illustrated in FIG. 1, the image of house 150, which is captured byuser device 110, may be transmitted to server 120 and stored in house model database 218. The images stored in house model database 218 may also be stored in association with metadata, such as the focal length of the camera that was used to capture the image, a resolution of the image, a date and / or time that the image was captured, global positioning coordinates, or the like. In some implementations, the images stored in house model database 218 may depict top-down views of physical structures. In other implementations, house model database 218 may store images depicting ground-level views of houses, which can be evaluated using the end-to-end model, solver network, or other machine learning models described herein. Processor 200 may process the images to generate descriptors, for example, by detecting a set of keypoints (e.g., up to 14 keypoints, or more) within an image.

[0071] The images and / or 3D models stored in house model database 218 may serve as inputs to machine-learning or artificial-intelligence models. The images and / or the 3D models may be used as training data to train the machine-learning or artificial -intelligence models or as test data to generate predictive outputs. Machine-learning or artificial -intelligence models may include supervised, unsupervised, or semi-supervised machine-learning models.

[0072] Image processor 211 may be used to receive and pre-process 2D images of a house for which a semantic geometry will be predicted. Image processor 211 may receive and pre- process one or more 2D images for a house. The 2D images may be received from a remote device, such as the native or web application executing on user device 110. Additionally, or alternatively, the 2D images may be accessed from a database, such as house model database 218.

[0073] In some embodiments, the 2D images may be preprocessed to evaluate one or more parameters for the images and / or to determine the suitability of the images for use in generating a semantic geometry. This may be calculated as an angular perspective score derived from an angle between a line or ray from the focal point of the camera to the point and the orientation of the surface or feature on which the point lies. The angle between the focal point of the camera and the surface of the physical structure informs a degree of depth information that can be extracted from the resulting captured image. For example, an angle of 45 degrees between the focal point of the camera and the surface of a physical structure can provide optimal image data for segmentation information, which can improve the use of the techniques described herein to produce semantic geometries of the physical structure using 2D images.

[0074] In some cases, 2D images may be preprocessed in real-time as the image is captured (e.g., on and by user device 110). In other cases, evaluation of the 2D image may occur after capture, such as on and by user device 110 or by image processor 211 after receiving the 2D image from user device 110. The angular perspective score may indicate a proximity of the angle between the focal point of the camera and the surface of the physical structure to a particular value, such as 45 degrees. Various techniques may be used to evaluate the angular perspective score based on the angle, such as the absolute value of an angle difference from 45 degrees, a percentage difference of the angle compared to 45 degrees, etc.

[0075] In some examples, the angular perspective score may be compared with a threshold value or range and a notification can be generated if the angular perspective score is outside of threshold value or range to indicate that an additional 2D image should be captured at a different camera position with respect to the physical structure, such as to improve the use of the image in generating semantic geometries for the structure. In some examples, the notification can provide instructions to the user for a direction or distance to move the camera before capturing the additional 2D image. In some examples, an instructive prompt can be generated while a user is framing the physical structure within a viewfinder or on a display of the user device 110, thereby guiding the user to capture images with an optimal or improved angle relative to the surface of the physical structure. For example, a native application executing on a user device may provide an augmented reality (AR) output and indicate an angular perspective score or other suitability score, indicating whether a camera position falls in or outside of a target threshold position or range, an angular perspective score threshold value or range, or other suitability score threshold value or range, or indicate instructions for adjusting a position of the camera to improve the position, angular perspective score, or other suitability score.

[0076] Segmentation engine 212 may be used to generate semantic information for a building from one or more images of the building. In some embodiments, semantic information for a building includes one or more 2D segmentation maps generated by segmentation engine 212. Segmentation engine 212 may use any one or a combination of: feature extraction techniques, semantic segmentation techniques, boundary detection techniques, region merging techniques, labeling, and visualization to generate 2D segmentation maps from 2D images of a building. For example, segmentation engine 212 may detect target physical structures within a 2D image of a building. Detecting target physical structures can include performing one or more image segmentation techniques,which include inputting the 2D image into a trained classifier to detect pixels relating to the target physical structures of a building. When the target physical structures are detected, segmentation engine 212 can determine the dimensions of bounding boxes and render the bounding boxes around the target physical structures. The bounding boxes may be convex hulls or quadrilaterals that contain the image data of the target physical structures.Additionally, or alternatively, segmentation engine 212 can generate segmentation masks for each image and apply the segmentation masks to a respective 2D image. The segmentation masks may be trained separately to detect certain objects in an image. The bounding boxes and / or segmentation masks for each image may then be combined to form 2D segmentation maps, or silhouettes, of the building from the perspective of the camera when it took each respective image.

[0077] Latent space module 213 may select or update latent codes for decoding by INF decoder 214 and GNN decoder 217. As described further herein, latent codes may be selected from a latent space defined by a plurality of latent variables. Each unique instantiation of the plurality of latent variables in the latent space may be referred to as a latent code or a point in the latent space. In some embodiments, each latent code is a 512 dimensional vector. In other words, the dimensions of the latent space may be defined by the 512 variables as well as the ranges of possible values for each variable. In turn, each variable that makes up the latent space may inform one or more real-world parameters that can be used to define the structure of a building. In total, each point, or unique instantiation of the plurality of latent variables, within the latent space may represent a plausible geometry for a building, even if for a hypothetical building. As used herein, a plausible geometry for a building may be one that is completely parameterized (has no omissions in parameter values that translate as unobserved portions) and is informed on actual buildings observed during training (e.g., adhere to traditional building standards, concepts, or designs).

[0078] During training of INF decoder 214, latent space module 213 may learn the latent space based on losses (e.g., from loss module 216) between known silhouettes (e.g., from house model database 218) and silhouettes predicted by INF decoder 214. For example, latent space module 213 may iteratively adjust the values within a latent code for decoding by INF decoder 214 as well as adjusting the weights of the INF decoder 214 until the loss between a resulting silhouette and a known silhouette meets one or more predefined threshold criteria. Subsequently, latent space module 213 may store a record of the resulting latent code as a known or learned latent code that corresponds to a 3D model in house model database 218.Latent space module 213 may then repeat the process until a plurality of known or learned latent codes corresponding to each 3D model in house model database 218 have been learned. Latent space module 213 may identify additional latent codes, such as by adjusting values within the known or learned latent codes, fitting a distribution to the resultant latent space comprising the known or learned latent codes and the additionally generated ones, or combinations therein such as fitting a distribution to the latent space created by the plurality of known or learned latent codes and interpolating or extrapolating additional latent codes based on the distribution.

[0079] Latent space module 213 may be used in a similar fashion during runtime to identify a final latent code for a building captured in one or more 2D images. For example, latent space module 213 may select, or otherwise sample, an initial latent code from the latent space for decoding by INF decoder 214 to produce a semantic object from an inferred representation. In some embodiments, the semantic object is a silhouette. In some embodiments, the initial latent code is selected at random from the latent space, in some embodiments the initial latent code is sampled according to a cluster within the distribution fit, in some embodiments the sampled initial latent code corresponds to a known or learned latent code from training (and therefore reflects a previously observed building with a reconstructed model). Based on losses (e.g., from loss module 216) between the inferred silhouette and a silhouette predicted from a 2D image (e.g., by segmentation engine 212), as well as one or more probability distributions of the known latent codes, latent space module may iteratively select a new latent code, or update the initial latent code, until the loss between a resulting inferred silhouette and the predicted silhouette meets one or more predefined threshold criteria. While referred to herein as selecting, or sampling, an initial latent code and iteratively selecting, or sampling, new latent codes, the process described above may also be referred to in terms of a latent code object that may be generated with a unique latent code. Subsequently, the latent code object may be updated by replacing the unique latent code with a new unique latent code selected or sampled from the latent space. Updating a latent code may also be by applying a vector change, such as a matrix value, to the previous latent code.

[0080] INF decoder 214 may be used to generate one or more representations of a building from a latent code and a specified pose or perspective. The one or more representations may include a volumetric density rendering from which various semantics may be extracted such as a 2D segmentation mask, implicit surfaces, and / or a silhouette. INF decoder 214 maymodel a continuous field, such as the shape or contour of a building, by learning a neural network that implicitly represents the field without explicitly modeling the coordinates or discrete elements within the field. As described further herein, INF decoder 214 may be trained on known semantics like silhouettes determined for 3D models in house model database 218. Once trained, INF decoder 214 may take, as inputs, a latent code and one or more poses (e.g., determined from a 2D image of a building), and produce one or more silhouettes or other representations of a building from each of the one or more poses. In some embodiments, the known silhouettes determined for 3D models in house model database 218 are used to train INF decoder 214 and predict an encoding for a series of plausible geometries as latent codes. During execution, latent space module 213 may select, or sample, a latent code from the latent space for decoding by INF decoder 214 into the corresponding plausible geometry.

[0081] Probability engine 215 may be used to fit one or more distributions to the latent space. For example, after INF decoder 214 has been trained and a plurality of known latent codes corresponding to the 3D models in house model database 218 have been learned, probability engine 215 may apply one or more probability distribution models to the plurality of known latent codes to define one or more probability distributions for the latent space. In some embodiments, probability engine 215 applies a Mixture of Gaussians (MoG) model to the known latent codes to combine multiple Gaussian distributions within the latent space, or generate additional latent codes varying the values within any one known latent code. Probability engine 215 can use the resulting distribution to identify clusters of latent codes that result in similarities between plausible building geometries when decoded by INF decoder 214. The clusters of latent codes may represent plausible geometries that are more likely to exist in the real world based on their similarities with the geometries of the 3D models in house model database 218, or corresponding latent codes in the latent space. In other words, the further from any one cluster a latent code exists, the more likely its decoded representation will produce a loss when compared to image data.

[0082] During runtime, the probability distributions generated by probability engine 215 may be used by latent space module 213 to update or select a new latent code between iterations of the INF decoder 214. For example, based on the loss (e.g., from loss module 216) and a probability distribution, latent space module 213 can use Gradient Decent (GD) to iteratively update or select new latent codes along a path of latent codes in the latent space that will ultimately result in matching semantics with received images.

[0083] Loss Module 216 may be used to compute losses between semantics, such as silhouettes. The resulting loss may represent the differences between the two silhouettes and / or may quantify how well INF decoder 214 can reconstruct a silhouette from a latent code compared to a known silhouette corresponding to the latent code. Loss module 216 may compute one or more losses between compared semantics, including a Normalized Object Coordinate Space (NOCS) loss, a mask loss, a depth loss, a normal consistency loss, an eik loss, a cross-entropy loss, and a Generative Adversarial Network (GAN) loss. During training of INF decoder 214, loss module 216 may be used to compute losses between a known silhouette for a 3D model in house model database 218 and an inferred silhouette produced from INF decoder 214. Subsequently, the loss may be used by latent space module 213 to update the values within latent code used by the INF decoder 214, thereby learning to encode the latent prior values. During runtime, loss module 216 may determine a loss between a silhouette of a building predicted from a 2D image (e.g., by segmentation engine 212) and a silhouette of a building inferred from a latent code (e.g., by INF decoder 214). Subsequently, the loss may be used by latent space module 213 to update the latent code used for a next iteration of INF decoder 214 as the network traverses a latent space to locate a latent code that best fits received evidence (in the form of captured 2D images).

[0084] In some embodiments, loss module 216 is used to compute losses between semantic geometries predicted from an image and semantic geometry predicted from a latent code. The resulting loss may represent how well a semantic geometry of a building produced by GNN decoder, described further herein, fits or otherwise aligns with semantics predicted (e.g., by segmentation engine 212) from a 2D image of the building. During runtime, the loss may be used to make minor adjustments to the semantic geometry, such as removing cloned aspects or orienting lines to parametric expectations (like rooflines parallel to the ground, or lines meeting at right angles).

[0085] GNN decoder 217 may be used to generate one or more semantic geometries of a building from a latent code. As described further herein, GNN decoder 217 may include one or more graph neural networks trained on the 3D models in house model database 218. The one or more semantic geometries may represent and / or model spatial relationships and semantic information for physical features of a building, such as interior or exterior wall edges, roof edges, the vertexes where adjacent edges meet, distances and / or angles between edges and / or nodes, edge and / or node classifications, or the like. In some embodiments, the semantic geometries include semantic spatial graphs, or wireframes. The semantic spatialgraphs may include a plurality of nodes connected by edges, where the edges represent the physical edges of a building and the nodes represent the points where edges start, intersect, or end. In some embodiments, a final semantic geometry is derived from the outputs of GNN decoder 217, including: a plurality of node locations and the associated probabilities of each node corresponding to a physical feature of the building; the probability of any two nodes being adjacent (e.g., connected by an edge); and edge classification probabilities. For example, a final semantic geometry may include the nodes having a probability higher than a first predefined threshold probability and the associated edges between the resulting nodes with adjacency probabilities higher than a second predefined threshold probability.

[0086] FIG. 3 is a diagram illustrating an example of a process flow for training a system to generate semantic geometries for a building based on 2D images of the building, according to certain aspects of the present disclosure. The training process may be performed, at least in part, by or using any of the components illustrated in FIGS. 1-2, such as user device 110, server 120, house model database 218, latent space module 213, INF decoder 214, loss module 216, GNN decoder 217, and / or probability engine 215.

[0087] The training process may begin by defining latent space 312 and obtaining known representations 304 (e.g., from which silhouettes may be extracted) and known semantic geometries 308 from house model database 218. As described above, house model database 218 may store one or more 3D models for a plurality of existing physical structures, such as houses, apartment buildings, office buildings, outbuildings, or the like. Non-limiting examples of 3D models of a physical structure include a CAD model, a 3D shape including a cuboid base with an angled roof, a pseudo-voxelized volumetric representation, mesh geometric representation, a graphical representation, a 3D point cloud, or any other suitable 3D model of a virtual or physical structure. The one or more 3D models for a physical structure may include a plurality of parameters describing a building’s underlying geometries. In some embodiments, and as specifically referred to herein, the underlying geometries are for the exterior of a house. However, additional geometries, such as interior geometries, may be used as well as geometries for other types of buildings. In some embodiments, the plurality of parameters includes roof ridge lengths, roof facet widths, and eave information. For embodiments directed to interiors, floor plans and floor areas, wall heights, door locations, and the like are the underlying parameters.

[0088] Based on the 3D models, known representations 304 (such as silhouettes) and known semantic geometries 308 may be generated for each model in house model database 218. Known representations 304 for each model may include one or more 2D segmentation maps, 3D (semantic) point clouds, density rendering, implicit surfaces, and / or wireframes; additional outputs may be produced from the known representation, such as a rendered silhouette. Known geometries 308 for each model may include one or more semantic spatial graphs and / or wireframes including a set of vertices and a set of edges with associated classification labels, such as ridge, eve, or the like.

[0089] Latent space 312 may serve as a lower-dimensional representation of the various real -world parameters and / or measurements associated with a physical structure’s geometry. The size of latent space 312 may be based on the number of variables used to represent the real-world parameters, as well as the number of possible values for each variable. In some embodiments, latent space 312 includes 512 variables, such that each point in latent space 312, referred to as a latent code, is a 512-dimensional vector with values for each of the 512 variables. Each unique point in latent space 312 may represent a unique combination of values for each of the latent variables. As such, each latent code may represent a unique hypothetical geometry for a building.

[0090] The training process may proceed by training INF decoder 214 and producing known latents 320 for each model in house model database 218. As described above, latent space module 213 may initially assign random values to a latent code for each model in house model database 218. Latent space module 213 may then provide such initial latent code assigned to a particular model to INF decoder 214 to produce inferred representation 316 for the particular model. Loss module 216 may then compare inferred representation 316 with known representation 304 for the particular model to determine a loss, such as by comparing silhouettes of the two representations relative to a common perspective or pose.

[0091] FIG. 4 illustrates exemplary comparisons between known silhouettes for a building and silhouettes inferred from a latent code. It will be appreciated that such distinct silhouettes shown in FIG. 4 are unlikely in earlier iterations of training, when inferred representations 316 are likely to output as indecipherable masses by the network. Nonetheless, even in early iterations of training, the outputs are compared and variables within the initial latent code and weights of the decoder network 214 are adjusted to produce an incrementally closer match between the two. As described above, known silhouettes may be generated for a particularmodel from one or more different perspectives or poses. For example, and as illustrated, model 404 may be used to generate front silhouette 412 from front perspective 428 and side silhouette 424 from left-side perspective 416. As further described above, a single latent code may be used to generate one or more inferred silhouettes. For example, and as further illustrated, a particular latent code may result in side-inferred silhouette 432 from left-side perspective 416 and front-inferred silhouette 420 from front perspective 428. By comparing (e.g., using loss module 216) front silhouette 412 with front inferred silhouette 420 and / or side silhouette 424 with side-inferred silhouette 432, a reconstruction loss may be determined for the latent code used to generate the inferred silhouettes.

[0092] Referring back to FIG. 3, and based on the loss from loss module 216, latent space module 213 may adjust the variables with the initial latent code, or for a batch of latent codes, for the particular model or batch of models and provide the updated, or new, latent code(s) to INF decoder 214 when again the variables of the latent codes used as input and the weights of the decoder network to generate an output are adjusted for a subsequent iteration. This cycle may be repeated as many times as necessary until loss module 216 and / or latent space module 213 determines that the loss meets one or more predefined threshold criteria. In other words, loss module 216 and / or latent space module 213 may determine that silhouettes from inferred representation 316 and silhouettes from known representation 304 begin to converge and match each other sufficiently.

[0093] FIG. 5 illustrates exemplary inferred silhouettes generated by subsequent iterations of an INF decoder. As described above, when learning a latent code for a model, such as model 404, latent space module 213 may initially select random values for latent code from latent space 312, illustrated as Zo 504. INF decoder 214 may then use Zo 504 to generate initial silhouette 508. Initial silhouette 508 may then be compared with a known silhouette for model 404 to compute a loss. Based on the loss, latent space module 213 may adjust the values of latent code and / or otherwise update Zo 504 to Zi 512 as well as any weights of INF decoder 214. INF decoder 214 may then use Zi 512 to generate intermediate silhouette 516 and a new loss may be computed. This cycle may be repeated until a final latent code, illustrated as Z* 520, is used to generate final silhouette 524 with a loss that meets one or more predefined threshold criteria.

[0094] Referring back to FIG. 3, once a final latent code has been determined for a particular model, it may be associated with the particular model and stored in known latents320. Latent space module 213 may then proceed to repeat this process for each model in house model database. It should be understood that, while known latents 320 may correspond to previously known and / or existing physical structures, once trained, INF decoder 214 can use any latent code selected from within latent space 312, which may include additional latent codes generated by latent space module 213 that do not directly correlate to buildings in house model database 218, to and generate a silhouette for a building even hypothetical ones based on the additional latent codes. In other words,

[0095] With known latents 320 determined by training INF decoder 214, the training process may proceed to train GNN decoder 217. For example, as described above, with each latent code of known latents 320 being assigned to a 3D model in house model database 218, known geometries 308 for the 3D models may be provided along with their associated latent codes to GNN decoder 217. GNN decoder 217 may then train itself on the combination of the two sets of data to generate semantic geometries given a latent code from known latents 320 as input. Loss module 216 may assist in training GNN decoder 217 by, for example, computing reconstruction, regularization, and / or total losses.

[0096] The training process may end by executing probability engine 215 on known latents 320 to determine probability distribution 324 for latent space 312. As described above, probability engine 215 may apply one or more probability distribution models to known latents 320 to define probability distribution 324 for latent space 312. In some embodiments, probability distribution 324 is a Mixture of Gaussians (MoG) model that results in a plurality of distributions, or clusters, within latent space 312 that correspond to plausible building structures and / or plausible geometries that are more likely to exist in the real world.

[0097] FIG. 6 illustrates exemplary renderings of a probability distribution fit to a learned latent space. As illustrated, side perspective rendering 604 and top perspective rendering 608 include peaks, or clusters, in the probability distribution. In some embodiments, peaks or clusters in the probability distribution represent locations within the latent space where latent codes have a higher likelihood of conforming to distributions of parameters identified in the models of house model database 218. For example, and as further illustrated, latent codes selected from within first cluster 612 may result in first silhouettes 628 with a similar square shape. Latent codes selected from within second cluster 616 may result in second silhouettes 624 with “L” shapes. Lastly, latent codes selected from within third cluster 620 may result in third silhouettes 632 with “F” shapes.

[0098] FIG. 7 is a diagram illustrating an example of a process flow for generating semantic geometries for a building based on 2D images of the building, according to certain aspects of the present disclosure. The process may be performed, at least in part, by or using any of the components illustrated in FIGS. 1-2, such as user device 110, server 120, image processor 211, segmentation engine 212, latent space module 213, INF decoder 214, GNN decoder 217, and / or loss module 216.

[0099] The process may begin by receiving images 704 at image processor 211. As described above, images 704 may include one or more 2D images of a building. For example, images 704 may include 2D images of a building from different perspectives or viewpoints such that images 704 include images of the building from all sides. While described as images of a physical building that exists in the real world, embodiments described herein may also accept hypothetical images of a building, such as an architectural drawing, an artist’s rendition, or the like. Likewise, while described herein in reference to images of exterior surfaces of a building, images 704 may also include 2D images of a building’s interior, such as an image of a room including one or more interior walls, a floor, and / or a ceiling.

[0100] As further described above, image processor 211 may receive images 704 from a remote device, such as user device 110, used to capture images 704. Additionally, or alternatively, image processor 211 may obtain images 704 from a local or remote data store of building images. In some embodiments, image processor 211 may preprocess images 704 to evaluate their suitability for use in generating semantic geometries. For example, image processor 211 may determine whether images 704 depict a building structure with enough clarity and detail. One or more computer vision techniques may be used to make this determination, such as object detection techniques, object classification techniques, bounding box techniques, or the like. Additionally, or alternatively, image processor 211 may evaluate images 704 and determine a perspective, pose, or viewpoint from which images 704 were captured or from which a building is depicted.

[0101] Segmentation engine 212 may then analyze images 704 to produce semantic information for the building depicted in images 704, such as predicted silhouettes 708. As described above, predicted silhouettes 708 may include one or more 2D segmentation maps, each corresponding to an image of images 704. Segmentation engine 212 may use any one or a combination of: feature extraction techniques, semantic segmentation techniques, boundarydetection techniques, region merging techniques, labeling, and visualization to generate 2D segmentation maps from 2D images of a building.

[0102] The process may proceed with latent space module 213 sampling latent code 712 to represent the building of interest depicted in images 704. As described herein, latent code 712 may be a mutable object that stores a unique latent code comprising unique values for each latent variable in the latent space. As described above, latent space module 213 may initially sample latent code 712 at random from the latent space. Additionally, or alternatively, latent space module 213 may sample latent code 712 based, at least in part, on probability distribution 324. For example, latent space module 213 may select a latent code determined from probability distribution 324 to correspond to a most common plausible building structure, or even a known latent code 320 corresponding a known structure.

[0103] The process may proceed with INF decoder 214 generating inferred representation 716, such as a silhouette, from sampled latent code 712. As described above, INF decoder 214 may take, as inputs, latent code 712 and one or more poses or perspectives from which silhouettes of inferred representations 716 may be generated. For example, the one or more poses or perspectives may be determined from images 704 (e.g., by image processor 211 and / or segmentation engine 212) and may correspond to the poses or perspectives of predicted silhouettes 708. As further described above, inferred representation 716 may also include 2D segmentation masks (e.g., a binary mask, a multi-class mask, a probability mask, etc), implicit surfaces, contours, or the like.

[0104] In serial, or in parallel, with generating silhouettes of inferred representations 716, GNN decoder 217 may generate semantic geometries 720 from latent code 712. Additionally, or alternatively, GNN decoder 217 may generate semantic geometries 720 after latent code 712 has been updated for a final time, as described further herein. As described above, semantic geometries 720 may represent and / or model spatial relationships and semantic information for physical features of a building, such as interior or exterior walls, roofs, or the like. For example, semantic geometries 720 may include a semantic spatial graph comprising a plurality of nodes and edges, classifications for each and / or measurements for each. In some embodiments, semantic geometries 720 are derived from the outputs of GNN decoder 217, including: a plurality of node locations and the associated probabilities of each node corresponding to a physical feature of a building; the probability of any two nodes being adjacent (e.g., connected by an edge); and edge classification probabilities.

[0105] FIG. 8 illustrates exemplary data outputs from a GNN decoder for a given latent code. As illustrated, GNN decoder 217 may accept, as input, latent code 712 and generate semantic geometries 720. As further illustrated, semantic geometries may include or be based on node locations 804, node probabilities 808, adjacency probabilities 812, and edge class probabilities 816. In some embodiments, GNN decoder 217 generates a predefined number of nodes and uses latent code 712 to determine their associated node locations 804. Node locations 804 may be defined with reference to any suitable 3D reference frame.

[0106] As further described above, nodes may be used to represent locations where physical edges of a building’s structure start, end, and / or intersect with other physical edges (e.g., physical vertexes). As such, the number of vertexes for a given building may be less than the predefined number of nodes generated by GNN decoder 217. To determine which nodes correspond to a physical vertex, GNN decoder 217 may determine node probabilities 808 for each node. In other words, each node may be associated with a value indicating the probability of that particular node actually representing a vertex of a building.

[0107] FIG. 9 illustrates an exemplary rendering of a preliminary semantic geometry 900 from the outputs of a GNN decoder. As illustrated, and based on the plurality of nodes and their associated node locations generated by GNN decoder 217, preliminary semantic geometry 900 may be generated with a plurality of nodes including first nodes 904 and second nodes 908. As further illustrated, first nodes 904 may be identified as corresponding to physical vertexes and included in the final semantic geometry for the building while second nodes 908 may be excluded. Determining that first nodes 904 should be included and second nodes should be excluded may be based on node probabilities 808 associated with each individual node. For example, each probability may be compared to one or more predefined threshold criteria to determine whether the respective node should be included or excluded in the final semantic geometry.

[0108] As described above, the edges of a semantic geometry may represent physical edges of a building that separate adjacent vertexes. Accordingly, and referring back to FIG. 8, GNN decoder 217 may compute adjacency probabilities 812 to determine which of the plurality of nodes are connected by an edge. Adjacency probabilities 812 may be determined for each unique pair of nodes. The adjacency probability for a pair of nodes may represent the likelihood that an edge connecting the pair of nodes will represent a physical edge of a building. In some embodiments, identifying the final set of edges for a semantic geometryincludes identifying pairs of nodes where both nodes have been selected for inclusion in the final geometry (e.g., based on node probabilities 808), and the adjacency probability for the pair meets one or more predefined threshold criteria.

[0109] Referring again to FIG. 9, semantic geometry 900 may include a plurality of edges connecting pairs of nodes. For ease of understanding, a subset of the possible edges has been included in the illustration of semantic geometry 900. However, it should be understood that a possible edge may be created between each unique pair of nodes. As illustrated, eave edges 912, rake edges 916, and ridge edge 920 may be included in the final semantic geometry while erroneous edges 924 may be excluded based on the adjacency probabilities between the nodes each of them connect.

[0110] Finally, referring back to FIG. 8, GNN decoder 217 may generate edge class probabilities 816 for the edges. Edge class probabilities 816 may represent the likelihood of a given edge representing a predefined physical edge type. For example, when generating semantic geometries for a roof, edges may be classified as a ridge, a rake, an eave, a hip, a valley, or other similar edge types. As another example, when generating semantic geometries for an interior space, edges may be classified as a wall line, a floor line, a ceiling line, or the like. In some embodiments, GNN decoder 217 determines an edge class probability for each edge class that a particular edge may represent. Additionally, or alternatively, GNN decoder 217 may generate a single classification and provide a probability of that classification being correct.[oni] In some embodiments, edge class probabilities 816 are based at least in part the surrounding geometries of an edge. For example, and as illustrated in FIG. 9, edges that are substantially parallel with a ground plane (e.g., horizontal), such as eave edges 912 and ridge edge 920, may be more likely to be classified as eaves or ridges while edges that nonhorizontal, such as rake edges 916, may be more likely to be classified as hips, valleys, or rakes. As another example, an edge extending from a node with edges extending from the node on both sides may be more likely to be classified as hips or valleys.

[0112] Referring back to FIG. 7, the process may proceed with loss module 216 determining a loss between predicted silhouettes 708 and either or both of silhouettes of inferred representations 716 and semantic geometries 720. In some embodiments, a loss is computed for each silhouette of predicted silhouettes 708. For example, based on the number of images 704 provided to image processor 211, each with a unique perspective or pose, thesame number of predicted silhouettes 708 may be produced by segmentation engine 212. Loss module 216 may then determine a loss between each silhouette of predicted silhouettes 708 and an extracted silhouette from inferred representations 716 from a same or similar perspective or pose. Based on the various losses, loss module 216 may compute a combined loss for the current instantiation of latent code 712.

[0113] The process may continue with latent space module 213 updating latent code 712 based on the loss from loss module 216 and / or probability distribution 324. While described as updating latent code 712, latent space module 213 may additionally, or alternatively, select a new latent code from within the latent space to provide to INF decoder 214 and / or GNN decoder 217 to generate a new set of inferred representations 716 and semantic geometries 720. The loss from loss module 216 may represent, or otherwise be used to derive, a gradient direction in which latent space module 213 should move within the latent space to update latent code 712. For example, based on the loss, latent space module 213 may use Gradient Decent (GD) to iteratively update latent code 712, reducing the loss between the inferred silhouettes and the predicted silhouettes at each iteration until a minimum loss is achieved. Additionally, latent space module 213 may use probability distribution 324 to update latent code 712 between successive iterations to ensure that latent code 712 is updated along a path of latent codes in the latent space that will result in a likely building representation.

[0114] FIG. 10 illustrates exemplary iterations of comparisons between inferred silhouettes and semantic geometries generated from a latent code and a predicted silhouette generated from an image of a building. As described above, predicted silhouette 1002 may be generated (e.g., by segmentation engine 212) from one or more images of a building. As further described above, initial latent code 1004 may be sampled (e.g., by latent space module 213) at random from within the latent space and / or as guided by a probability distribution fit to the latent space. Initial latent code 1004 may then be used to generate initial inferred silhouette 1008 (e.g., through INF decoder 214) and initial semantic geometry 1012 (e.g., by GNN decoder 217). Based on the loss between initial inferred silhouette 1008 and / or initial semantic geometry 1012, initial latent code 1004 may be modified, or a new latent code may be selected, to produce updated latent code 1016 and its associated updated inferred silhouette 1020 and / or updated semantic geometry 1024. Updated latent code 1016 may be iteratively updated, as described above, until the loss between final inferred silhouette 1032 and / or final semantic geometry 1024 generated from final latent code 1028 and predicted silhouette 1002 meets one or more predefined threshold criteria. While described above asgenerating a new semantic geometry with each iteration of latent code 712, some embodiments may only generate a semantic geometry once a latent code is identified that results in a minimum loss between the inferred silhouettes for the latent code and the predicted silhouettes for a building.

[0115] FIG. 11 illustrates exemplary movement within the latent space between iterations of an INF decoder. As described above, a probability distribution, such as probability distribution 1104, fit to the learned latent space may be used to update latent codes between iterations of an INF decoder, such as INF decoder 214. As illustrated, probability distribution 1104 may include one or more areas or regions within the latent space from which latent codes are more likely to represent building structures and / or architecture matching those captured in images 704. Additionally, or alternatively, probability distribution 114 may include one or more areas or regions within the latent space from which latent codes represent plausible building structures that are more likely to represent existing or common building structures. Based on these areas or regions, latent codes may be iteratively updated, or a new latent code selected, such that the resulting representations are more likely to represent a plausible and / or common building structure, such as may be observed in images 704.

[0116] For example, and as further illustrated, an initial latent code selected from first region 1112 of the latent space may be iteratively updated along direction 1116 toward second region 1124, based on loss between inferred representations of the latent codes in region 1112 and observed data of received images those inferred representations are compared to. As a result, and with each iteration, the inferred silhouettes may gradually transition from silhouettes similar to first silhouette 1108 toward silhouettes similar to second silhouette 1120. Subsequently, the latent code may be iteratively updated along direction 1128 toward third region 1130 until a loss between final silhouette 1132 and a predicted silhouette meets one or more predefined threshold criteria.

[0117] Referring back to FIG. 7, once a minimum loss has been achieved by latent code 712, semantic geometries 720 generated by GNN decoder 217 may be output as the requested semantic geometry. For example, semantic geometries 720 may be transmitted to the user device that provided images 704. A client application may then render semantic geometries 720 on a display of the user device, or a semantic spatial graph for the building captured in the images 704 saved to house model database 218. For example, semantic geometries 720may be rendered superimposed with images 704 to demonstrate the fit between semantic geometries 720 and images 704. As another example, semantic geometries 720 may be displayed alone with one or more measurements for the building depicted in images 704, such as the lengths of edges, angles between adjacent edges, angles between an edge and one or more axes, area measurements for surfaces enclosed by edges, or the like.

[0118] In some embodiments, INF decoder 214 and / or GNN decoder 217 modify semantic geometries 720 before superimposing them onto images 704. For example, INF decoder 214 and / or GNN decoder 217 may determine that one or more portions of a semantic geometry would not be visible from a particular perspective or pose of an image onto which the semantic geometry would be superimposed. As such, these portions may be occluded or otherwise made transparent when rendering the semantic geometry onto the image.

[0119] FIG. 12 is a flowchart illustrating an example of a process 1200 for generating a semantic geometry for a building from a digital image of the building. Process 1200 may be performed, at least in part, by any of the components illustrated in FIGS. 1-2, such as user device 110 and server 120. Process 1200 may begin at block 1210, where a digital image depicting a building is received. The digital image may be received from a user device, such as user device 110. For example, a digital camera integrated into and / or coupled with user device 110 may capture the digital image. The digital image may be received by an image processor, such as image processor 211. The digital image may include a set of pixels that depict the building observed from a first perspective. For example, the set of pixels may depict the building from a front, a side, a rear, or the like. In some embodiments, the building may be a residential house. While described as a single digital image, block 1210 may optionally include receiving multiple images depicting the building. For example, the digital image may be received with one or more additional images depicting the building from different perspectives or poses.

[0120] At block 1220, the digital image may be transformed into a predicted silhouette, or other semantically meaningful representation, of the building. For example, an image segmentation engine, such as segmentation engine 212, may process the set of pixels in the digital image depicting the building to generate a silhouette for the building. As described above, the silhouette for the building may be a 2D segmentation map, a collection of segmentation masks, a 2D point cloud, a bounding polygon, or the like, that represents the physical contours of the building depicted in the digital image. As described above, morethan one digital image depicting the building may be received at block 1210. Accordingly, block 1220 may optionally include transforming each additional digital image depicting the building into a separate predicted silhouette.

[0121] At block 1230, a latent code that corresponds to a point in a learned latent space, comprising known latent codes and additional latent codes generated from the known latent codes, may be sampled. As described above, the learned latent space may be defined by a plurality of latent variables with predefined number of values, such that each point, referred to as a latent code, in the learned latent space may represent a unique combination of values for each of the plurality of latent variable. As further described above, each point in the latent space may represent a unique combination of hypothetical building parameters in the real world that could be used to define the physical attributes for a hypothetical building. In some embodiments, the latent code is sampled at random. Additionally, or alternatively, the latent code may be sampled based on a probability distribution fit to the latent space, as described further above. For example, the latent code may be generated based on the latent code in the latent space determined to correspond to the most likely building structure based on a plurality of known models from which the latent space is learned.

[0122] At block 1240, the latent code may be decoded into an inferred representation for a hypothetical building, such as a density rendering. In some embodiments, the latent code is decoded by a trained INF decoder, such as INF decoder 214. In some embodiments, a semantically meaningful asset is extracted from the inferred representation, such as a silhouette. As described above, the INF decoder may be trained on a plurality of known building models to infer representations of a building from a predefined perspective using a latent code within the latent space. As further described above, the extracted semantically meaningful asset from the inferred representation may be a 2D segmentation map, a collection of segmentation masks, a 2D point cloud, a bounding polygon, or the like, that represents the physical contours of the inferred building from a same or similar perspective as the digital image depicting the building.

[0123] At block 1250, a loss between the predicted silhouette of the building and the inferred silhouette for the hypothetical building may be determined. As described above, the loss may represent the differences between the predicted silhouette and the inferred silhouette. Visually, the loss may represent how closely the contours of the building in the predicted silhouette align with the contours of the hypothetical building in the inferredsilhouette when projected onto each other (e.g., from the same perspective). In other terms, the loss may represent the probability that the inferred silhouette accurately represents the predicted silhouette. As further described above, the loss may be based on a Normalized Object Coordinate Space (NOCS) loss, a mask loss, a depth loss, a normal consistency loss, an eik loss, a cross-entropy loss, and a Generative Adversarial Network (GAN) loss.

[0124] At block 1260, it may be determined whether a convergence between the predicted silhouette and the inferred silhouette has been detected. As described herein, a convergence may be detected when the loss between the predicted silhouette and the inferred silhouette determined at block 1250 meets one or more predefined threshold criteria. For example, a convergence may be detected when a loss between a predicted silhouette and an inferred silhouette is less than a predefined loss value. As another example, a convergence may be detected when a latent code has been identified as resulting in the lowest possible loss compared to other latent codes within the latent space. If at block 1260, a convergence is detected, process 1200 may proceed to block 1280. On the other hand, if at block 1260, a convergence has not been detected, process 1200 may proceed to block 1270.

[0125] At block 1270, the latent code may be updated based on the loss between the predicted silhouette and the inferred silhouette. For example, based on the difference between the predicted silhouette and the inferred silhouette, the loss may indicate a direction within the latent space towards which an updated latent code selected in that direction will result in lesser amount of loss or cost. In some embodiments, updating the latent code includes updating a single latent variable of the latent code at a time. Additionally, or alternatively, the differences between successive updates may be greater or larger based on the loss. For example, a greater amount of loss between the predicted silhouette and the inferred silhouette may result in a larger modification of the latent code. On the other hand, smaller losses, indicating that the desired latent code that will result in a minimal amount of loss is in close proximity to the current latent code, may result in smaller modifications to the latent code to ensure that the desired latent code is not passed.

[0126] In some embodiments, updating the latent code is additionally based on one or more probability distributions fit to the latent space. As described above, the probability distribution may be fit to the latent space based on distributions of known latent codes corresponding to known models, or buildings that exist in the real world. As such, updatingthe latent code based on the probability distribution, as well as the loss, may ensure that the updated latent code represents a likely building structure in the real world.

[0127] As illustrated, process 1200 may return to block 1240 after updating the latent code at block 1270 by decoding the updated latent code into a new inferred representation and extracting its silhouette. The loop including blocks 1240 through block 1270 may then be repeated any suitable number of times until it is determined at block 1260 that a convergence between the predicted silhouette and the latest inferred silhouette has been detected.

[0128] At block 1280, a semantic geometry for the building may be decoded from the latent code. As described above, a semantic geometry for a building may include a semantic spatial graph or wireframe representation of the edges and vertexes of the building. For example, a semantic spatial graph may include a plurality of nodes connected by edges where the edges represent the physical edges of a building and the nodes represent the points where edges start, intersect, or end. In addition, the semantic spatial graph may store one or more dimensions, measurements, or classifications for the edges in the graph. For example, each edge may include a length and angle relative to a reference, such as a gravity vector, another edge, a node separating the edge from other edges, or the like. Additionally, or alternatively, each edge may be assigned a classification representing a physical building edge type. For example, in the case of a semantic geometry for a roof, each edge may be classified as an eave, rake, ridge, valley, hip, or the like. As another example, in the case of a semantic geometry for a building interior, each edge may be classified as a wall line, floor line, ceiling line, or the like.

[0129] As described further below, the semantic geometry generated for a building depicted in a digital image may be used in one or more applications. For example, the semantic geometry may be rendered as an overlay on one or more images of the building from the same perspective to allow a user to evaluate the accuracy with which the semantic geometry reconstructs the physical features of the building. As another example, the semantic geometry may be used to determine and display one or more measurements, such as a total length of gutters (e.g., from combined lengths of the edges classified as eaves) that may be required for the building, a total surface area that may need to be covered by a construction material (e.g., shingles, paint, siding, etc.), or the like.

[0130] FIG. 13 is a flowchart illustrating an example of a process 1300 for superimposing a semantic geometry of a house’ s roof over an image of the house for display on a mobiledevice. Process 1300 may be performed, at least in part, by any of the components illustrated in FIGS. 1-2, such as user device 110 and server 120. Process 1300 may begin at block 1310, where one or more instructions for capturing one or more digital images of an exterior of a house are displayed in a user interface of a mobile device. The mobile device may be a user device, such as a smart phone, tablet, digital camera, or the like with an integrated display and one or more input mechanisms, such as a touch screen. The one or more instructions may include instructions for how a user of the mobile device can capture the one or more digital images using a camera of the mobile device and / or access the one or more digital images through a file and / or photo browser of the mobile device. For example, the user interface may include a selectable option to capture the one or more digital images using the camera of the mobile device. Additionally, or alternatively, the one or more instructions may describe or guide a user through the process and / or requirements for framing the exterior of the house within the digital images to minimize the capture of background or foreground details while ensuring that the contours of the building are detectable from the one or more images.

[0131] At block 1320, the one or more digital images captured by the mobile device may be received via the user interface in response to the one or more instructions. For example, the one or more digital images may be received in response to selections of a selectable option to capture an image using a camera integrated in, or otherwise connected to, the mobile device. As another example, the one or more digital images may be received in response to selections of the one or more images through a file and / or photo browser functionality of the mobile device.

[0132] At block 1330, a semantic spatial graph of the roof of the house may be generated by the mobile device using the one or more digital images. In some embodiments, the mobile device generates the semantic spatial graph of the roof of the house by executing process 1200 described above. Additionally, or alternatively, the mobile device may transmit the one or more digital images to a remote server, such as server 120, for execution of process 1200 described above. In response, the remote server may transmit the semantic spatial graph to the mobile device for subsequent processing.

[0133] At block 1340, an image of the house from a first perspective may be displayed via the user interface. The image of the house may be selected from the one or more images by an application executing on the mobile device. For example, the application may select a first image of the house taken from a particular perspective for initial display, such as a frontperspective view of the house. Additionally, or alternatively, a user of the mobile device may select the image from the one or more images via a user interface displayed by mobile device.

[0134] At block 1350, one or more portions of the semantic spatial graph that are visible in the image may be determined by the mobile device based on the first perspective. As described above, while a semantic spatial graph of a building may represent the 3D contours of the building, when depicted in a 2D image, various structures or portions of the 3D representation may not be visible in the 2D image. For example, given a 2D image depicting a house taken from a front of the house, physical features on the back of the house may not be visible in the 2D image. As such, based on the perspective of the image displayed at block 1340, portions of the semantic spatial graph that are and are not visible in the 2D image may be determined. An INF decoder, such as INF decoder 214, may be used to make such determinations. For example, by providing the latent code used to generate the semantic spatial graph to the INF decoder along with the perspective of the image and the coordinates of a point represented by the semantic spatial graph, the INF decoder may determine whether the point would be occluded by other structures of the building. Based on the portions determined to be occluded, the remaining portions of the semantic spatial graph may then be determined to be visible from the perspective of the 2D image.

[0135] At block 1360, the one or more visible portions of the semantic spatial graph may be displayed via the user interface overlaying the image of the house. For example, the image of the house may be displayed with one or more lines overlaying the visible edges of the roof. In some embodiments, the one or more lines are distinguished from each other based on the underlying classification for the edge used to render a particular line. For example, lines corresponding to eaves may be rendered with a first color and / or weight while lines corresponding to a ridge may be rendered using a second color and / or weight. Additional visual distinctions may be used to highlight various visible features of the semantic spatial graph, such as textual representations of edge or surface dimensions, as described further below.

[0136] FIG. 14 is a flowchart illustrating an example of a process 1400 for superimposing measurements for a house’ s roof over an image of the house for display on a mobile device. Process 1400 may be performed, at least in part, by any of the components illustrated in FIGS. 1-2, such as user device 110 and server 120. Process 1400 may begin at block 1410, where one or more instructions for capturing one or more digital images of an exterior of ahouse are displayed in a user interface of a mobile device. The mobile device may be a user device, such as a smart phone, tablet, digital camera, or the like with an integrated display and one or more input mechanisms, such as a touch screen. The one or more instructions may include instructions for how a user of the mobile device can capture the one or more digital images using a camera of the mobile device and / or access the one or more digital images through a file and / or photo browser of the mobile device. For example, the user interface may include a selectable option to capture the one or more digital images using the camera of the mobile device. Additionally, or alternatively, the one or more instructions may describe or guide a user through the process and / or requirements for framing the exterior of the house within the digital images to minimize the capture of background or foreground details while ensuring that the contours of the building are detectable from the one or more images.

[0137] At block 1420, the one or more digital images captured by the mobile device may be received via the user interface in response to the one or more instructions. For example, the one or more digital images may be received in response to selections of a selectable option to capture an image using a camera integrated in, or otherwise connected to, the mobile device. As another example, the one or more digital images may be received in response to selections of the one or more images through a file and / or photo browser functionality of the mobile device.

[0138] At block 1430, a semantic spatial graph of the roof of the house may be generated by the mobile device using the one or more digital images. In some embodiments, the mobile device generates the semantic spatial graph of the roof of the house by executing process 1200 described above. Additionally, or alternatively, the mobile device may transmit the one or more digital images to a remote server, such as server 120, for execution of process 1200 described above. In response, the remote server may transmit the semantic spatial graph to the mobile device for subsequent processing.

[0139] At block 1440, area measurements for each face of the roof may be determined by the mobile device using the semantic spatial graph. A face of the roof may correspond to any planar surface of the roof. Faces may be bounded by three or more edges that form a closed polygon and are each coplanar with each other. In some embodiments, the faces of the roof are identified from the semantic spatial graph using one or more graph traversal algorithms to detect a plurality of closed polygons in the semantic spatial graph. Once the plurality of closed polygons are detected, the polygons formed by coplanar edges and / or vertices (e.g., asdetermined from coordinates associated with each node or edge) may be identified as corresponding to the faces of the roof. Additionally, or alternatively, a guided traversal algorithm may use the coordinates of the edges and / or nodes to efficiently traverse the semantic spatial graph. For example, after traversing two edges, the traversal may only proceed with subsequent edges and / or nodes that are coplanar with the first two edges. Additionally, or alternatively, selecting edges for traversal may be based on the semantic information contained within the semantic spatial graph, such as edge and / or node classifications. For example, based on the starting edge’s classification (e.g., an eave), an edge that is connected to the starting edge and has a different classification (e.g., a rake, a hip, or a valley) may be selected for the second edge traversal. Subsequently, edges may be selected that are coplanar with the two starting edges. While described in reference to roofs, similar methods may be applied to semantic geometries for other types of structures, such as interior or exterior walls of a building. Once the polygons that correspond to the faces of the roof are identified, the surface area measurements may be calculated using the dimensions of the nodes and edges of each polygon, such as the edge lengths and angles between adjacent edges at each node.

[0140] At block 1450, one or more faces of the roof that are visible in an image of the house are determined by the mobile device. The user of the mobile device may select the image from the one or more images via a user interface displayed by mobile device. The one or more visible faces may be determined based at least in part on the perspective of the selected image. For example, by providing the latent code used to generate the semantic spatial graph to an INF decoder along with the perspective of the image and coordinates of one or more points on the plane formed by a corresponding polygon in the semantic spatial graph, the INF decoder may determine whether the surface would be totally or partially occluded or obstructed by other structures of the building between the vantage point of the image and the points on the surface. Based on the surfaces that would be partially or totally occluded, the remaining faces derived from the semantic spatial graph may be identified as the visible faces of the roof from the perspective of the 2D image.

[0141] At block 1460, the area measurements for the one or more visible faces are displayed via the user interface overlaying the image of the house. For example, the image of the house may be displayed with a text representation of each surface area measurement overlaying the corresponding face of the roof depicted in the image. Additionally, or alternatively, semi-transparent polygons may be displayed overlaying their correspondingroof faces to identify the regions covered by a surface area measurement. In this way, a user can quickly and easily evaluate the accuracy of a surface area measurement by verifying that the transparent polygon aligns with the visible face in the image.

[0142] The foregoing description of the embodiments, including illustrated embodiments, has been presented only for the purpose of illustration and description and is not intended to be exhaustive or limiting to the precise forms disclosed. Numerous modifications, adaptations, and uses thereof will be apparent to those skilled in the art.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method, comprising: receiving a digital image, wherein the digital image comprises a set of pixels that depict a building observed from a first perspective; transforming the set of pixels that depict the building into a first semantic representation; sampling a latent code that corresponds to a first point in a learned latent space, wherein points in the learned latent space represent a distribution of building structures according to encoded parameters; decoding the latent code into a second semantic representation according to the first perspective; determining a loss between the first semantic representation and the second semantic representation; updating the latent code to correspond to a second point in the learned latent space based on the loss; decoding the updated latent code into a third semantic representation according to the first perspective; detecting a convergence between the first semantic representation and the third semantic representation; and decoding the updated latent code into a semantic geometry for the building, wherein the semantic geometry represents classifications for one or more structural features of the building.

2. The computer-implemented method of claim 1, wherein the first semantic representation comprises a two-dimensional segmentation map.

3. The computer-implemented method of claim 1 , wherein the first semantic representation comprises a silhouette of the building.

4. The computer-implemented method of claim 1, wherein decoding the latent code into a second semantic representation further comprises providing an implicit neural representation of the latent code, wherein the implicit neural representation has a density volume.

5. The computer-implemented method of claim 4, further comprising providing a silhouette of the implicit neural representation.

6. The computer-implemented method of claim 1, wherein determining a loss between the first semantic representation and the second semantic representation comprises detecting a difference between a projection of the first semantic representation and the second semantic representation.

7. The computer-implemented method of claim 1, wherein updating the latent code comprises selecting a new latent code from the learned latent space.

8. The computer-implemented method of claim 1, wherein updating the latent code comprises modifying the sampled latent code.

9. The computer-implemented method of claim 1, wherein updating the latent code based on the loss comprises applying a gradient direction derived from the loss to the sampled latent code.

10. The computer-implemented method of claim 1, wherein updating the latent code based on the loss comprises traversing the latent space according to a distribution fit to the learned latent space.

11. The computer-implemented method of claim 1, wherein decoding the updated latent code into a third semantic representation further comprises providing an updated implicit neural representation of the updated latent code, wherein the implicit neural representation has a density volume.

12. The computer-implemented method of claim 11, further comprising providing an updated silhouette of the updated implicit neural representation.

13. The computer-implemented method of claim 1, wherein detecting a convergence between the first semantic representation and the third semantic representation comprises detecting a substantial similarity between a projection of the first semantic representation and the third semantic representation.

14. The computer-implemented method of claim 1, wherein detecting a convergence between the first semantic representation and the third semantic representationcomprises detecting a subsequent updated latent code does not produce a lower loss between a projection of the first semantic representation and the third semantic representation.

15. The computer-implemented method of claim 1, wherein decoding the updated latent code into the semantic geometry comprises executing a graph neural network on the updated latent code.

16. The computer-implemented method of claim 1, wherein the semantic geometry comprises a graph structure.

17. The computer-implemented method of claim 16, wherein the graph structure comprises: a plurality of edges that correspond to edges on a roof of the building; and a plurality of nodes that correspond to intersections between the edges on the roof of the building.

18. The computer-implemented method of claim 17, wherein decoding the updated latent code into the semantic geometry for the building comprises classifying each edge of the plurality of edges as one of a plurality of edge types.

19. A computer-implemented method, comprising: receiving a digital image, wherein the digital image comprises a set of pixels that depict a building observed from a first perspective; transforming the set of pixels that depict the building into a two-dimensional (2D) segmentation map; sampling a latent code from a learned latent space comprising a plurality of latent codes, wherein any one latent code may be sampled as a point from the learned latent space; updating the latent code through one or more iterations until a convergence is detected between the 2D segmentation map and an inferred segmentation map decoded from the latent code, wherein each iteration comprises: decoding the latent code into the inferred segmentation map for a hypothetical building structure according the first perspective; determining a loss between the 2D segmentation map and the inferred segmentation map, wherein the loss represents differences between the 2Dsegmentation map and the inferred segmentation map when projected onto each other; and modifying the latent code based on the loss; and decoding the modified latent code into a semantic geometry for the building, wherein the semantic geometry represents classifications for one or more structural features of the building.

20. A computer-implemented method for learning a latent code for parameters of a building geometry, comprising: providing a multi-dimensional representation of a building structure, the multidimensional representation having a parameterized geometry and image data associated with the building structure; generating a first semantic representation of the image data associated with the building structure; generating a vector of randomized values associated with the parameterized geometry for each building structure; inputting the vector into a first network to output a second semantic representation; determining a loss between the first semantic representation and the second semantic representation; and iteratively optimizing, based on the loss, the random values associated with the parameterized geometry of the vector and one or more weights of the first network until an output of the iteratively optimized first network executed upon an iteratively optimized vector converges with the first semantic representation.

21. The computer-implemented method of claim 20, further comprising training a second network by: providing a second network the iteratively optimized vector; generating a third semantic representation; determining a loss between the parameterized geometry of the multidimensional representation of the building structure and the third semantic representation; and iteratively optimizing, based on the loss, one or more weights of the second network until an output of the iteratively optimized second network executed upon theiteratively optimized vector converges with the parameterized geometry of the multidimensional representation of the building structure.

22. The computer-implemented method of claim 21, further comprising: generating a fourth semantic representation based on the image data associated with the building structure; inputting the iteratively optimized vector into the first network to output an updated second semantic representation and the second network to output an updated third semantic representation; determining an updated density loss between the first semantic representation and the updated second semantic representation; determining an updated semantic loss between the updated third semantic representation and the fourth semantic representation; and iteratively re-optimizing, based on the updated density loss and the updated semantic loss, one or more values associated with the parameterized geometry of the iteratively optimized vector and one or more weights of the first network and one or more weights of the second network until an output of the iteratively re-optimized first network and iteratively re-optimized second network executed upon an iteratively re-optimized vector converges with the first semantic representation and fourth semantic representation respectively.

23. The computer-implemented method of claim 20, wherein generating a first semantic representation comprises generating a silhouette of the building structure based on the image data.

24. The computer-implemented method of claim 20, wherein generating the vector of randomized values comprises generating a 512-dimensional vector.

25. The computer-implemented method of claim 20, wherein the first network is an implicit neural representation decoder network and the second semantic representation is an implicit neural representation having a volume density.

26. The computer-implemented method of claim 21, wherein the second network is a graph neural network and the third semantic representation is a semantic spatial graph of geometry inferred from the provided vector.

27. The computer-implemented method of claim 22, wherein the fourth semantic representation is a two-dimensional semantic segmentation of the image data.

28. The computer-implemented method of claim 20, further comprising: learning a latent code for a plurality of multi-dimensional building structures based on the iteratively optimized vector; and arranging a latent space comprising the plurality of known latent codes associated with a plurality of multi-dimensional building structures.

29. The computer-implemented method of claim 28, further comprising fitting a probability distribution to the plurality of known latent codes, wherein the probability distribution comprises additional points in the latent space corresponding to latent codes for hypothetical building structures among the known latent codes associated with the plurality of multi-dimensional building structures.

30. The computer-implemented method of claim 20, wherein the multidimensional representation is a three-dimensional geometrical model.

31. The computer-implemented method of claim 20, wherein the multidimensional representation is a two-dimensional floorplan.

32. A system comprising: a database comprising a plurality of multi-dimensional building models, each multi-dimensional building model having a parameterized geometry and image data associated with the building structure; an encoder network capable of generating encoded values for the parameterized geometry into a latent code, wherein the latent code is a multidimensional vector associated with a respective multi-dimensional model and the encoder network if further capable of iteratively optimizing the encoded values based on losses provided by one or more decoder networks; an image segmentation network configured to provide one or more semantic outputs based on the image data; a first decoder network capable of decoding a provided latent code and generating a first semantic output for comparison with the one or more semantic outputs of the image segmentation network; anda second decoder network capable of decoding a provided latent code and generating a second semantic output for comparison with the parameterized geometry or one or more outputs of the image segmentation network.