Feature extraction via annotation of synthetic images generated by depth-guided novel view synthesis

By projecting two-dimensional annotations onto depth maps generated from synthetic images, the model addresses the challenge of converting two-dimensional data into accurate three-dimensional models, enhancing geospatial feature extraction efficiency and accuracy.

WO2026081007A1PCT designated stage Publication Date: 2026-04-231000786269 ONTARIO INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
1000786269 ONTARIO INC
Filing Date
2025-10-14
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for geospatial feature extraction, whether manual or automated, struggle to efficiently convert two-dimensional annotations into accurate three-dimensional models, particularly for large-scale projects.

Method used

A machine learning model combining neural radiance fields and depth-guided novel view synthesis is used to generate synthetic images with depth maps, allowing two-dimensional annotations to be projected onto three-dimensional structures, enabling automated three-dimensional feature extraction.

Benefits of technology

This approach enables highly automated and scalable three-dimensional semantic feature extraction, suitable for large-scale geospatial projects, with improved accuracy and efficiency by leveraging depth maps for three-dimensional positional information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025051352_23042026_PF_FP_ABST
    Figure CA2025051352_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for generating three-dimensional models of target features based on synthetic imagery are provided. An example method involves accessing source images that depict a target three-dimensional feature from multiple points of view, encoding each of at least some of the source images each into one or more corresponding feature maps, decoding a synthetic image based on a target view and the feature maps into which the source images were encoded, wherein decoding the synthetic image involves generating one or more depth maps corresponding to the target view, applying a machine learning model to the synthetic image to generate a set of two-dimensional annotations that represent the target three-dimensional feature as depicted in the synthetic image, projecting the set of two-dimensional annotations onto at least one of the depth maps, and reconstructing a three-dimensional model of the target three-dimensional feature based on the projected two-dimensional annotations.
Need to check novelty before this filing date? Find Prior Art

Description

FEATURE EXTRACTION VIA ANNOTATION OF SYNTHETIC IMAGES GENERATED BY DEPTH-GUIDED NOVEL VIEW SYNTHESISBACKGROUND

[0001] Vector maps can be manually extracted from imagery using software platforms that allow individuals to manually annotate images through a user interface. A common use case is the annotation of aerial, satellite, street-view, and other imagery, to produce two-dimensional and three-dimensional vector maps of landcover and land use features, such as road and building features. Automated approaches to geospatial feature extraction that involve the application of machine learning models have been proposed.

[0002] Neural radiance fields (NeRFs) are a recent development in the field of novel view synthesis. The approach generally involves training deep neural networks to model the visual radiance of scenes as continuous functions in three-dimensional space. Neural radiance fields can be used to synthesize highly realistic images of scenes from arbitrary viewpoints.BRIEF DESCRIPTION OF THE DRAWINGS

[0003] FIG. 1 is a schematic diagram of an example machine learning model for novel view synthesis that involves a local attention guidance process.

[0004] FIG. 2 is a schematic diagram of an example machine learning model for novel view synthesis, similar to that of FIG. 1 , shown in greater detail, wherein the local attention guidance process involves a depth map feature projection process.

[0005] FIG. 3 is an illustration of an example depth map feature projection process for guiding local attention, for use in a machine learning model for novel view synthesis, such as that shown in FIG. 2.

[0006] FIG. 4 is a schematic diagram of an example machine learning model for feature extraction that operates by generating sequences of annotation tokens that represent target features to be extracted from source images.

[0007] FIG. 5 is a schematic diagram of an example system for generating geometric models that represent target features to be extracted from source images, which includes a machine learning model similar to that of FIG. 4, shown in greater detail.

[0008] FIG. 6 is a schematic diagram of an example system for generating a three- dimensional geometric model that represents a three-dimensional target feature to be extracted from source imagery, which operates based on a machine learning model for novel view synthesis, similar to described in FIG. 1 to FIG. 3, in combination with a machine learning model for feature extraction, similar to those described in FIG. 4 to FIG. 5.

[0009] FIG. 7A is an example illustration of how a set of two-dimensional annotations of a three-dimensional target feature extracted from a two-dimensional synthetic image can be projected onto a three-dimensional depth map that corresponds to the synthetic image to obtain the three-dimensional structure of the three-dimensional target feature, as shown in the example use case where the three-dimensional target feature is a pitched roof structure. FIG. 7B is an example illustration similar to that of FIG. 7A in which the three-dimensional target feature is an exterior side wall of a building.

[0010] FIG. 8 is a flowchart of an example method for generating a three-dimensional geometric model that represents a three-dimensional target feature to be extracted from source imagery.

[0011] FIG. 9 is a schematic diagram of an example system for capturing source imagery of a three-dimensional target feature and generating three-dimensional model data that represents the three-dimensional target feature, with example hardware components described for illustrative purposes.

[0012] FIG. 10 is another example system for capturing source imagery of a three- dimensional target feature and generating three-dimensional model data that represents the three-dimensional target feature, similar to that of FIG. 9, with the addition of at least semimanual user image annotation.DETAILED DESCRIPTION

[0013] As described above, manual image annotation can be used to perform two- dimensional and three-dimensional feature extraction based on source imagery that can include aerial, satellite, street-view, and other imagery. Automated approaches to geospatial feature extraction that involve the application of machine learning models have been proposed, including in U.S. Patent Application No. 17 / 731 ,769, filed April 28th, 2022, entitled MACHINE LEARNING FOR VECTOR MAP GENERATION (the 769 Application, incorporated herein by reference in its entirety). The 769 Application describes how an autoregressive machine learning model with an attention mechanism similar to the transformer model can be trained to produce sequences of annotation operations that mimic the way a human annotator would perform annotation operations for feature extraction from single images.

[0014] As also described above, neural radiance fields can be used to synthesize highly realistic images of scenes from arbitrary viewpoints. Another previous disclosure, U.S. Patent Application No. 18 / 769,041 , filed July 10th, 2024, entitled GENERALIZABLE NOVEL VIEW SYNTHESIS GUIDED BY LOCAL ATTENTION MECHANISM (the ‘041 Application, incorporated herein by reference in its entirety), describes how a machine learning model can be trained to generate synthetic images of a scene from arbitrary points of view, wherein eachsynthetic image is generated along with a depth map that contains three-dimensional positional information about the scene and which aids in the image synthesis process.

[0015] The present disclosure combines and extends these two previous disclosures for the use case of three-dimensional feature extraction. In particular, this disclosure proposes that the depth maps that are generated as a biproduct of the image synthesis process described in the ‘041 Application can be leveraged for the three-dimensional positional information therein to enable three-dimensional feature extraction even while using the two-dimensional feature extraction approach described in the 769 Application. In particular, the proposal is that the two- dimensional annotations extracted from synthetic images can be projected onto the corresponding depth maps to obtain the three-dimensional structure of the extracted features.

[0016] Following the initial capture of source imagery, this approach can enable a highly- automated machine learning pipeline through which three-dimensional semantic feature information can be extracted from imagery at scale in a way that may be particularly well-suited for large-scale feature extraction projects in the geospatial industry. The annotation process can also involve a team of human image annotators using an image annotation system for training and quality control purposes.

[0017] The ‘041 Application proposes an encoder-decoder architecture that is generalizable across scenes and which relies solely on the 2D neural features encoded from source images to decode a target image from a novel viewpoint. In this approach, three-dimensional detail is captured at the decoder through a series of attention mechanisms that includes a local attention mechanism, which involves generating depth maps that aid in the image decoding process. Although the ‘041 Application is incorporated in its entirety herein by reference, certain elements of that disclosure are recast here, for greater understanding, throughout FIG. 1 to FIG. 3.

[0018] As shown in FIG. 1 , a machine learning model 100 for novel view synthesis comprises an encoder 110 and a decoder 120 in a deep learning architecture. The decoder 120 applies attention over the encoder 110 through at least one global attention mechanism 112 and one or more local attention mechanisms 114, which will be described in greater detail below.

[0019] The encoder 110 is configured to process a plurality of source images 102 from multiple points of view, which depict some arbitrary scene to be modeled, to generate a series of multiscale feature maps for each of the source images 102. A reference to the source images 102 may include a reference to the raw pixel data (e.g., RGB) as well as the metadata thereof, such as camera parameters (e.g., focal length, lens distortion, camera pose, resolution), geospatial projection information (e.g., latitude and longitude position), or other data. The source images 102 may contain one or several batches of imagery covering the area, and may have been captured on the same date, or on different dates.

[0020] The encoder 110 comprises some series of encoder layers, omitted here for simplicity, that progressively encode each source image 102 into a series of intermediate feature representations, which tend to increase in number of channels and decrease in resolution, or, in other words, a series of multiscale feature maps. For example, if a source image 102 is an aerial image captured at native ground resolution 0.25m, the series of multiscale feature maps may correspond to features extracted at 0.25m, 0.5m, and 1.0m resolutions. The series of multiscale feature maps for each source image 102 culminates in what will be referred to herein as a final feature map for the source image 102. A series of multiscale feature maps is generated for each source image 102.

[0021] The series of encoder layers may include one or more convolutional layers, one or more self-attention layers, one or more feed-forward neural layers, or any other suitable encoding layers capable of extracting and encoding key features from the source images 102, with downsampling layers as appropriate. The encoder 110 may further include an embedding component that embeds the camera parameters for the source images 102. These camera parameters refer to the parameters used in a camera model to describe the mathematical relationship between the 3D coordinates of a point in the scene to the 2D coordinates of its projection onto an image plane (whether according to a pinhole camera model, pushbroom camera model, fisheye camera model, orthographic projection, or other camera model). Further, it should be understood that one or more blocks of such components may be arranged in a deep learning architecture. A more detailed description of one example architecture is described in FIG. 2, further below.

[0022] Turning to the decoder 120, the decoder 120 is configured to decode a representation of a target view 104, which describes some arbitrary pose relative to the scene, into a target image 106 of the scene. The decoder 120 comprises some series of decoder layers, omitted here for simplicity, that progressively decode a representation of the target view 104 into a series of intermediate feature representations, which tend to decrease in the number of channels and increase in resolution, and which may be referred to as a series of multiscale feature maps. In keeping with the above example of an aerial image captured at native ground resolution 0.25m, the series of multiscale feature maps may correspond to features decoded at 1 ,0m, 0.5m, and 0.25m resolutions. The series of multiscale feature maps for each target view 104 culminates in what will be referred to herein as a final feature map, from which the pixel colors for the target image 106 can be determined (e.g., by some final activation function).

[0023] It should also be noted that the camera parameters that define the target view 104 need not necessarily match the same camera model used to capture the source images 102. For example, the source images 102 may have been captured through a fisheye camera model, whereas the target view 104 may call for a pinhole camera model or an orthographic projection.Therefore, the machine learning model 100 may be used to generate target images 106 with a preferred camera model for the scene.

[0024] The series of decoding layers may include one or more convolutional layers, one or more self-attention layers, one or more feed-forward neural layers, or any other suitable decoding layers capable of decoding key features for the target image 106, with upsampling layers as appropriate. The decoder 120 may further include an embedding component that embeds camera parameters for the target view.

[0025] Notably, the decoder 120 includes at least one global attention layer, which is depicted here as engaging in the global attention mechanism 112, and one or more local attention layers, each of which engages in a respective local attention mechanism 114. It is through these cross-attention mechanisms that the features of the target image 106 are progressively decoded by attending to feature information encoded from the source images 102.

[0026] As mentioned above, the decoder 120 applies attention over the encoder 110 through at least one global attention mechanism 112 and one or more local attention mechanisms 114. The global attention mechanism 112 is applied at or near the top of the decoder 120, to attend to the higher-level (i.e., lower resolution) features at the encoder 110, where an attention calculation is relatively inexpensive. Since global attention is applied at the top of the decoder 120, the global attention mechanism 112 applies attention over each feature of the final encoded feature map for each source image 102.

[0027] However, further down the decoder 120, to attend to the lower-level (i.e., higher resolution) features at the encoder 110, where attention calculations are more expensive, the decoder 120 applies a form of local attention, indicated here as the local attention mechanisms 114. These local attention mechanisms 114 are guided or assisted by a local attention guidance process 130 which incorporates an understanding of the spatial (i.e., topographical or geometric) features of the scene. One example of such a local attention guidance process 130 is one which involves back-projecting pixel information through the target view to determine the relevant feature information for use in the image decoding process, as described in FIG. 3, further below.

[0028] In terms of training, the machine learning model 100 may be trained on a dataset comprising a plurality of sets of source images 102 depicting a plurality of scenes, for generalizability. In terms of the objective function, the machine learning model 100, including the form of the local attention guidance process 130 that involves depth map pixel re-projection, may be trained solely on image loss. The training process would typically involve selecting some of the images of each scene to serve as the ground truth images against which the synthesized images are measured to determine image loss. Therefore, the machine learningmodel 100 can be trained end-to-end in a self-supervised manner without the need for annotated training data. Furthermore, the encoder 110 and decoder 120 may thereby learn to encode the features of the source images 102 and decode a target image 106 without being tied to the structure of any given scene. The depth map generation and re-projection process need not be trained separately, and may be learned implicitly as part of decoding the target image 106, without the need for annotated training data.

[0029] The machine learning model 100, including the trained learned neural network weights, biases, activation functions, and other architectural components and functionality, may be embodied in non-transitory machine-readable programming instructions, and executable by one or more processors of one or more computing devices, which include memory to store programming instructions that embody the functionality described herein and one or more processor to execute the programming instructions.

[0030] In FIG. 2, a machine learning model 200, which may be understood to be an example of the machine learning model 100 shown in greater detail, comprises an encoder 210 and a decoder 220 in a deep learning architecture.

[0031] As in FIG. 1 , the encoder 210 of FIG. 2 is configured to process each source image 202 through a series of convolutional layers 212 into a series of progressively encoded feature maps at multiple scales. In the present example, the encoder 210 includes three sets of convolutional and downsampling layers (indicated as conv / down layers 212), which produce two intermediate encoded feature maps 214C and 214B followed by a final encoded feature map 214A. For example, if the source images 202 include aerial images captured at native ground resolution 0.25m, the features 214C may correspond to the highest-resolution 0.25m features, whereas the features 214B may correspond to the lower-resolution 0.5m features, and the features 214C may correspond to the lowest-resolution 1.0m features. It should be noted that each convolutional layer may be applied in accordance with any known techniques, including the use of several convolutional layers of varying kernel size, and that each downsampling layer may be applied in accordance with any known techniques.

[0032] It should also be noted that whereas the raw pixel data of the source images 202, indicated here as RGB input 201 , flow through the conv / down layers 212, the corresponding camera parameters for each source image 202, indicated here as source image camera parameters 203, are embedded and used to form part of the key matrix in the following attention calculations. Each set of source image camera parameters 203 is passed through embedding layer 205, resulting in an embedded representation of each corresponding source view, indicated here as Pi0, where / denotes a source image 202, and where 0 indicates that Pi0represents the camera parameter embeddings added at the final stage of encoding. For each source image 202, the camera parameter embedding Pi0is concatenated with its corresponding final encoded feature map 214A to form a key matrix, indicated here as K°, which will be used inthe global attention calculation at the decoder 220, as described further below. The embedding layer 205 may be specialized to the type of camera model being used (e.g., pinhole, pushbroom, fisheye, orthographic), or may be generalized for any camera model.

[0033] Turning to the decoder 220, as above, the decoder 220 is configured to decode a representation of a target view, which describes some arbitrary pose and projection parameters relative to the scene (i.e., camera parameters), to decode a target image 206 of the scene. The decoder 220 decodes this target view through a series of convolutional and upsampling layers situated between attention layers. In keeping with the above example of the source images 202 including aerial images captured at native ground resolution 0.25m, the decoded features may be progressively decoded through 1.0m, 0.5m, and 0.25m resolutions.

[0034] In the present example, the representation of the target view comprises a set of target view camera parameters 204, passed through embedding layer 205 (similar or identical to the embedding layer 205 at the encoder 210), resulting in a camera parameter embedding Px° where x denotes the target view. The resulting camera parameter embedding Px° is concatenated with a learnable parameter 207, to form a query Q°, that will be used in the global attention calculation, as described below.

[0035] The first layer of the decoder 220 comprises a global attention layer 222. The global attention layer 222 computes attention based on Q°, derived from the target view as described above, and K°, derived from the source images 202 as described earlier in this disclosure. Attention may be computed in any suitable manner, such as performing scaled dot product between cross-attention values (see, e.g., Attention may be computed as described in Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017)). At this stage, the global attention layer 222 attends to all of the features of all of the final encoded feature maps 214A of all of the source images 202 to produce an initial decoded feature map 224A (e.g., 1.0m resolution features). At this stage in the decoding process, a global attention calculation is relatively inexpensive, and therefore is justifiable so that the target image 206 can be decoded with the benefit of attention applied globally across the source images 202.

[0036] Following the global attention layer 222, the initial decoded feature map 224A is further processed by a first convolutional layer and upsampling layer (indicated as conv / up 226- 1) to produce a further decoded feature map 224B (e.g., 0.5m resolution features). The feature map 224B is further processed by a first local attention layer 228-1 . The first local attention layer 228-1 computes attention based on Q1, a query derived from concatenating the previously decoded feature map 224B with an embedded representation of the target view, indicated here as Px1, and with a key matrix denoted as K1, which corresponds to a limited set of features selected from the encoder 210, concatenated with an embedded representation of thecorresponding source view. The limited set of features to which the first local attention layer 228-1 attends is determined by a depth map feature projection process 230, described below.

[0037] It should be noted at this stage that each successive query (i.e., Q°, Q1, Q2) may be derived based on the original query and upsampled to the increased resolution of the larger dimension feature map. For example, Q1can be derived by upsampling Q° for feature map 224B. Conversely, each successive key (i.e., K°, K1, K2) at the encoder 210 matches the resolution of the corresponding feature map.

[0038] The depth map feature projection process 230 involves generating a depth map at an intermediate stage of the decoding process, and using the depth map to capture three- dimensional information, and to narrow the set of features involved in the attention calculations at the decoder 220. The depth map feature projection process 230 involves three major steps for determining the limited set of features to which a local attention layer 228 attends. Applying the depth map feature projection process does not alter the scale of the decoded features, but rather, fills in the detail (with reference to the source image feature maps) missing from the previously upsampled feature map.

[0039] First, at step 232, a depth map for the scene is predicted based on an intermediate decoded feature map that precedes the local attention layer. For example, in the case of determining the limited set of features to which the local attention layer 228-1 attends, the immediately preceding feature map 224B is used to generate a depth map. The depth map may be generated by any suitable technique for generating depth maps based on a feature map extracted from an image. For example, the depth map may be generated through a series of convolutional layers and / or other neural layers. The resolution of the depth map may correspond to the scale of the features from which it was derived (e.g., 0.5m in the case of attention layer 228-1).

[0040] Second, at step 234, for each feature of the feature map used to predict the depth map, the corresponding point on the depth map is projected onto each intermediate encoded feature map, for each source image, of the corresponding scale (e.g., 0.5m). For example, in the case of local attention layer 228-1 , for each feature of the feature map 224B, a corresponding point on the depth map is projected onto the feature map 214B for each source image 202. In some cases, projection may fail, for example, where a point on the depth map cannot be geometrically projected to the image plane of a source image 202. In such a case, the attention calculation will be limited to the features of the remaining source images 202 to which projection is geometrically possible.

[0041] It should also be noted that, in some cases, the projection of a point on the depth map to a source image 202 may be occluded by another section of the depth map. In some implementations, this occlusion may be automatically detected, and the occluded source image202 may be excluded from the local attention calculation. However, in other implementations, it is expected that the model may learn to implicitly account for such occlusions, and to generate a low attention score for the features of the occluded source image 202, without the need for a separate process to handle the occlusion.

[0042] Third, at step 236, each feature on the source image feature maps to which a point on the depth map was projected is selected to be included in the limited set of features to which the local attention layer attends. For example, in the case of local attention layer 228-1 , each feature on the feature map 214B, of each source image 202, to which a point on the depth map was projected, is attended to by the local attention layer 228-1 . The resulting attention calculation produces the next intermediate decoded feature map 224BB.

[0043] This depth map feature projection process is illustrated for greater clarity with respect to a single pixel in FIG. 3. In FIG. 3, a depth map feature projection process 300 begins with a target pixel 312 of an intermediate decoded feature map 310 of any suitable resolution (e.g., 0.5m).

[0044] First, a depth map 320 is predicted based on the intermediate decoded feature map 310, by any suitable technique. The scale of the generated depth map 320 corresponds to the scale of the intermediate decoded feature map 310 (e.g., 0.5m). The point on the depth map 320 that corresponds to the pixel 312 of the intermediate decoded feature map is indicated as target point 322. The depth map 320 is illustrated in a monochromatic gradient to represent the depth of the surface structure of the scene. Areas of the depth map 320 that appear darker should be understood to be further away from the camera, whereas areas of the depth map 320 that appear lighter should be understood to be closer to the camera.

[0045] Next, the target point 322 is projected back to each source image 330 (i.e., source image 330-1 , 330-2, and 330-3), or, more precisely, back to the corresponding feature maps (e.g., 0.5m feature map) thereof. Thus, the target point 322 is projected to the intermediate encoded feature map 332-1 derived from the source image 330-1 , to the intermediate encoded feature map 332-2 derived from the source image 330-2, and to the intermediate encoded feature map 332-3 derived from source image 330-3.

[0046] Finally, a limited set of features selected from the intermediate encoded feature maps 332-1 , 332-2, and 332-3, based on the projection of the target point 322, is selected. This limited set of features will be used in the relevant local attention mechanism. In an extreme case, only the precise features onto which the target point 322 was projected may be included in the selection (indicated as target features 334-1 , 334-2, 334-3). However, this selection may be overly narrow, and may miss important information surrounding the feature to which the point 322 was projected. Preferably, a small selection of surrounding features may also be included in the attention calculation. For example, the set of directly adjacent features (indicated here asthe local features 336-1 , 336-2, and 336-3, for each source image 332-1 , 332-2, and 332-3, respectively) may also be included in the local attention calculation. In other cases, even larger sets of local features may be included (e.g., features within two or three spaces of the projected feature), as computational resources permit.

[0047] Although depicted for illustrative purposes for a single target pixel 312, it is to be understood that the depth map feature projection process 300 is to be repeated for each pixel of the intermediate decoded feature map 310.

[0048] Returning back to FIG. 2, the local attention layer 228-1 produces the further decoded feature map 224BB based on the previous decoded feature map 224B with attention to the limited set of features selected from the encoded feature map 214B, through a depth map feature projection process, as described above.

[0049] The feature map 224BB is further processed by a second convolutional layer and upsampling layer (indicated as conv / up 226-2) to produce feature map 224C, which is in turn processed by a second local attention layer 228-2. As with the first local attention layer 228-1 , the second local attention layer 228-2 computes attention based on Q2, a query derived from concatenating the previously decoded feature map 224C with an embedded representation of the target view, indicated here as Px2and a key matrix denoted as K2, a limited set of features selected from the encoder 210 in accordance with the depth map feature projection process 230, concatenated with embedding camera parameters for the corresponding source view. The second local attention layer 228-2 produces the final decoded feature map 224CC. Finally, the decoder 220 applies a final activation function, such as the Softmax function, to classify the features of the final decoded feature map 224CC into image pixels, indicated here as RGB output 225.

[0050] Thus, the target image 206 is decoded from the target view camera parameters 204, drawing feature information from the encoder 210 through a global attention layer 222 and two local attention layers 228-1 , 228-2, at multiple scales. However, it should be understood that the model 200 is simplified for illustrative purposes, and that in practice, additional layers (corresponding to higher or lower resolutions) may be used. For example, the decoder 220 may include two or more global attention layers (at or near the top of the decoder 220), the decoder 220 may include several more local attention layers (toward the bottom of the decoder 220), and the encoder 210 may include several more convolutional layers.

[0051] In general, each attention layer at the decoder 220 corresponds to the layer of feature maps at the encoder 210 of matching feature resolution. However, in some cases, an attention layer may be configured to attend to a feature map higher or lower in the encoder 210 that may not match in resolution. Further, in some cases, an attention layer may be configuredto attend to multiple layers of feature maps up and down the encoder 210, if computational resources permit.

[0052] As mentioned above with regard to FIG. 1 , the machine learning model 200 of FIG. 2 may be trained on a diverse range of scenes and may be made generalizable across scenes. The machine learning model 200 may be trained solely on image loss (e.g., L2 rendering loss), with the capability to decode the three-dimensional structure of a scene as an intermediate process, drawing solely from the two-dimensional features extracted from the source images 202.

[0053] Further, the machine learning model 200, including the trained learned neural network weights, biases, activation functions, and other architectural components and functionality, may be embodied in non-transitory machine-readable programming instructions, and executable by one or more processors of one or more computing devices, which include memory to store programming instructions that embody the functionality described herein and one or more processor to execute the programming instructions.

[0054] In terms of applications, the machine learning models described above may be applied in any use case for novel view synthesis, including, for example, novel view synthesis of objects, interior scenes, exterior scenes, even including large outdoor scenes comprising large structures such as buildings, roads, and landscapes. Indeed, the computational efficiency achieved through the use of the local attention mechanism described herein lends itself to modeling large scenes with sufficient computational resources and / or efficiently modeling smaller scenes where less computational resources are available.

[0055] Furthermore, as described above, the images synthesized using the techniques described above may be used as the source imagery for automated machine learning feature extraction. These synthetic images may be used for two-dimensional feature extraction such as for roads and building footprints, and also, leveraging the three-dimensional positional information provided by the depth maps that are generated as an intermediate product, in the use case of three-dimensional feature extraction, such as for extracting complex roof geometry and other detailed urban and natural landcover features.

[0056] The 769 Application proposes the use of an autoregressive machine learning model with an attention mechanism similar to the transformer model that can be trained to produce sequences of annotation operations that mimic the way a human annotator would perform annotation operations for feature extraction on single images. The model can produce annotation data in the form of sequences of coordinate tokens and operation tokens that together represent the series of drawing actions that can be performed to annotate target features in the referenced image. The output tokens can also encode for spatial constraints, such as parallelism among edges of the extracted polygon, and can also encode for particulardrawing actions, such as straight lines versus curved lines. These sequences of tokens can then be interpreted directly into geometric models in the form of vector maps that contain positional coordinate information that describes the extracted feature. Although the 769 Application is incorporated in its entirety herein by reference, certain elements of that disclosure are recast here, for greater understanding, throughout FIG. 4 to FIG. 5.

[0057] FIG. 4 is a schematic diagram of an example machine learning model 400 for generating a sequence of annotation tokens that represents a target feature to be extracted from a source image. The machine learning model 400 is an autoregressive model comprising an encoder 410 and a decoder 450 in a deep learning architecture.

[0058] The encoder 410 is to process a source image 412 to generate a feature map 414. A reference to the source image 412 may include the raw pixel data (e.g., RGB) as well as metadata such as camera parameters (e.g., focal length, lens distortion, camera pose, resolution), geospatial projection information (e.g., latitude and longitude position), or other data. The source image 412 may contain one or several batches of imagery covering the area, which may have been captured on the same dates or on different dates.

[0059] The feature map 414 encodes key features of the source image 412. For example, where the machine learning model 400 is to extract land cover and land use features from remote imagery, the feature map 414 will encode for geometric information about the various land features depicted in the imagery (e.g., shape, spatial constraint), and if applicable, the feature types associated with such features (e.g., building footprint, grassland, etc.).

[0060] The encoder 410 may include any suitable encoding layers, such as a self-attention layer (that applies attention among the elements of the input sequence), a convolutional neural network (CNN), a combination thereof, or other type of encoding layer capable of encoding key information about the features depicted in the source image 412. The encoder 410 may comprise a block of several of such encoding layers (Nx) stacked on top of one another.

[0061] The decoder 450 is to decode the feature map 414 into an output sequence of annotation tokens 452, which is a tokenized representation of the target features to be extracted from the source image 412. The decoder 450 is autoregressive in that it uses both the feature map 414 and any previously-generated elements of the output sequence of annotation tokens 452, depicted here as the autoregressive feed 454, to generate the output sequence of annotation tokens 452.

[0062] The decoder 450 may include any suitable decoding layers, such as a self-attention layer, a cross-attention layer (that applies attention between the elements of the input sequence and the output sequence), a deconvolution layer, or a combination thereof. The decoder 450 may comprise a block of several of such decoding layers (Nx) stacked on top of one another.

[0063] The machine learning model 400 may include additional components such as embedding layers, positional encoding, additional neural layers, output activation functions, and other components.

[0064] The output sequence of annotation tokens 452 may be a sequence of coordinate tokens and operation tokens which may be interpreted to represent the features extracted from the source image 412. Following further processing into a vector map by an interpretation module, the output sequence of annotation tokens 452 may be used in conventional GIS software, Computer-Aided Design (CAD) software, and the like.

[0065] In terms of training, the machine learning model 400 will generally need to be trained using labeled annotation data, collected from manual image annotators, who are trained to annotate single images with polygonal outlines of target features. The resulting two-dimensional annotation data can be used to train the machine learning model 400 to extract similar features.

[0066] The machine learning model 400 (and any of its subcomponents) may be embodied in non-transitory machine-readable programming instructions and executable by one or more processors of one or more computing devices, which include memory to store programming instructions that embody the functionality described herein and one or more processor to execute the programming instructions.

[0067] In the use case where the machine learning model 400 is to be applied to extract target features from a synthetic image that was generated through the techniques described in FIG. 1 through FIG. 3, the encoder 410 may be omitted, since one of the feature maps that were decoded as part of the generation of the synthetic image may be used as the feature map 414 that is input to the decoder 450. Generally, the highest-resolution decoded feature map may be used in this way. Alternatively, the encoder 410 may still be used, as described above, to encode the synthetic image into the feature map 414, for example, to obviate the need to store the previously decoded feature maps for the synthetic images, which may be beneficial for reduced memory storage or for other workflow purposes.

[0068] FIG. 5 is a schematic diagram of a system 500 for generating a geometric model that represents a target feature to be extracted from a source image, which includes a machine learning model similar to the machine learning model 400 of FIG. 4, shown in greater detail.

[0069] The machine learning model 502 is another autoregressive model comprising an encoder 510 and decoder 550 in a deep learning architecture. The encoder 510 is to process a source image 512 to generate a feature map 514. The feature map 514 encodes key features of the source image 512, including the geometry (i.e., shape, spatial constraint) of land cover and land use features visible in the source image 512, and their associated feature types.

[0070] The encoder 510 includes a convolutional neural network (CNN) 516 as the primary encoding layer. The CNN 516 is applied over the source image 512 to extract the target features. The encoder 510 may comprise a block of several of such CNN layers (Nx) stacked on top of one another.

[0071] The decoder 550 is to decode the feature map 514 into an output sequence of annotation tokens 552 of the target features to be extracted from the source image 512. The decoder 550 is autoregressive in that it uses both the feature map 514 and any previously- generated elements of the sequence of annotation tokens 552, depicted here as the autoregressive feed 562, to generate the sequence of annotation tokens 552.

[0072] The decoder 550 includes a self-attention layer 554 (to apply attention among the elements of the autoregressive feed 562), a cross-attention layer 556 (to apply attention between the elements of the autoregressive feed 562 and the feature map 514), and a feedforward layer 558 for further processing, similar to a transformer model. The decoder 550 may comprise a block of several of such decoding layers (Nx) stacked on top of one another. The output of the decoder 450 is fed into a softmax function 560 or other suitable activation function. Prior to input into the decoder 550, the autoregressive feed 562 is converted into an output embedding by output embedding layer 564.

[0073] The machine learning model 502 may include additional components such as skip connections, additional neural layers, and other components. In some examples, the various components of the machine learning model 502 may be rearranged where appropriate. The attentive layers may apply attention in accordance with any known techniques, including full / global attention, local attention, efficient attention using clustering, and other techniques. The CNN 516 may be applied in accordance with any known techniques, including the use of several convolutional layers of varying kernel size, and the like.

[0074] The sequence of annotation tokens 552 may comprise a sequence of coordinate tokens and operation tokens which may be interpreted to represent the target features extracted from the source image 512. The coordinate tokens may represent the coordinates of the various geometric entities extracted from the source image 512 (e.g., vertices of building rooftop features, points along the centerline of a road). The operation tokens represent the types of annotation operations that are used to assemble the vertices together into the resulting geometric entities. These may include simply “start” and “stop” tokens, or may also include more tokens that represent particular drawing actions that connect the vertices (e.g., straight line segments, curved line segments, two-dimensional primitives, etc.) and the spatial constraints among them (e.g., parallelism). As such, the sequence of annotation tokens 552 may accurately encode for detailed design elements and spatial constraints among geometric entities (e.g., the vertices of a building footprint polygon are to be joined by straight lines with certain lines perpendicular and / or parallel to one another, whether roof ridge lines are parallel orperpendicular to one another, whether the points along the centerline of a road are to be joined by curved lines, etc.). The use of operation tokens in this manner enables the machine learning model 502 to annotate or “draw” the target features to be extracted from the source image 512 as a more detailed and accurate reflection of the ground truth than by simply outputting a set of vertices.

[0075] In the example shown, each output token represents either an annotation operation, or a single dimensional coordinate of a point of a geometric entity (that is, each dimensional coordinate is output one at a time). Thus, the sequence of annotation tokens 552 as shown, in progress, begins with a “start building” token (to indicate that a building rooftop is being drawn), followed by a “draw line” token (to indicate the building rooftop will begin with the drawing of one or more straight lines), followed by the X-coordinate of the first point of the building rooftop, followed by the Y-coordinate of the first point of the building rooftop, and so on. The output sequence of annotation tokens 552 may proceed in this manner to produce an entire geometric model of the building rooftop, at least as viewed from the perspective of the source image 512, shown here as geometric model 572, comprising points A, B, and C (depicted in progress). Although not shown, once the geometric model 572 is completed, the output sequence of annotation tokens 552 may output a “close feature” token that indicates that the preceding vertices are to be grouped together as a single target feature representing the completed building rooftop (or at least the portion extracted from this source image 512). The sequence of annotation tokens 552 may also include tokens that represent individual features of a larger structure (e.g., draw rooftop outline, draw roof ridge line, valley, etc.).

[0076] It is noteworthy that, since the decoder 550 applies self-attention among the elements of the output sequence of annotation tokens 552, the machine learning model 502 may learn to apply different annotation techniques in different annotation scenarios. For example, the machine learning model 502 may be more likely to sample parallel and perpendicular lines following a “start ridge line” token (because roof ridge lines drawn with straight lines that are perpendicular and / or parallel to one another), than it would when drawing a “start valley line” token (because roof valley lines are rarely parallel or perpendicular to other roof features). Thus, the machine learning model 502 may learn to opt for different drawing techniques to represent different feature types, which may produce a more detailed and accurate reflection of the ground truth, than if it were limited to modeling the geometry as a set of vertices.

[0077] As mentioned previously, the sequence of annotation tokens 552 may be made interpretable by conventional GIS software, CAD software, and the like, after further processing by an interpretation module, shown here as interpreter 570. The interpreter 570 is configured with a set of rules that provides a complete set of instructions for how to interpret the various tokens produced by the machine learning model 502. In other words, the interpreter 570 isconfigured to convert, translate, decode, or otherwise interpret the output sequence of annotation tokens 552 as a set of points, lines, and / or polygons representative of the target features extracted from the source image 412, into a format that is suitable for CAD software, GIS, and the like. As discussed herein, this geometric model 572 that may then be projected onto the depth map corresponding to the source image 512 to determine the three-dimensional structure of the extracted features.

[0078] As with the machine learning model 400, the encoder 510 of the machine learning model 502 may be omitted for use with synthetic images for which an underlying feature map may already be readily accessible.

[0079] The functionality of the system 500 (and any of its subcomponents) may be embodied in programming instructions and executable by one or more processors of one or more computing devices, such as servers in a cloud computing environment, which include memory to store programming instructions that embody the functionality described herein and one or more processors to execute the programming instructions. It is emphasized that system 500 (with appropriate modifications if applicable) may be applied to extract vector data representing any sorts of features from any sort of imagery captured by any sort of image capture device.

[0080] One way to combine the methods for feature extraction described above in FIG. 4 to FIG. 5 with the methods for novel view synthesis described in FIG. 1 to FIG. 3 for the use case of two-dimensional feature extraction is to generate synthetic images in an orthographic projection and to use these orthographic images in the feature extraction process. Orthographic images have properties that are particularly useful in geospatial applications and lend themselves particularly well to geospatial feature extraction (i.e., uniform scale, geometric accuracy, reduced or eliminated relief displacement). Vector maps that are extracted from orthographic images inherit these advantages.

[0081] Another way to combine the above-described methods is the approach described above for the use case of three-dimensional feature extraction. Although the machine learning models for feature extraction described above operate on single images, and therefore produce annotation data which is, at least initially, generated within the two-dimensional confines of the synthetic images, this annotation data can be projected through three-dimensional space onto the three-dimensional depth maps that were generated as a bi-product of producing the synthetic images, thereby producing a three-dimensional model of the extracted feature. This approach is illustrated in FIG. 6, below.

[0082] In FIG. 6, a system 600 includes a machine learning model for novel view synthesis, denoted here as NeRF Model 610, and a system that involves a machine learning model for feature extraction, denoted here as Annotation Model 650. It is to be understood that the NeRFModel 610 may be similar to those machine learning models for novel view synthesis described in FIG. 1 through FIG. 3, and that the Annotation Model 650 may be similar to the system described in FIG. 5. For further description of the NeRF Model 610 and the Annotation Model 650, reference may be had to the descriptions of the figures referenced above.

[0083] The NeRF Model 610 accesses source images 602 and a target view 604, and generates a target image 606 along with a corresponding depth map 608. The Annotation Model 650 receives the target image 606 and generates annotation data 682 that represents a target feature in the target image 606. A projection module 690 then receives the depth map 608 and the annotation data 682, and projects the annotation data 682 onto the depth map 608, to produce 3D model data 692, which contains the three-dimensional positional information that results from projecting the two-dimensional annotation data 682 onto the three-dimensional depth map 608. Since the depth map 608 was generated as part of the process of generating the target image 606, the two-dimensional annotation data 682, which were also extracted from the target image 606, when projected onto the depth map 608, should fall directly onto the three-dimensional points occupied by the target feature. This result is illustrated also in FIG. 7, in the example use case where the target feature to be extracted is a three-dimensional pitched roof structure.

[0084] In FIG. 7A, a synthetic image 704A is generated in an orthographic projection based on captured overhead imagery, and this synthetic image 704A is annotated with two- dimensional annotations 702A (whether automatically, manually, or semi-manually). These two- dimensional annotations 702A are projected onto its corresponding depth map 706A, in accordance with the methods described above. As a result, each of the two-dimensional points in the two-dimensional annotations 702A fall onto a corresponding three-dimensional point in the depth map 706A, thereby providing a three-dimensional model 708A for the extracted feature. Thus, the two-dimensional annotation data is transformed into a three-dimensional wireframe mesh.

[0085] It should be apparent at this stage that the above-described projection process enables three-dimensional features to be extracted using a two-dimensional image annotation process with the knowledge that the third dimension will be accounted for in the corresponding depth map. Thus, a user, or a machine learning model, may be trained to annotate the roof outline, ridge lines, valleys, hips, and other components of a roof, each of which may be located at different heights from one another may traverse three-dimensional space in ways that may not be directly apparent from the synthetic image, by nevertheless annotating these features as two-dimensional annotations, since these two-dimensional annotations will be projected onto an underlying depth map that contains the relevant positional information in the third dimension. A machine learning model that is to perform feature extraction in this way should be trained to generate two-dimensional annotations of three-dimensional features accordingly. A typical setof training data could be based on three-dimensional wireframe models projected onto flat two- dimensional images, and the task for the machine learning model would be to reproduce two- dimensional annotations that, when projected onto a corresponding depth map, will result in the three-dimensional model.

[0086] It also should be noted at this stage that it may be advantageous to use synthetic images that are generated in an orthographic projection for the feature extraction process, since orthographic imagery eliminates or reduces problems associated with inconsistent scale, relief displacement, and occlusions. Thus, a three-dimensional model of a pitched roof structure may be annotated on an orthographic image that is synthesized from a point of view that is directly overhead of the roof structure, thereby providing visibility of all roof facets with correct geometric accuracy.

[0087] It should also be emphasized at this stage that the above-described methods are not limited to extracting three-dimensional features that are to be extracted from an overhead point of view. Although the above-described methods may be particularly useful in such use cases in the geospatial industry, it is emphasized that the above-described methods could equally be applied to extract three-dimensional features from ground-level imagery, or any other type of imagery. For example, as shown in FIG. 7B, street-view imagery may be used to perform feature extraction of building wall details. In FIG. 7B, a synthetic image 704A is generated in an orthographic projection based on captured street-view imagery of the side of a building, and the synthetic image 704B is annotated with two-dimensional annotations 702B that outline the side walls, door frames, and windows visible on the side of the building. These two-dimensional annotations 702B are projected onto a corresponding depth map 706B to produce three- dimensional model 708B. As another example, images of an interior scene captured from a smartphone camera may be collected and used to synthesize images and extract three- dimensional models of objects located in the room.

[0088] FIG. 8 is a flowchart of an example method 800 for generating a three-dimensional geometric model that represents a three-dimensional target feature to be extracted from source imagery, in accordance with the above-described techniques. The method 800 may be employed with any of the systems and / or devices described herein, including those of FIG. 6, FIG. 9, and FIG. 10, or other systems and / or devices. At operation 802, a plurality of source images that depict a scene from multiple points of view are accessed. At operation 804, at least some of the plurality of source images are each encoded into one or more corresponding feature maps. At operation 806, a synthetic image is decoded based on a target view of the scene, and based on the feature maps into which the at least some of the plurality of source images were encoded, wherein decoding the synthetic image involves generating at least one depth map for the target view as an intermediate product. At operation 808, a machine learning model is applied to the synthetic image to generate a set of annotations that outline a targetfeature visible in the at least some of the plurality of source images. At operation 810, the set of annotations is projected onto the at least one depth map. At operation 812, a three-dimensional structure of the target feature is reconstructed based on projecting the set of annotations onto the depth map. The steps of the method 800 may be organized into one or more functional processes and embodied on a non-transitory machine-readable storage medium in programming instructions executable by one or more processors in any suitable configuration, including the computing devices of the systems described herein.

[0089] FIG. 9 is a schematic diagram of an example system 900 for extracting a three- dimensional structure of a target feature, in accordance with the techniques described above, with additional hardware components described for illustrative purposes. The system 900 includes a machine learning model for novel view synthesis, denoted here as NeRF Model 950, and a system for r feature extraction, denoted here as Annotation Model 960. It is to be understood that the NeRF Model 950 may be similar to the NeRF Model 610 of FIG. 6, and that the Annotation Model 960 may be similar to the Annotation Model 650 of FIG. 6. For further description of the NeRF Model 950 and Annotation Model 960, reference may be had to the descriptions of the figures referenced above.

[0090] The system 900 includes one or more image capture devices 910 to capture image data 912 covering a scene 902 (i.e., an area of interest) that contains a target feature 904. For example, an image capture device 910 may include any suitable camera system capable of capturing geospatial imagery (e.g., aircraft, satellite) or other overhead imagery (e.g., drone, balloon). As another example, an image capture device 910 may include any suitable camera system capable of capturing ground-level imagery (e.g., street-view vehicle). An image capture device 910 may also include any suitable mobile device similarly capable of capturing images (e.g., smartphone).

[0091] The types of target features 904 may include any natural landcover features, such as forests, grass, bare land, shrubs, trees, water, and the like, or any manmade land use features such as buildings, roofs, roads, bridges, railways, driveways, crosswalks, sidewalks, parking lots, pavement, and the like. A common use case for the system 900, which is illustrated here, is to extract the three-dimensional structure of a pitched roof structure, such as the rooftop of a common residential home.

[0092] The image data 912 may include the raw image data (e.g., 3-band or 4-band imagery) in any suitable format that is made available by the image capture devices 910. The image data 912 may further include metadata associated with such imagery, including camera parameters (e.g., focal length, lens distortion, camera pose), geospatial projection information (e.g., latitude and longitude position), and other data. For three-dimensional feature extraction, the image data 912 should contain such raw image data and metadata for a collection of multiview imagery of the target feature 904 from multiple different perspectives or points of view.

[0093] The system 900 further includes one or more computing devices 920 to process the image data 912 as described herein. In particular, the computing devices 920 are configured to process the image data 912 through the NeRF Model 950 and Annotation Model 960 as described herein to generate three-dimensional model data 926 that represents the three- dimensional structure of the target feature 904.

[0094] The computing devices 920 may include one or more computing devices, containing computer processors (e.g., CPUs and / or GPUs) such as servers in a local or cloud-computing environment. The computing devices 920 may include one or more communication interfaces to receive / obtain / access the image data 912 and to output / transmit the resulting three-dimensional model data 926 through one or more computing networks and / or telecommunications networks such as the internet. The computing devices 920 may include memory to store the resulting three-dimensional model data 926 and to store executable programming instructions that embody the functionality described herein.

[0095] The computing devices 920 may also run a combination of ancillary software programs, machine learning training tools, data quality control tools, and other software used to train and configure the machine learning models and other software modules to perform the functionality described herein that is involved in processing the image data 912 into three- dimensional model data 926.

[0096] The three-dimensional model data 926 may contain vector data comprising sets of points, lines, polygons, and / or wireframe meshes, and may also contain the associated geometric constraints among those geometric elements, that represent the structure (i.e., geometry) of the target feature 904 to be extracted from the image data 912. These three- dimensional model data 926 can be stored and converted into any suitable format (e.g., .shp, .cad or other file type) to be imported into any suitable software application such as a computer- aided design (CAD) system or geographic information system (GIS) for viewing and / or further manipulation. The three-dimensional model data 926 may be attributed with additional information such as scale or geospatial projection information. For example, the three- dimensional model data 926 that correspond to a building that was extracted may be attributed with location information (e.g., GPS coordinates), scale information, address data, or other pertinent information that may be available either from the image data 912 (i.e., information contained in, or derived from, the camera parameters) or other data sources.

[0097] After generation, the three-dimensional model data 926 may be transmitted to one or more end user devices 930, which may be used to store, view, manipulate, and / or otherwise use such three-dimensional model data 926 as three-dimensional models 934 (either directly as incorporate into a particular filetype such as .CAD or .OBJ). For this purpose, the end user devices 930 may store, host, access, run, or execute one or more software programs that process such three-dimensional model data 926 (e.g., a GIS viewer), indicated here as an enduser application 932. The end user devices 930 may communicate with the computing devices 920 to access the three-dimensional model data 926 through any suitable means, such as through an application programming interface (API), access through a website, or similar.

[0098] Users of the end user devices 930 may use the three-dimensional models 934 for any such purposes as for city planning, land use planning, architectural and engineering work, property insurance risk assessments, environmental assessments, automated vehicle navigation, or for use in virtual reality or augmented reality systems, for the generation of a digital twin of a city, and the like. As one particular example, the end user application 932 may be configured to process a data file comprising the three-dimensional model data 926 to generate a property report, which may contain a three-dimensional building rendering and building measurements generated based on the three-dimensional model data 926. In the case where the target feature 904 is a pitched roof structure of a common residential building, for example, the building measurements could include an estimate of the square footage of the building footprint of the building, a square footage of the roof structure of the building, a height of the building (at the base of the pitched roof structure), measurements of particular roof features such as total ridge length, roof perimeter, or other measurements. Such details about the structure and measurements of the roof of the building may be useful in use cases such as insurance claims adjustment or underwriting activities.

[0099] FIG. 10 is a schematic diagram of another example system 1000 for extracting a three-dimensional structure of a target feature, in accordance with the techniques described above, which may be similar to the system 900 of FIG. 9, with the exception that the system 1000 includes one or more human image annotators 1062 who provide user input 1064 to an annotation model 1060 in combination with, or as an alternative to, the automated feature extraction performed by the annotation model 960 of FIG. 9. For a discussion of the elements of the system 1000, including the scene 1002, target feature 1004, image capture devices 1010, image data 1012, computing devices 1020, NeRF Model 1050, Annotation Model 1060, 3D model data 1026, end user devices 1030, end user application 1032, and 3D model 1034, reference may be had to the description of FIG. 9.

[0100] In addition, the system 1000 provides one or more tools that could be used by image annotators 1062, either to further refine the annotation data produced by an automated machine learning model as part of the annotation model 1060, or to generate annotation data directly. In either case, the annotation data and / or refinements are referred to as user input 1064. The provided annotation tools could include those annotation tools disclosed in the 769 Application, or in U.S. Patent Application No. 18 / 437,477, filed February 9th, 2024, entitled PLATFORM FOR DISTRIBUTED LANDCOVER FEATURE DATA REVIEW, the entirety of which is incorporated herein by reference.

[0101] Thus, it should be seen that two-dimensional and three-dimensional target features can be extracted directly from synthetically generated images in accordance with the techniques described above. This approach comes with several advantages over conventional approaches to three-dimensional feature extraction.

[0102] First, this approach does not require a separate data source that provides three- dimensional information, as is common in the geospatial industry. A typical approach for extracting three-dimensional buildings, for example, may involve extracting two-dimensional rooftop polygons based on the provided imagery and then attributing the rooftop polygons with height information obtained from LiDAR. In such an approach, it can be challenging to ensure alignment of the different data sources, especially when the vintages of the data sources do not necessarily match over the entire area of interest. The approach proposed herein obviates the need to align and reconcile different data sources. Rather, the semantic feature information and the three-dimensional positional information can both be extracted from the same data source - the source images themselves.

[0103] Second, as compared to constructing three-dimensional models with reference to multiview imagery directly, our process enables three-dimensional feature extraction directly from single images. Although the initial collection and encoding of source images requires multiview imagery, the following step of feature extraction can be performed on single images. Furthermore, the initial process of encoding the multiview source imagery and generating the required synthetic images (and corresponding depth maps) can be trained entirely end-to-end on image loss, requiring no annotated training sets - and therefore is highly scalable. Moreover, the following step of feature extraction, since it is performed on single images rather than on multiview images, is also much more highly scalable. It is much easier to train human annotators to annotate two-dimensional features than it is to generate three-dimensional models, and likewise it is much easier to gather two-dimensional annotation data than it is to gather three-dimensional models for training data.

[0104] Third, as compared to monocular depth estimation, estimating depth using multiview imagery and novel view synthesis is likely to produce more accurate depth maps. Although methods for monocular depth estimation have been proposed, such methods do not benefit from the additional perspectives provided by multiview imagery. Thus, it is expected that the techniques proposed herein will produce more accurate three-dimensional positional information for extracted features than approaches that involve monocular depth estimation.

[0105] Therefore, the techniques described herein provide for a highly-automated machine learning pipeline through which highly-accurate two-dimensional and three-dimensional semantic feature information can be generated based on source images alone. It should be understood that the features and aspects of the various examples provided above can be combined into further examples that also fall within the scope of the present disclosure. Thescope of the claims should not be limited by the above examples but should be given the broadest interpretation consistent with the description as a whole.

Claims

CLAIMS1 . A method comprising: accessing a plurality of source images that depict a target three-dimensional feature from multiple points of view; encoding each of at least some of the plurality of source images into one or more corresponding feature maps; decoding a synthetic image based on a target view of the target three-dimensional feature, and further based on the corresponding feature maps into which the source images were encoded, wherein decoding the synthetic image involves generating one or more depth maps corresponding to the target view as an intermediate product; applying a machine learning model to the synthetic image to generate a set of two- dimensional annotations that represent the target three-dimensional feature as depicted in the synthetic image; projecting the set of two-dimensional annotations onto at least one of the depth maps; and reconstructing a three-dimensional model of the target three-dimensional feature based on the projected two-dimensional annotations.

2. The method of claim 1 , wherein: the plurality of source images comprises at least one of: an aerial image, a satellite image, and a street-view image.

3. The method of claim 1 , wherein the target three-dimensional feature comprises a pitched roof structure of a building.

4. The method of claim 1 , wherein: decoding the synthetic image comprises: decoding one or more feature maps corresponding to the synthetic image, including at least one feature map from which pixel information of the synthetic image can be determined; and applying the machine learning model comprises: applying cross-attention across one or more of the feature maps corresponding to the synthetic image.

5. The method of claim 1 , wherein applying the machine learning model comprises:encoding the synthetic image into a feature map corresponding to the synthetic image; and applying cross-attention across the feature map corresponding to the synthetic image.

6. A method comprising: generating a synthetic image of a scene from a target view based on a plurality of source images, wherein generating the synthetic image involves generating a depth map for the target view to assist with generating the synthetic image; accessing a set of annotations that outline a target feature that appears in the synthetic image; projecting the set of annotations onto the depth map; and reconstructing a three-dimensional model of the target feature based on projecting the set of annotations onto the depth map.

7. The method of claim 6, wherein: generating the synthetic image involves: encoding the plurality of source images each into a corresponding series of multiscale feature maps; applying global attention, based on the target view, across a set of higher-level features of the series of multiscale feature maps of at least some of the source images that were encoded, to produce a first set of decoded features; applying a convolutional and upsampling layer to the first set of decoded features to produce a second set of decoded features; generating the depth map based on the second set of decoded features; for each point on the depth map corresponding to a pixel that should be rendered in the synthetic image, back-projecting the point through the target view to determine a set of lower-level features of the series of multiscale feature maps to be used to decode the pixel; and applying local attention across the set of lower-level features to produce a third set of decoded features from which pixel information of the synthetic image can be determined.

8. The method of claim 6, wherein generating the synthetic image of the scene from the target view comprises generating the synthetic image based on an embedded representation of a set of camera parameters corresponding to the target view.

9. The method of claim 6, wherein the set of annotations that outline the target feature were generated manually through an image annotation system.

10. The method of claim 6, wherein the set of annotations that outline the target feature were generated automatically at least in part by a machine learning model.11 . The method of claim 6, wherein the target feature comprises at least a component of a three-dimensional structure of a building.

12. A method comprising: accessing a synthetically generated image of a scene and a corresponding depth map; applying a machine learning model to the synthetically generated image to generate a set of annotations that represent a target feature visible in the synthetically generated image; projecting the set of annotations onto the corresponding depth map; and reconstructing a three-dimensional model of the target feature based on projecting the set of annotations onto the depth map.

13. The method of claim 12, wherein: applying the machine learning model involves: accessing a feature map corresponding to the synthetically generated image; and decoding a sequence of tokens that represents the target feature, autoregressively, token-by-token, wherein decoding each token comprises: applying cross-attention across the feature map; applying self-attention across any previously decoded tokens in the sequence of tokens; and generating the token based on at least the cross-attention and the selfattention; and interpreting the sequence of tokens as a vector map representation of the target feature; and wherein the sequence of tokens comprises a combination of coordinate tokens and operation tokens, wherein each coordinate token represents one or more coordinates of a vertex of the vector map, and wherein at least one operation token represents a drawing action to be performed to connect at least two vertices of the vector map.

14. The method of claim 13, wherein at least one operation token represents a selection operation by which a constraint on the coordinates of particular vertex of the vector map is defined with reference to a previously generated group of vertices of the vector map.

15. The method of claim 13, wherein the target feature comprises a pitched roof structure.

16. A system comprising one or more computing devices configured to: access a plurality of source images that depict a target three-dimensional feature from multiple points of view; encode each of at least some of the plurality of source images into one or more corresponding feature maps; decode a synthetic image based on a target view of the target three-dimensional feature, and further based on the corresponding feature maps into which the source images were encoded, wherein decoding the synthetic image involves generating one or more depth maps corresponding to the target view as an intermediate product; apply a machine learning model to the synthetic image to generate a set of two- dimensional annotations that represent the target three-dimensional feature as depicted in the synthetic image; project the set of two-dimensional annotations onto at least one of the depth maps; and reconstruct a three-dimensional model of the target three-dimensional feature based on the projected two-dimensional annotations.

17. A system comprising one or more computing devices configured to: generate a synthetic image of a scene from a target view based on a plurality of source images, wherein generating the synthetic image involves generating a depth map for the target view to assist with generating the synthetic image; access a set of annotations that outline a target feature that appears in the synthetic image; project the set of annotations onto the depth map; and reconstruct a three-dimensional model of the target feature based on projecting the set of annotations onto the depth map.

18. A system comprising one or more computing devices configured to: access a synthetically generated image of a scene and a corresponding depth map;apply a machine learning model to the synthetically generated image to generate a set of annotations that represent a target feature visible in the synthetically generated image; project the set of annotations onto the corresponding depth map; and reconstruct a three-dimensional model of the target feature based on projecting the set of annotations onto the depth map.

19. At least one non-transitory machine-readable storage medium comprising instructions that when executed cause one or more processors to: access a plurality of source images that depict a target three-dimensional feature from multiple points of view; encode each of at least some of the plurality of source images into one or more corresponding feature maps; decode a synthetic image based on a target view of the target three-dimensional feature, and further based on the corresponding feature maps into which the source images were encoded, wherein decoding the synthetic image involves generating one or more depth maps corresponding to the target view as an intermediate product; apply a machine learning model to the synthetic image to generate a set of two- dimensional annotations that represent the target three-dimensional feature as depicted in the synthetic image; project the set of two-dimensional annotations onto at least one of the depth maps; and reconstruct a three-dimensional model of the target three-dimensional feature based on the projected two-dimensional annotations.

20. At least one non-transitory machine-readable storage medium comprising instructions that when executed cause one or more processors to: generate a synthetic image of a scene from a target view based on a plurality of source images, wherein generating the synthetic image involves generating a depth map for the target view to assist with generating the synthetic image; access a set of annotations that outline a target feature that appears in the synthetic image; project the set of annotations onto the depth map; and reconstruct a three-dimensional model of the target feature based on projecting the set of annotations onto the depth map.

21. At least one non-transitory machine-readable storage medium comprising instructions that when executed cause one or more processors to: access a synthetically generated image of a scene and a corresponding depth map; apply a machine learning model to the synthetically generated image to generate a set of annotations that represent a target feature visible in the synthetically generated image; project the set of annotations onto the corresponding depth map; and reconstruct a three-dimensional model of the target feature based on projecting the set of annotations onto the depth map.

Citation Information

Patent Citations

  • Pose prediction method and device, electronic equipment and computer readable medium

    CN115311641A

  • Three-dimensional target detection method, system and equipment based on point cloud and image and medium

    CN118097123A

  • Data synthesis using three-dimensional modeling

    US11272164B1