Apparatus and methods for 3D semantic geometry generation
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-11-14
- Publication Date
- 2026-05-21
Smart Images

Figure EP2024082360_21052026_PF_FP_ABST
Abstract
Description
[0001] APPARATUS AND METHODS FOR 3D SEMANTIC GEOMETRY GENERATION
[0002] FIELD OF THE INVENTION
[0003] The present disclosure relates to the field of generative machine learning models, in particular to a method of generating semantically labelled 3D geometry.
[0004] BACKGROUND
[0005] It is advantageous in many different fields of technology to be able to generate 3D geometry, or 3D scenes, having semantic labelling. A 'semantic label' refers to a high-level feature in computer vision tasks that maps pixel-level features, e.g., individual pixels, with specific classes. Semantic labelling can be performed on an image through classification algorithms such as segmentation models. In this disclosure, reference to ‘3D semantic data’ or 3D semantic geometry’ refers to some representation of a 3D scene that contains semantic labels on particular regions or points of the 3D geometry. For example, the 3D geometry pay be represented as a point cloud, and a semantic label may be applied to each point such that each point is categorised as representing a particular object.
[0006] One field in which 3D semantic data 3D semantic data is especially helpful, if not essential, is self-driving technology. 3D semantic data is crucial for enhancing the reliability and robustness of various downstream tasks, such as scene understanding, perception, 3D completion, planning, and decision-making. However, producing or obtaining large-scale geometrically and semantically coherent data presents significant challenges. On the one hand, labelling data manually is a time-consuming task, even for 2D image data. Labelling a large 3D scene, e.g., an outdoor scene containing many different objects, would be a prohibitive task even with the assistance of an automated model.
[0007] On the other hand, models may be used to generatively obtain 3D semantic data. Producing semantic 3D geometry currently necessitates the use of highly complex models that can grasp the intricate relationships between the semantics and geometry of real-world scenes. However, this task is not straightforward because training these complex models requires vast amounts of annotated 3D data. Annotated 3D data is often less readily available compared to 2D data, such as images. Moreover, the problem is then one of manually producing annotated 3D data for training the models, which is itself a highly time-consuming task.
[0008] As 2D generative diffusion models gain popularity, recent approaches have sought to adapt their architectures for generating 3D semantic scenes. Diffusion models are a class of generative models that learn the distribution of data, and, hence, sampling generated data, by iteratively denoising random noise into structured outputs. Their working principle is similar to a diffusion process where they successively corrupt structured data into noise and train a neural network to reverse this process. Diffusion models have been successfully applied in various domains like text and audio, but are still typically limited to simpler 2D representation.
[0009] The scarcity of large, annotated 3D datasets compared to their 2D counterparts hinders the performance of current diffusion-based methods. In one respect, the scarcity of annotated 3D datasets would make 3D diffusion models less effective and less generalizable across varied data distributions, since the training data is highly limited. This makes them wholly unsuitable for producing semantic 3D data for self-driving technology, because the 3D data self-driving technology must itself be diverse in order to be representative of the vast differences that exist between different real-world urban, sub-urban, and rural environments. Furthermore, such 3D diffusion models are by their nature complex, and so take a prohibitive amount of time and computing resources to train. For example, it would not be viable to train or deploy such a 3D diffusion model on consumergrade hardware or GPUs. In the field of 3D diffusion, text-to-3D methods are a class of frameworks that train 3D models (i.e., representations) by distilling the knowledge of 2D diffusion models into a 3D representation. They achieve this through the idea of Score Distillation Sampling (SDS). The main idea is to combine text-to-image and 3D representation by optimizing a 3D model via aligning the distribution of rendered images with the target distribution that is derived from the 2D diffusion model. Such methods can generate models of small 3D objects, but suffer from the disadvantage that they fail to generalize to the more complex large-scale outdoor scenes and fail completely to produce scenes having a plurality of objects. Thus, such text-to-3D methods are not suitable as generative methods for producing semantic 3D geometry for the training of self-driving technology.
[0010] Furthermore, the generated semantic grids from these methods are often restricted to lower, coarse resolutions, which may limit their applicability for downstream tasks, especially for self-driving technology in which the detail must be representative of the level of detail at which sensors on real vehicles can operate.
[0011] It is, therefore, an object of the present disclosure to solve the aforementioned problems, amongst other things, that are inherent in the current state-of-the-art models for generating 3D semantic data. The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.
[0012] SUMMARY
[0013] In view of the scenarios above, improved apparatus and methods for generating 3D semantic geometry are needed that are generalisable and so enable varied outputs. Improved apparatus and methods for generating 3D semantic geometry are also needed in ways which are not prohibitive to train or deploy. This can be done by providing, inter alia, off-the-shelf (i.e., pretrained) generative models and semantic segmentation models hat iteratively refine a partial representation of a 3D geometry, without the need to rely on 3D annotated data. This is advantageous because, amongst other things, it avoids the cost of training and avoids the need to rely on a small set of pre-labelled data (which would result in underfitting to that small set and so result in a model that is not generalisable or capable of producing a varied output).
[0014] A first aspect of the present disclosure provides an apparatus for generating semantically labelled 3D geometry from a partial 3D representation, wherein the partial 3D representation represents a scene comprising a plurality of objects, the apparatus configured to:
[0015] obtain the partial 3D representation of the scene, wherein at least a portion of the partial 3D representation comprises undefined geometry and wherein at least a portion of the geometry in the partial 3D representation comprises semantic labelling;
[0016] iteratively refine the partial 3D representation to obtain the semantically labelled 3D geometry, the iterative refinement comprising, in dependence on each pose of a plurality of different poses selected from the partial 3D representation:
[0017] generating a 2D semantic representation at the pose;
[0018] in dependence on the 2D semantic representation, using a pre-trained generative neural network to generate a detailed 2D image, wherein the detailed image represents the perspective at the pose and comprises colour information for each pixel;
[0019] in dependence on the detailed 2D image, using a semantic segmentation model to generate detailed semantic labelling, the detailed semantic labelling comprising semantic labels for objects represented in the detailed 2D image;
[0020] in dependence on the detailed semantic labelling generated by the semantic segmentation model, update i) the partial semantic labelling of the partial 3D representation and ii) a geometry of the partial 3D representation, to thereby obtain an improved partial 3D representation;
[0021] use the improved partial 3D representation to generate a subsequent pose of the plurality of different poses used for the iterative refinement. In example implementations, the detailed 2D image comprises a 2D representation of geometry that was undefined in the partial 3D representation of the scene, and wherein the apparatus is configured to update the geometry of the partial 3D representation by determining a geometry for at least some of the geometry that was undefined in the partial 3D representation.
[0022] In example implementations, semantic segmentation model is configured to categorise at least a portion of the pixels of the detailed 2D image.
[0023] In example implementations, the detailed 2D image comprises a plurality of objects, and updating the partial semantic labelling of the partial 3D representation comprises labelling at least some objects in the partial 3D representation in dependence on objects labelled by the semantic segmentation model.
[0024] In example implementations, the apparatus is configured to generating the 2D semantic representation by using a neural radiance field model.
[0025] In example implementations, the neural radiance field model is configured to update the geometry of the partial 3D representation by inferring geometry from the detailed semantic labelling.
[0026] In example implementations, the pre-trained generative neural network is a 2D diffusion model configured to generate an RGB image in dependence on the 2D semantic representation.
[0027] In example implementations, the apparatus may be further configured to: use the semantically labelled 3D geometry to generate a plurality of semantic images and / or a semantic video for use in simulating a driving scene for training a self-driving vehicle.
[0028] In example implementations, the scene is an outdoor scene, and wherein the apparatus is configured to select each pose of the plurality of the different poses to be a drivable area within the outdoor scene.
[0029] In example implementations, the apparatus is configured to iteratively refine the partial 3D representation without using predetermined 3-dimensional semantic labelling.
[0030] In example implementations, the partial 3D representation is a 2D birds-eye-view of the scene, wherein the birds-eye-view comprises height data for a plurality of objects within the scene.
[0031] In example implementations, the apparatus is configured to obtain the partial 3D representation of the scene by:
[0032] obtaining a 2D birds-eye-view image comprising partial semantic labelling for the plurality of objects; obtaining a height map defining heights of at least some of the plurality of objects within the scene;
[0033] in dependence on the 2D birds-eye-view image and the height map, using at least one generative model to generate the partial 3D representation.
[0034] In example implementations, the apparatus is configured to generate the partial 3D representation by:
[0035] in dependence on the 2D birds-eye-view image and the height map, using a neural radiance field model to infer at least some 3D geometry and thereby generate a coarse 3D representation;
[0036] in dependence on the coarse 3D representation, using a pre-trained 2D diffusion model to generate a detailed 2D birds-eye-view image; in dependence on the detailed 2D birds-eye-view image, using a semantic segmentation model to categorise at least a portion of the pixels of the detailed 2D image to thereby generate the at most 2-dimensional semantic labelling.
[0037] A second aspect of the present disclosure provides a method of generating semantically labelled 3D geometry from a partial 3D representation, wherein the partial 3D representation represents a scene comprising a plurality of objects, the method comprising:
[0038] obtaining the partial 3D representation of the scene, wherein at least a portion of the partial 3D representation comprises undefined geometry and wherein at least a portion of the geometry in the partial 3D representation comprises semantic labelling;
[0039] iteratively refining the partial 3D representation to obtain the semantically labelled 3D geometry, the iterative refinement comprising, in dependence on each pose of a plurality of different poses selected from the partial 3D representation:
[0040] generating a 2D semantic representation at the pose;
[0041] in dependence on the 2D semantic representation, using a pre-trained generative neural network to generate a detailed 2D image, wherein the detailed image represents the perspective at the pose and comprises colour information for each pixel;
[0042] in dependence on the detailed 2D image, using a semantic segmentation model to generate detailed semantic labelling for geometry represented by the scene;
[0043] in dependence on the detailed semantic labelling generated by the semantic segmentation model, updating i) the partial semantic labelling of the partial 3D representation and ii) a geometry of the partial 3D representation to thereby obtain an improved partial 3D representation;
[0044] using the improved partial 3D representation to generate a subsequent pose of the plurality of different poses used for the iterative refinement.
[0045] A third aspect of the present disclosure provides an apparatus for generating semantically labelled 3D geometry from a partial 3D representation, wherein the partial 3D representation represents a scene comprising a plurality of objects, the apparatus configured to:
[0046] obtain a 2D image representing a scene, the 2D image comprising partial semantic labelling for a plurality of objects depicted in the image;
[0047] obtain height data defining a third dimension for the plurality of objects within the scene;
[0048] use a neural radiance field model to infer, in dependence on the 2D birds-eye-view image and the height data, at least some 3D geometry within the scene and thereby generate a coarse 3D representation;
[0049] use a first 2D diffusion model to generate, in dependence on the coarse 3D representation, a colour 2D birds-eye-view;
[0050] use a semantic segmentation model to categorise, in dependence on the colour 2D birds-eye-view image, at least some pixels of the detailed 2D image to thereby generate a partial 3D representation of the scene, wherein at least a portion of the partial 3D representation comprises undefined geometry and wherein the partial 3D representation comprises semantic labelling;
[0051] iteratively refine the partial 3D representation to obtain the semantically labelled 3D geometry, the iterative refinement comprising, in dependence on each pose of a plurality of different poses selected from the partial 3D representation:
[0052] generating noisy 2D semantic representation of the pose;
[0053] in dependence on the 2D semantic representation, using a pre-trained generative neural network to generate a colour 2D image comprising newly generated objects;
[0054] in dependence on the colour 2D image, using a semantic segmentation model to generate a detailed 2D semantic representation of the pose, the detailed 2D semantic representation comprising semantic labels for the newly generated objects; in dependence on the detailed 2D semantic representation, update i) partial semantic labelling of the partial 3D representation ii) a geometry of the partial 3D representation to thereby obtain an improved partial 3D representation ; and use the improved partial 3D representation to generate a subsequent pose of the plurality of different poses for use in a subsequent step of the iterative refinement.
[0055] A fourth aspect of the present disclosure provides a computer program stored in non-transitory form and including code instructions which, when executed on one or more processor, causes the one or more processor to execute any of the methods disclosed herein.
[0056] BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The present disclosure is described by way of example, with reference to the accompanying drawings, in which:
[0058] Figure 1 shows an overview of a method of generating 3D semantic data starting from a simple 2D layout, according to presently disclosed embodiments;
[0059] Figure 2 shows a first phase involved in generating 3D semantic data in which a 2D layout is refined using a neural radiance field to provide a birds-eye-view initialisation having partial 3D information;
[0060] Figure 3 shows a method of generating 3D semantic data using iterative improvement of a partial representation of a 3D geometry using a neural radiance field model, according to presently disclosed embodiments;
[0061] Figure 4 illustrates a practical example of refinements for a scene according to Figure 2;
[0062] Figure 5 illustrates a practical example of refinements for a scene according to Figure 3; and
[0063] Figure 6 is a flowchart illustrating a method of generating 3D semantic data according to figure 3, and presently disclosed embodiments.
[0064] DETAILED DESCRIPTION
[0065] As mentioned above, in order to generate consistent large-scale 3D semantic data or grids, prior art methods typically rely on repurposing classical 2D diffusion modes to adapt them to the 3D case. For example, pyramid diffusion introduces a Pyramid of Discrete Diffusion (PDD) model that leverages 3D U-Nets (a convolution neural network used for image segmentation) to generate large semantic scenes in a coarse-to-fine manner. This approach enhances the detail and resolution of the generated 3D scenes by progressively refining the scene representation. In another example, ‘Urban Architect’ introduces a framework for generating 3D urban scenes that allows for user-control through 3D layout specifications (i.e., 3D shape primitives). It provides this in part by reformulating the Variable Score Distillation (VSD) that is used in typical text-to-3D models.
[0066] However, all known methods suffer from several disadvantages that make them unsuitable or prohibitively costly (in terms of computing or training time) to use. The most fundamental limitation is that they all require pre-determined 3D labelled data, e.g., annotated 3D data. This is much less abundant compared to 2D data. Indeed, collecting and labelling 3D data is inherently more complex and resource-intensive in the first place and so new 3D annotated data cannot readily be obtained. By comparison, 2D images can be easily obtained from various sources, including cameras and online databases.
[0067] Moreover, 3D data acquisition involves specialised processes such as LiDAR (Light detection and ranging) scanning or photogrammetry, which are less accessible and more costly compared to cameras. The result is a limited diversity of the trained model, which in turn limits the data generated by the model to the distribution of the dataset that the model was trained on. As is well known in machine learning fields, models trained on a small set of data, or on data that is not diverse, do not generalise well and risk being ‘overfit’. This makes such models unsuitable for tasks such as generating scenes for training self-driving vehicles, where it is critical to provide a broad and diverse range of scenes and environments in order to explore and expose the limits of self-driving algorithms. Furthermore, known methods necessitate the training of diffusion models from scratch, which is computationally expensive. Due to their iterative noising-denoising nature, the training process can take days or even weeks. Furthermore, latent diffusion models require the additional training burden of the autoencoder which, depending on the data size, might also take several days. Training may not even be possible on consumer-grade hardware and GPUs due to the large memory and bandwidth that such models need in order to store the vast amounts of 3D geometry data.
[0068] Known methods also produce generations with limited resolution, which may be inadequate for certain downstream tasks. In fact, generating fme-grained, detailed scenes is preferable to low-resolution ones, as it is generally more beneficial to downsample a highly detailed scene than to upscale a coarse one. Consequently, downstream tasks that require low-resolution scenes can achieve better results with downsampled, high-resolution scenes compared to those generated at low resolutions.
[0069] Consequently, the inventors have established that such problems can be solved by iteratively refining the partial / coarse 3D representation of a scene containing some semantic labels but undefined geometry, to thereby obtain the semantically labelled 3D geometry. This can be performed by sampling a plurality of poses from the partial / coarse 3D representation and, for each pose of a plurality of different poses selected from the partial 3D representation: generate a 2D semantic representation, use a pre-trained generative neural network to generate a detailed 2D image from the 2D semantic representation, and then use a semantic segmentation model to generate detailed semantic labelling from the detailed 2D image. Once the more detailed 2D semantic representation is obtained (which, on the basis of the pre-trained generative neural network, contains semantic labels for new details or objects) the partial semantic labelling and the geometry (e.g., the undefined geometry) of the partial 3D representation can be updated. Thus improved / updated partial 3D representation is then used to generate a subsequent pose of the plurality of different poses used for the iterative refinement, and the iterative refinement continuous until a detailed 3D model with detailed semantic labelling is produced.
[0070] In one example, producing highly detailed 3D semantic scenes as outlined above involves using off-the-shelf (i.e., pre-trained) generative 2D diffusion models. The method involves an iterative refinement phase where 2D-generated images (preferably colour images) enhance the initialized scene. The 3D representation is implicitly learned within a semantic NeRF model. Optionally, the method may also involve a first initialisation phase where a coarse geometry is derived from a bird’s-eye- view (BEV) 2D layout to produce the partial / coarse 3D representation having undefined geometry.
[0071] Examples in the present disclosure pertain to generating a complex outdoor scene that is suitable for generating training data for use with a self-driving vehicle. Thus, the examples in this disclosure pertain to producing a scene based on a ‘bird’s eye view’ (BEV) initialisation. This is because a BEV is a particularly helpful starting point to obtain a coarse representation containing roads, streets, trees, and an overall layout of an urban environment. For example, a BEV initialisation representation would be more appropriate than a street-level view initialisation representation, in which other roads, buildings, and objects would be occluded. For this reason, the initialisation geometry is referred to as the BEV initialisation. It should therefore be appreciated that this is specific only to the example of producing 3D semantic data for a typical urban driving scene. The skilled person would appreciate that embodiments of the present disclosure are not restricted to initialisation scenes that are birds-eye views. An advantage of the present disclosure is that complex 3D scenes comprising a plurality of objects, e.g., any outdoor scene, can be generatively obtained alongside detailed semantic labels. Thus, the skilled person would appreciate that the methods described herein are applicable to a wide variety of viewpoints, scenes, and environments other than BEV-type scenes.
[0072] Figure 1 provides a broad overview of the overall method 100, including the optional initialisation phase. The first (optional) phase is termed the ‘BEV’ initialisation phase 102, and the second phase is termed the street- level refinement phase. The first phase 102 is optional because the BEV initialisation geometry may be pre-determined or pre-obtained for use with phase 2. Detailed examples are provided below. The objective of the optional BEV initialisation phase 102 is to start with a simple / coarse 2D representation of a BEV 106 (also called a ‘2D BEV prior’ or ‘2D layout prior’) and transform it into a partial 3D representation 110 with semantic labels. This partial 3D representation 110 may also be referred to as a ‘2.5D’ semantic model. The 2.5D semantic model 110 is a coarse model in that it contains some undefined geometry, and the geometry it does contain lacks details. For example, the 2.D BEV may contain only a few objects, and the only 3-dimensional aspect of the data is the height data associated with those objects. As indicated in the figure, the 2.5D model is represented using a neural radiance field (NeRF) model. The NeRF model is also used to iteratively transform the simple 2D BEV prior 106 into the 2.5 scene (the details of which are described below with respect of Figure 2).
[0073] The second phase is referred to as a street-level refinement phase 114, in which more detailed objects, complete geometry, and more semantic labels are generated. In this phase, the 2.5D model 110 is refined into a fine-grained semantic NeRF model 112, which is 3D. This is a fine-grained semantic NeRF model 112 is the end result, i.e., the ‘semantically labelled 3D geometry’ mentioned above. The fine-grained semantic NeRF model 112 is obtained by the iterative method outlined above, i.e., including pose sampling 116 from the NeRF model and generating street-level renderings 114 using a pre-trained generative model.
[0074] In general, the method takes as input a coarse 2D BEV semantic layout 106 of the scene to be generated. This layout can be obtained in many different ways, e.g., manually sketched, obtained from a simulator, or generated using a generative model, e.g., a text-to-image diffusion model used to generate BED scenes. In the case of a text-to-image BEV initialization, the input to the method could simply be some text describing the scene and its components.
[0075] Figure 2 shows a detailed schematic illustrating the steps of the first ‘BEV initialisation’ phase. The objective of the first phase is to establish the initial form of the 3D geometry of the street refinement phase. The initialisation phase 102, therefore, does not need to generate a complete set of geometry, but rather ‘sets the scene’ such that large-scale features are present, alongside some height or depth dimensions for at least some objects within the scene.
[0076] As mentioned above, the initial input for the first initialisation phase 102 is a ‘BEV coarse semantic prior’ 204, which in essence is a simple 2D BEV. This ‘BEV coarse semantic prior’ 204 is equivalent to the ‘2D layout prior’ 106 in Figure 1. This can be a sketch or a simple rendering of a BEV or plan view of a scene. For example, for an outdoor urban environment, the BEV coarse semantic prior may illustrate several streets, trees, pavements, and cars in basic outline, absent details. The BEV coarse semantic prior is a semantic image, and so regions / pixels of the prior are categorised as representing a particular object.
[0077] The BEV scene is initialised in more detail using a neural radiance field (NeRF) model, and the prior 204 contains the semantic supervision for the NeRF. A NeRF is a radiance field that implicitly represents a 3D scene in a continuous manner. A typical NeRF is a multidimensional model where, in each 3D location / point of a 3D scene, a colour and a density value are learned by a neural network. Different views may be synthesised by querying 5D coordinates that define a projected ray (represented by a 3 -dimensional position coordinate [x, y, z], and a 2-dimensional viewing angle [0, <[>]) and using a volume rendering technique to project the output colours and densities into an image. By accumulating the local information along rays defined by a camera using volume rendering, the final colour and geometry of the scene can be generated. Radiance fields are a powerful tool for Novel View Synthesis (NVS) as their model within the 3D representation enables some light effects such as secularity. However, radiance fields remain challenging to train on large unbounded areas, especially where multiple different objects exist within the scene, and are costly to use for rendering.
[0078] Turning back to the phase 1 process, the BEV prior 204 is used to initialise a scene using a semantic NeRF 108. This is a NeRF that is trained to include semantic information from the views it synthesises. In other words, the NeRFs used in the present disclosure have been extended to incorporate semantic understanding (beyond the standard colour field), which is utilized to render novel views of the scene. The semantic NeRF 108 uses ray marching to generate two new images based on the viewpoint of the BEV Coarse semantic prior: a rendered noisy semantic 206 image and a rendered noisy depth image 208. The rendered noisy semantic 206 does not comprise RGB information: it contains only labels, for at least some pixels in the image (preferably all pixels) that categorise those pixels as being particular objects. It will nevertheless be appreciated that, to convey the semantic labels, each semantic category within the 2D image may be rendered in a different colour. The rendered noisy depth image 208 comprises depth / height information for at least some objects, as inferred by the NeRF. At the initial stage, the NeRF may not have any height or depth information: this is introduced at step 214.
[0079] The rendered noisy semantic 206 is input into a 2D diffusion model 202. At step 210, the 2D diffusion model generates a new 2D RGB image based on the rendered noisy semantic 206. The 2D diffusion model thus uses the semantic data contained in the rendered noisy semantic 206 to condition / supervise the generation of the image. For example, regions labelled as ‘tree’ in the noisy semantic 206 will be rendered in full RGB by the 2D diffusion model. In short, the objective of the 2D diffusion model is to generate more nuanced detail, and even new objects, into the BEV scene. Once the detailed / RGB BEV image 210 is generated, it is fed into a 2D semantic segmentation model 212 to revert the image back into a semantic representation. Thus, the 2D semantic segmentation model 212 generates a new detailed semantic image from the generated RGB BEV image 210. This is because the semantic NeRF is interested only in the semantic data, and not the RGB information (which does not form part of the NeRF model), and so only semantic information is used to train the NeRF.
[0080] In more detail, the scene intended to be generated in the present disclosure is a complex scene from a single BEV semantic image. Inferring the complex geometry from such an initial would be exceedingly challenging using a typical NeRF (which stores colour information, i.e., 3-dimensional RGB information), i.e., it would take a very long time to train. Moreover, RGB information provides little to no additional value for use in training self-driving vehicles, since the semantic information is the critical factor. Thus, the present NeRF model discards the higher dimensions of the RGB information and uses semantics instead.
[0081] However, relying on semantics only may cause a NeRF to overfit to a single BEV image without accurately representing the actual scene geometry. The inventors therefore have established that, in order to address this issue, a height map prior should also be introduced for training the NeRF. This height map prior is introduced at step 214. In more detail, a height map prior assigns a predefined height for various structures (e.g., buildings, trees, etc) in the scene based on their class. The height map prior and the semantic information, derived from the RGB generated image 210 using the 2D semantic segmentation model 212, are thus used as input to train the NeRF model.
[0082] In order to implement this height map prior, a depth loss is applied, which is often commonly applied in depth- supervised training when only depth information is available. When a NeRF is supervised using depth obtained from an image, the depth that is obtained is not in the same scale as the scene (e.g., an object that 1 m from the camera will likely not have 1 m depth). Thus, depth supervision techniques are used that eliminate the need for this scale information. In one example, a scale- invariant depth loss is used, such as the one in “MegaDepth: Learning Single-View Depth Prediction from Internet Photos” by Li et al.
[0083] 2018.
[0084] The training step of the NeRF uses a loss function 216 which thus takes as input: the rendered noisy semantic 206 image, the rendered noisy depth image 208, the height-map prior, and the new detailed semantic image (generated by the 2D semantic segmentation model 212). Since what is being predicted is probabilities (i.e., of semantic categories) and not actual values, we cannot use MSE (mean square error) loss or the like. Thus, a cross-entropy loss function is used, which is a known loss function used in semantic segmentation tasks. The loss function thus calculates a loss between the semantic information generated by the 2D semantic model 212 and the semantic data in the rendered noisy semantic image 206, which is used to train the neural network underpinning the NeRF. In one example, the type of loss function used to train the semantic NeRF can be found in "In-Place Scene Labelling and Understanding with Implicit Scene Representation" by Zhi et al. 2021. Likewise, another aspect of the loss function calculates an objective loss between the height prior (shown in 214) and the rendered noisy depth 208 generated by the semantic NeRF model.
[0085] In summary of the phase 1 scene initialisation: using the coarse BEV prior 204 and per-class depth priors (shown in 214), a ‘2.5D’ semantic NeRF representation 110 of the scene is trained. Although it contains 3D information, this model is not fully 3D trained, as some geometry remains undefined, and fine details remain unobservable during the supervision phase. For example, the BEV may indicate the tops of buildings or trees, but other structures (such as tree trunks, or smaller objects occluded by the buildings) may not be defined. The result of this step is a coarsely trained NeRF that will be utilized to render 2D semantic control images for the refinement process shown in Figure 3.
[0086] Optionally, the model for generating the detailed RGB image using the pre-trained 2D diffusion model is additionally guided by ControlNet. ControlNet is a strategy by which the capabilities of a generative diffusion model are enhanced by allowing for more precise control (i.e., conditioning) over the generation process. It achieves this by leveraging zero convolutions. ControlNet operates by creating trainable copies of the original diffusion model, allowing these copies to learn from the conditioning data while maintaining the original model's generative power.
[0087] Figure 3 shows the street-refinement phase used to transform the coarse, and partially undefined, ‘2.5D’ semantic NeRF representation into a complete 3D semantic geometry representing a complete scene, of which the original BEV prior 204 represented a single birds-eye-view viewpoint. The goal of this refinement phase, in general, is to complex 3D semantic scenes that outperform known methods in terms of resolution, details and diversity, and which do not require any pre-determined 3D annotated data.
[0088] Step 306 represents the starting point, i.e., the ‘2.5D’ semantic NeRF representation that may, e.g., be produced from the iterations described in respect of Figure 2. The ‘2.5D’ semantic NeRF representation, also called a partial 3D representation of the scene, acts as an initialisation semantic map for generating new 3D semantic labels and geometry. The partial 3D representation may also be referred to as an ‘implicit’ representation. The implicit representation might be generated from by method described in respect of Figure 2 using a NeRF trained on a coarse BEV representation. Alternatively, the implicit representation may be generated using any of: a trained text-2-image BEV RGB or a semantic model, or a real- world mapbased rasterizer. The implicit representation 306 is, preferably, a drivable area of an outdoor scene, for use in generating diverse 3D semantic data for self-driving vehicle training tasks.
[0089] Step 308 represents a step of sampling a pose from the partial 3D representation. This pose sampling generated a plurality of 6-diminensional inputs (comprised of 3D position coordinate (x, y, z), and a 3D orientation from that position). The pose sampling is what allows the NeRF to generate images, and subsequently iteratively refine the NeRF. At each given sampled pose, a semantic image 312 is rendered from the implicit BEV NeRF model 110 (e.g. , as produced by the initialisation in Figure 2). For a given pose, HxW rays are generated, i.e., commensurate with the desired size of the resultant HxW image.
[0090] The semantic NeRF model 110 is not trained to output RGB information. Typically, a NeRF model outputs 4-dimensional data, i.e., RGB pixel information and depth information. This would be prohibitively difficult to train in the current scenario which is intended, preferably, to generate complex 3D scenes (e.g., outdoor scenes) comprising a large number of objects. Moreover, in present embodiments, the focus is on generating detailed and accurate semantic data, and detailed RGB information is not necessary. The present NeRF model therefore uses output, effectively 2-dimensional data, i.e., semantic information and depth information.
[0091] Step 312 represents a noisy semantic image output by the NeRF at a particular sampled pose. As mentioned above, this semantic image 312 represents one of the poses sampled in the pose sampling stage 308. The rendered noisy semantic is called a 2D semantic representation elsewhere in the specification. The noisy semantic image does not comprise RGB information: it contains only labels, for at least some pixels in the image (preferably all pixels) that categorise those pixels as being particular objects. It will nevertheless be appreciated that, in order to convey the semantic labels, the 2D image may be rendered in colour (e.g., block colour) using a plurality of colours, where each colour of the plurality of colours represents a particular category / label. This colourised representation of semantic labelling is convenient for a user to understand and may be used as an input for the diffusion model 304.
[0092] Step 314 represents a noisy depth output by the NeRF at the particular sampled pose. For each training iteration, the rendered semantic 312 and the rendered depth 314 represent the same pose. This 2D depth image again contains no RGB information, though as above the depth / height data may be represented using block colours. The noisy depth output therefore contains depth data for at least some of the objects that are contained within the representation of the sampled pose.
[0093] The 2D noisy semantic image 304 is then input into a 2D diffusion model. This diffusion model is substantially as described above with respect to the diffusion model 202 in Figure 2. As before, preferably, the diffusion model (which is preferably pretrained) is also guided by a ControlNet.
[0094] Step 316 represents the 2D noisy semantic image being input into the 2D diffusion model. The 2D diffusion model is preferably pre-trained (e.g., it may be trained on outdoor images containing various objects from a street-level viewpoint). The 2D diffusion model uses the semantic labels of the 2D noisy semantic image, generated by the NeRF model 110 to generate an RGB image 316. This RGB image 316 is equivalent to the 2D Generated street-level renderings 114 shown in Figure 1.
[0095] The RGB image generated by the 2D diffusion model does not have semantic labels, at this stage. Since the 2D diffusion model is a generative model, it is capable of generating new objects and details that are not represented in the 2D noisy semantic image, and which are not represented in the NeRF’s partial 3D representation. The 2D diffusion model may also ‘de-noise’ the 2D noisy semantic image. At this stage, the RGB image generated by the diffusion model is merely a 2-dimensional depiction of a scene that has been generated based on the guidance of the semantic labels in the rendered noisy semantic 312. Thus, although the RGB image 316 may contain new objects not (yet) represented in the NeRF’s model, the 2D image cannot define geometry. The geometry of objects, or newly generated details of existing objects, can only be revealed / inferred by the NeRF during the training step 320.
[0096] Since the newly generated RGB image contains new details and / or objects, but no labels, a 2D segmentation 318 model is used to revert the detailed RGB image back to a purely semantic representation. A mono-depth estimation function or model is also applied to the 2D RGB image 316 in order to provide depth data from training the NeRF (i.e., by comparing it to the rendered noisy depth data 314 using a suitable loss function). Step S318 involves a 2D semantic segmentation model generating new semantic labels for the RGB image created in step S316 by the 2D diffusion model. In other words, the 2D semantic segmentation generates a new, and more detailed, 2D semantic image, i.e., where preferably all pixels are categorized. Any new objects or details generated by the diffusion model 316 should also be categorised. The output of the 2D semantic model, therefore, may discard the RGB information, since this is not needed to train the NeRF. It should be appreciated that, although reference to RGB is made in respect of images produced by the 2D diffusion model, generally any colourised representation would be appreciated, as would be within the remit of the skilled person to derive. For example, the generated images could readily be represented in the YUV colour format.
[0097] At step 320, the NeRF model is updated based on the detailed 2D semantic image, and also the rendered noisy depth data output at 314. As with Figure 2, the loss function uses a combination of a cross-entropy loss and a scale-invariant depth loss. The result of the loss calculation is used to update the neural network underpinning the NeRF model, which involves updating the geometry of the NeRF model (e.g., inferring missing / undefined geometry) and updating the semantic information stored in the NeRF model 110.
[0098] To illustrate the operation of the street-level refinement, the following worked example of the above-described steps is provided, in which a NeRF model learns to depict the geometry of tree trunks from an initial coarse BEV in which only the tree canopies are indicated.
[0099] As step 306, the original implicit representation 306, which represents a BEV, is input, which contains several cars and the canopies of trees but not the tree trunks (since the tree trunks are occluded by the canopies).
[0100] At step 312, the 2D noisy semantic image generated by the NeRF model at step 312 will therefore contain a noisy label of a tree, in which the pixels representing the tree canopy may be labelled, but the area where the trunk should be may not be labelled (because the geometry of the trunk does not exist yet in the NeRF model).
[0101] At step 316, based on the semantic labelling, the 2D diffusion model will generate a colourised image that contains a depiction of a complete tree. In order to do this, preferably, the 2D diffusion model will have already been trained won a set of images containing trees, where the viewpoint of the images is at street level so that the trunks are visible. Advantageously, the 2D diffusion model is able to render the detail of the tree trunk in RGB, in addition to rendering the canopy.
[0102] At step 318, when this detailed colourised image is input into the 2D segmentation model, the output of the 2D segmentation model is a more detailed semantic image, devoid of colour data (e.g., RGB information), but which contains semantic information for a tree trunk.
[0103] When the more detailed semantic image is input back into the NeRF model in order to train the NeRF, the NeRF updates its semantic labelling and infers new geometry based on the new semantic labels in the detailed semantic image. Specifically, whereas the NeRF model did not initially contain the geometry for the tree trunk, the NeRF can infer a coarse geometry for a tree trunk based on the new semantic labels created by the semantic segmentation model. The NeRF model is thus able to infer new 3D geometry firstly because the 2D diffusion model is able to depict new objects, such as a tree trunk, based on guidance from i) the images on which it had been trained and ii) the noisy semantic data, and secondly because the 2D semantic segmentation model is able to label new objects generated by the 2D diffusion model, e.g., the tree trunk. Advantageously, therefore, the NeRF model can infer new 3D geometry without 3D semantic labels.
[0104] Figure 4 shows a worked example of refining a coarse BEV initial 204 according to the steps outlined in Figure 2. As shown in Figure 4, the starting coarse BEV initial 204 is represented in this case by a BEV semantic prior 400 which is a plan view of a street. The image is a purely semantic representation, i.e., where the different shades represent different object categories. In this example, the different categories depicted are roads, pavement / sidewalk, buildings, grassy regions, and trees. The per-class depth prior 402 is shown below. This depth prior represents exactly the same BEV viewpoint as the semantic prior. However, instead of semantic data, the depth prior indicates height data for certain objects. In this case, the height data indicated is applied to trees and buildings. As mentioned above, the height prior information 402 is used to train the initial NeRF representation (shown in steps 214 and 216 of Figure 2).
[0105] The result of the BEV initialisation 102 is a BEV-only trained NeRF model 110, also called the ‘2.5D semantic representation’ of the ‘partial 3D representation’ or the ‘implicit representation’ . The implicit representation 110 shown in Figure 4 is depicted from an oblique / perspective viewpoint, which serves to show buildings and trees from the side. However, since this NeRF model is trained only on a BEV (and not on multiple views from different angles, as is typical in known NeRF models), any geometry that is not visible from the BEV is undefined. For example, the 2.D implicit representation 110 does not define tree trunks or the main structure of the building. In this way, the 2.5D representation contains only partial 3D information (in this case height data) for particular regions of the scene. The role of the street-refinement phase described in Figure 3 is to learn, infer, and / or reveal this undefined geometry.
[0106] Figure 5 illustrates a worked example of refining the 2.5D representation implicit representation in the NeRF model to a finegrained and fully labelled 3D semantic representation 112. The starting point is the implicit / partial 2.5D representation 110. Figure 5 depicts a first iteration of street-level refinement 104 which outputs, from a particular pose which in this case is at street-level, a noisy semantic image 312 and a noisy depth image 314. Together, these represent the information 500 known to the NeRF model, at that pose, at the first iteration. As can be seen, the semantic image 312 does not contain any detail such as cars or lampposts, or well-defined road edges, and does not define trees or buildings well.
[0107] Figure 5 then depicts the noisy semantic image 312 being input into the 2D diffusion model 304, which outputs a detailed RGB image. Importantly, since the 2D diffusion model is a generative model, it is capable of introducing new objects and details in dependence on the corpus of images on which the diffusion model has been trained. The output of the diffusion model also depends on the semantic labels from the noisy semantic image 312. It should be appreciated that RGB image 316 illustrated in figure 5 is not exactly representative of the noisy semantic image 312 shown in Figure 5. The RGB image 316 in Figure is included merely to illustrate the concept that new objects, such as cars and trees, can be introduced by the diffusion model.
[0108] The newly-generated RGB image 316 is then input into a 2D semantic segmentation model 318, whose output (a semantic image) is input into the loss calculation together with the other input described above with reference to Figure 3. The loss is used to train the semantic NeRF model 110.
[0109] After one or more iterations, the representation of the NeRF model contains much richer geometric information and more detailed semantic information. Figure 5 thus shows an output 502 of a refined, fine-grained, NeRF model having been trained over a plurality of iterations. The output 502 shows a new semantic image 504 taken from the same pose as the first, noisy, semantic image 312. The output 502 also shows a new depth image 506 representative of the same pose as the first, noisy, depth image 314. As can be seen, the NeRF model now contains more detailed geometry: the pavement / sidewalk is now well-defined relative to the road, lampposts are now represented by the model, and buildings and trees have a better-defined shape.
[0110] Figure 6 shows a flowchart illustrating the method for generating 3D semantic geometry, an example of which is the streetlevel refinement described above with respect to Figure 3. The outlined method is for generating semantically labelled 3D geometry from a partial 3D representation, wherein the partial 3D representation represents a scene comprising a plurality of objects.
[0111] Step S600 comprises obtaining a partial 3D representation of the scene, wherein at least a portion of the partial 3D representation comprises undefined geometry and wherein at least a portion of the geometry in the partial 3D representation comprises semantic labelling. Steps S602 to S610 comprise iteratively refining the partial 3D representation to obtain the semantically labelled 3D geometry. The iterative refinement comprises the following steps, which depend on each pose of a plurality of different poses selected from the partial 3D representation.
[0112] Step S602 comprises generating a 2D semantic representation at the pose. This pose may simply be obtained in the first iteration, i.e., since the pose may already have been sampled e.g., at the pose sampling stage 308 shown in Figure 3. The 2D semantic representation may correspond to the rendered noisy semantic 312 in Figure 3 and Figure 5.
[0113] Step S604 comprises, in dependence on the 2D semantic representation, using a pre-trained generative neural network to generate a detailed 2D image, wherein the detailed image represents the perspective at the pose and comprises colour information for each pixel. The detailed 2D image may correspond to the semantic-conditions RGB generation 316 depicted in Figure 3, or Figure 5.
[0114] Step S606 comprises, in dependence on the detailed 2D image, using a semantic segmentation model to generate detailed semantic labelling, the detailed semantic labelling comprising semantic labels for objects represented in the detailed 2D image. The detailed semantic labelling may correspond to the detailed 2D semantic representation 318 produced by the 2D semantic segmentation model shown in Figure 3.
[0115] Step S608 comprises, in dependence on the detailed semantic labelling generated by the semantic segmentation model, updating i) the partial semantic labelling of the partial 3D representation and ii) a geometry of the partial 3D representation, to thereby obtain an improved partial 3D representation. This update step may correspond to the step of calculating the cross-entropy loss and the scale-invariant depth loss 320 shown in Figure 3, and subsequently using that calculated loss to update the geometry and semantics of the NeRF model 110.
[0116] Consequently, step S610 comprises using the improved partial 3D representation to generate a subsequent pose of the plurality of different poses used for the iterative refinement. After a plurality of iterations, the result should be a fine-tuned NeRF model 112 comprising a fully labelled 3D semantic representation of a scene.
[0117] In this disclosure, when the subject of a phase is described as being "configured to" or “arranged to”, followed by a term defining a condition or function, this is used to indicate that the subject of the phrase is in a state in which it has that condition, or is able to perform that function, without the subject being modified or further configured.
[0118] Some implementations may be described using the expressions “one / an embodiment” or “one / an implementation” or “one / an example”, along with their derivatives. These terms mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase “in some implementations” in various places in the specification are not necessarily all referring to the same embodiment. Moreover, unless otherwise noted the features described above are recognized to be usable together in any combination. Thus, any features discussed separately may be employed in combination with each other unless it is noted that the features are incompatible with each other.
[0119] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present disclosure may consist of any such individual feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the appended claims.
[0120] The foregoing description of example embodiments has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Many modifications and variations are possible in light of this disclosure. It is intended that the scope of the present disclosure be limited not by this detailed description, but rather by the claims appended hereto. Future filed applications claiming priority to this application may claim the disclosed subject matter in a different manner and may generally include any set of one or more limitations as variously disclosed or otherwise demonstrated herein.
Claims
CLAIMS1. An apparatus for generating semantically labelled 3D geometry from a partial 3D representation, wherein the partial 3D representation represents a scene comprising a plurality of objects, the apparatus configured to:obtain the partial 3D representation of the scene, wherein at least a portion of the partial 3D representation comprises undefined geometry and wherein at least a portion of the geometry in the partial 3D representation comprises semantic labelling;iteratively refine the partial 3D representation to obtain the semantically labelled 3D geometry, the iterative refinement comprising, in dependence on each pose of a plurality of different poses selected from the partial 3D representation:generating a 2D semantic representation at the pose;in dependence on the 2D semantic representation, using a pre-trained generative neural network to generate a detailed 2D image, wherein the detailed image represents the perspective at the pose and comprises colour information for each pixel;in dependence on the detailed 2D image, using a semantic segmentation model to generate detailed semantic labelling, the detailed semantic labelling comprising semantic labels for objects represented in the detailed 2D image;in dependence on the detailed semantic labelling generated by the semantic segmentation model, update i) the partial semantic labelling of the partial 3D representation and ii) a geometry of the partial 3D representation, to thereby obtain an improved partial 3D representation;use the improved partial 3D representation to generate a subsequent pose of the plurality of different poses used for the iterative refinement.
2. The apparatus of claim 1, wherein the detailed 2D image comprises a 2D representation of geometry that was undefined in the partial 3D representation of the scene, and wherein the apparatus is configured to update the geometry of the partial 3D representation by determining a geometry for at least some of the geometry that was undefined in the partial 3D representation.
3. The apparatus of claim 1 or 2, wherein semantic segmentation model is configured to categorise at least a portion of the pixels of the detailed 2D image.
4. The apparatus of claim 3, wherein the detailed 2D image comprises a plurality of objects, and updating the partial semantic labelling of the partial 3D representation comprises labelling at least some objects in the partial 3D representation in dependence on objects labelled by the semantic segmentation model.
5. The apparatus of any preceding claim, wherein the apparatus is configured to generating the 2D semantic representation by using a neural radiance field model.
6. The apparatus of claim 5, wherein the neural radiance field model is configured to update the geometry of the partial 3D representation by inferring geometry from the detailed semantic labelling.
7. The apparatus of any preceding claim, wherein the pre-trained generative neural network is a 2D diffusion model configured to generate an RGB image in dependence on the 2D semantic representation.
8. The apparatus of any preceding claim, wherein the scene is an outdoor scene, and wherein the apparatus is configured to select each pose of the plurality of the different poses to be a drivable area within the outdoor scene.
9. The apparatus of any preceding claim, wherein the apparatus is configured to iteratively refine the partial 3D representation without using predetermined 3-dimensional semantic labelling.
10. The apparatus of any preceding claim, wherein the partial 3D representation is a 2D birds-eye-view of the scene, wherein the birds-eye-view comprises height data for a plurality of objects within the scene.
11. The apparatus of claim 10 , wherein the apparatus is configured to obtain the partial 3D representation of the scene by:obtaining a 2D birds-eye-view image comprising partial semantic labelling for the plurality of objects; obtaining a height map defining heights of at least some of the plurality of objects within the scene;in dependence on the 2D birds-eye-view image and the height map, using at least one generative model to generate the partial 3D representation.
12. The apparatus of claim 11 , wherein the apparatus is configured to generate the partial 3D representation by:in dependence on the 2D birds-eye-view image and the height map, using a neural radiance field model to infer at least some 3D geometry and thereby generate a coarse 3D representation;in dependence on the coarse 3D representation, using a pre-trained 2D diffusion model to generate a detailed 2D birds-eye-view image;in dependence on the detailed 2D birds-eye-view image, using a semantic segmentation model to categorise at least a portion of the pixels of the detailed 2D image to thereby generate the at most 2-dimensional semantic labelling.
13. A method of generating semantically labelled 3D geometry from a partial 3D representation, wherein the partial 3D representation represents a scene comprising a plurality of objects, the method comprising:obtaining the partial 3D representation of the scene, wherein at least a portion of the partial 3D representation comprises undefined geometry and wherein at least a portion of the geometry in the partial 3D representation comprises semantic labelling;iteratively refining the partial 3D representation to obtain the semantically labelled 3D geometry, the iterative refinement comprising, in dependence on each pose of a plurality of different poses selected from the partial 3D representation:generating a 2D semantic representation at the pose;in dependence on the 2D semantic representation, using a pre-trained generative neural network to generate a detailed 2D image, wherein the detailed image represents the perspective at the pose and comprises colour information for each pixel;in dependence on the detailed 2D image, using a semantic segmentation model to generate detailed semantic labelling for geometry represented by the scene;in dependence on the detailed semantic labelling generated by the semantic segmentation model, updating i) the partial semantic labelling of the partial 3D representation and ii) a geometry of the partial 3D representation to thereby obtain an improved partial 3D representation;using the improved partial 3D representation to generate a subsequent pose of the plurality of different poses used for the iterative refinement.
14. An apparatus for generating semantically labelled 3D geometry from a partial 3D representation, wherein the partial 3D representation represents a scene comprising a plurality of objects, the apparatus configured to:obtain a 2D image representing a scene, the 2D image comprising partial semantic labelling for a plurality of objects depicted in the image;obtain height data defining a third dimension for the plurality of objects within the scene;use a neural radiance field model to infer, in dependence on the 2D birds-eye-view image and the height data, at least some 3D geometry within the scene and thereby generate a coarse 3D representation;use a first 2D diffusion model to generate, in dependence on the coarse 3D representation, a colour 2D birds-eye-view;use a semantic segmentation model to categorise, in dependence on the colour 2D birds-eye-view image, at least some pixels of the detailed 2D image to thereby generate a partial 3D representation of the scene, wherein at least a portion of the partial 3D representation comprises undefined geometry and wherein the partial 3D representation comprises semantic labelling;iteratively refine the partial 3D representation to obtain the semantically labelled 3D geometry, the iterative refinement comprising, in dependence on each pose of a plurality of different poses selected from the partial 3D representation:generating noisy 2D semantic representation of the pose;in dependence on the 2D semantic representation, using a pre-trained generative neural network to generate a colour 2D image comprising newly generated objects;in dependence on the colour 2D image, using a semantic segmentation model to generate a detailed 2D semantic representation of the pose, the detailed 2D semantic representation comprising semantic labels for the newly generated objects;in dependence on the detailed 2D semantic representation, update i) partial semantic labelling of the partial 3D representation ii) a geometry of the partial 3D representation to thereby obtain an improved partial 3D representation ; and use the improved partial 3D representation to generate a subsequent pose of the plurality of different poses for use in a subsequent step of the iterative refinement.
15. A computer program stored in non-transitory form and including code instructions which, when executed on one or more processor, causes the one or more processor to execute the method according to claim 13.17