Generating a 2d image of a 3D scene

Through the scene encoder and generated image model in machine learning methods, the problem of inaccurate generation of 2D images in the prior art is solved, and high-quality 2D images that consider 3D environment factors are generated, which improves the generation efficiency and image quality.

CN120337335APending Publication Date: 2025-07-18DASSAULT SYSTEMES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510065948.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing solutions for generating 2D images in 3D scenes fail to fully consider the object structure and relationships in the 3D environment, resulting in inaccurate images, natural and immersive content.

Method used

Using machine learning methods, 2D images are generated using the layout and viewpoint of the 3D scene through scene encoder and generative image model, including layout encoder, camera encoder, floor encoder and transformer encoder, combined with diffusion model for training and generation, taking into account perspective, occlusion and lighting factors.

Benefits of technology

More accurate, natural and immersive 2D images are generated, which can take into account the perspective of the 3D scene and the occlusion effect between objects, and does not rely on large pre-trained models, improving generation efficiency and image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337335A_ABST
    Figure CN120337335A_ABST
Patent Text Reader

Abstract

The present disclosure particularly relates to a computer-implemented machine learning method for training a function configured to generate a 2D image of a 3D scene (hereinafter referred to as a machine learning method). Functions include a scene encoder and a generative image model. A scene encoder takes a layout and viewpoints of a 3D scene as inputs and outputs a scene encoding tensor. The generative image model takes the scene encoding tensor output by the scene encoder as input and outputs the generated 2D image. A machine learning method includes obtaining a dataset including respective layouts and viewpoints of a 2D image and a 3D scene. The machine learning method includes training a function based on the obtained data set. This machine learning approach forms an improved solution for generating 2D images of a 3D scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly, to a machine learning method, system, and program for training a function configured to generate 2D images of 3D scenes. Background Art

[0002] Many systems and programs for object design, engineering, and manufacturing are available on the market. CAD is an abbreviation for Computer-Aided Design, which, for example, involves software solutions for designing objects. CAE is an abbreviation for Computer-Aided Engineering, which, for example, involves software solutions for simulating the physical behavior of future products. CAM is an abbreviation for Computer-Aided Manufacturing, which, for example, involves software solutions for defining manufacturing processes and operations. In these computer-aided design systems, the graphical user interface plays an important role in terms of technical efficiency. These technologies can be embedded in a Product Lifecycle Management (PLM) system. PLM refers to a business strategy that, in the context of the extended enterprise concept, helps companies share product data, apply common processes, and leverage corporate knowledge for product development from concept to the end of the product lifecycle. The PLM solutions provided by Dassault Systèmes (trademarks include CATIA, ENOVIA, 3DVIA, and DELMIA) provide an engineering hub for organizing product engineering knowledge, a manufacturing hub for managing manufacturing engineering knowledge, and an enterprise hub for enabling the integration and connection of the enterprise with the engineering and manufacturing hubs. Overall, the system provides an open object model that links products, processes, and resources to enable dynamic, knowledge-based product creation and decision support, thereby driving optimized product definition, manufacturing readiness, production, and services.

[0003] In this context, applications for 3D scene creation are being developed. These applications typically propose to create, manipulate, and provide 3D scenes, particularly (but not limited to) for touch-sensitive devices (such as smartphones or tablets). One task of these applications is to generate realistic 2D images of 3D scenes.

[0004] In recent years, solutions for generating 2D images for 3D scenes have been developed, for example, using generative deep learning models. However, these solutions do not fully take into account the entire 3D environment of the imaging scene. In particular, these solutions do not allow the use of knowledge of the 3D structure and relationships of objects in the scene or environment. Therefore, they are unable to produce accurate, natural, and immersive content, especially since they do not consider, for example, perspective, occlusion, or lighting factors.

[0005] In this context, there is still a need for an improved solution for generating 2D images of 3D scenes. Summary of the Invention

[0006] Accordingly, there is provided a computer-implemented machine learning method for training a function configured to generate 2D images of 3D scenes (hereinafter referred to as the machine learning method). The function includes a scene encoder and a generative image model. The scene encoder takes as input the layout and viewpoint of a 3D scene and outputs a scene encoding tensor. The generative image model takes as input the scene encoding tensor output by the scene encoder and outputs a generated 2D image. The machine learning method includes obtaining a data set including 2D images and corresponding layouts and viewpoints of 3D scenes. The machine learning method includes training the function based on the obtained data set.

[0007] The machine learning method may include one or more of the following:

[0008] - The layout of each 3D scene includes:

[0009] o a set of bounding boxes representing objects in the 3D scene; and

[0010] o the boundary of the 3D scene;

[0011] - The scene encoder includes a layout encoder configured to encode the set of bounding boxes. For each bounding box in the set, the layout encoder takes as input parameters representing the position, size, and orientation of the object represented by the bounding box in the 3D scene, and optionally parameters of the class of the object represented by the bounding box;

[0012] - The scene encoder further includes a floor encoder configured to encode the boundary of the 3D scene;

[0013] - The scene encoder further includes a camera encoder configured to encode the viewpoint;

[0014] - For each given 2D image of a given 3D scene in the data set, the size and position of the object represented by the bounding box in the layout of the given 3D scene are defined in a coordinate system based on the position and orientation of the camera that captured the given 2D image. Each viewpoint includes the field of view and pitch angle of the camera;

[0015] - The scene encoder further includes a transformer encoder. The transformer encoder takes as input the concatenation of the set of bounding boxes encoded by the layout encoder, the viewpoint encoded by the camera encoder, and the 3D scene boundary encoded by the floor encoder. The transformer encoder outputs a scene encoding tensor;

[0016] - The generative image model is a diffusion model;

[0017] - The diffusion model has an architecture including a denoiser, and the denoiser includes blocks. At least one of the blocks is enhanced by cross-attention using a scene encoding tensor; and / or

[0018] - The diffusion model is configured to operate in the latent space. The diffusion model is trained to denoise the compressed latent representation of 2D images of a dataset.

[0019] There is also provided a method of using a function after machine learning according to a machine learning method (hereinafter referred to as the usage method). The usage method includes obtaining the layout of a 3D scene. The usage method includes applying the function to the layout of the 3D scene, thereby generating a 2D image of the 3D scene.

[0020] The usage method may include one or more of the following:

[0021] - The generative image model is a diffusion model; and / or

[0022] - The application of the function includes:

[0023] o Applying a scene encoder to the obtained layout, thereby outputting a scene encoding tensor; and

[0024] o Using the diffusion model conditioned on the output scene encoding tensor to generate a 2D image of the 3D scene.

[0025] There is also provided a computer program including instructions for executing the machine learning method and / or the usage method.

[0026] There is also provided a computer-readable storage medium having a computer program recorded thereon.

[0027] There is also provided a system including a processor coupled to a memory, and a computer program is recorded on the memory. The system may further include a graphical user interface coupled to the processor.

[0028] There is also provided a device including a data storage medium having a computer program recorded thereon.

[0029] The device may form or be used as a non-transitory computer-readable medium, such as on Software as a Service (SaaS) or other servers, cloud-based platforms, etc. The device may alternatively include a processor coupled to the data storage medium. Thus, the device may form a computer system in whole or in part (e.g., the device is a subsystem of the entire system). The system may further include a graphical user interface coupled to the processor. Description of the Drawings

[0030] Non-limiting examples will now be described with reference to the drawings, where:

[0031] - Figure 1 Flowchart showing examples of machine learning methods and usage methods;

[0032] - Figure 2 Illustration of an example of a scene encoder;

[0033] - Figure 3 Illustration of an example of camera coordinates;

[0034] - Figures 4 to 12 Illustration of the layout and viewpoints of a 3D scene and the resulting 2D images generated by a trained function;

[0035] - Figure 13 Illustration of the results of a quantitative evaluation for assessing a trained function; and

[0036] - Figure 14 Illustration of an example of a system. Detailed Description

[0037] A computer-implemented machine learning method is proposed, configured to generate a function for 2D images of 3D scenes (hereinafter referred to as the machine learning method). This function includes a scene encoder and a generative image model. The scene encoder takes the layout and viewpoints of a 3D scene as input and outputs a scene encoding tensor. The generative image model takes the scene encoding tensor output by the scene encoder as input and outputs the generated 2D image. The machine learning method includes obtaining a data set including 2D images and the corresponding layout and viewpoints of 3D scenes. The machine learning method includes training the function based on the obtained data set.

[0038] This machine learning method provides an improved solution for generating 2D images of 3D scenes.

[0039] It is worth noting that the machine learning method allows training a function that can automatically and efficiently generate 2D images of 3D scenes. In particular, this function is trained to generate (various realistic) 2D images from a high-level, abstract, and surrogate representation of a 3D scene (and thus is easy to define). In fact, this training enables the function to generate 2D images of a 3D scene only from the layout and viewpoints of the 3D scene. From these two inputs, the trained function can generate 2D images, which is particularly useful and interesting for showing objects in a 3D scene. It is worth noting that compared with providing an exact object model for each object in a 3D scene and then using traditional rendering methods, it is much easier for the user to provide these two inputs to the trained function. Therefore, the trained function enables the user to easily and quickly generate 2D images of the 3D scene he or she is building, simply by defining its layout and providing viewpoints for these images.

[0040] In particular, the function is trained to generate particularly realistic and relevant 2D images of 3D scenes. In fact, the function includes a scene encoder that allows the layout in the generated 2D images to be taken into account. Thus, the 2D images generated by the function take into account the perspective, lighting, and occlusions between objects in the 3D scene (the scene encoder allows this information to be considered when generating the 2D images). In other words, the 2D images generated by the trained function are 3D-aware, i.e., it takes into account the 3D environment of the 3D scene. It is worth noting that this method allows off-screen objects in the 3D scene to be considered when generating 2D images, i.e., objects not within the camera's field of view to affect the generated image (e.g., light from a window).

[0041] In addition, the function is trained to generate a variety of 2D images. In fact, for a given 3D layout, the function is able to generate various 2D images, including object styles, colors, etc., while still respecting the layout. Thus, it allows users to obtain multiple inspirations from a single input abstract layout.

[0042] Furthermore, the proposed machine learning method is trained end-to-end for the task in a single training phase. It does not rely on large pre-trained image generation models, pre-trained depth estimators, or other external modules.

[0043] The machine learning method and / or the usage method are computer-implemented. This means that the steps (or substantially all steps) of the machine learning method and / or the usage method are performed by at least one computer or any similar system. Thus, the steps of the machine learning method and / or the usage method are performed by a computer, possibly fully automatically or semi-automatically. In an example, the triggering of at least some steps of the machine learning method and / or the usage method can be performed through user-computer interaction. The required level of user-computer interaction may depend on the expected level of automation and be balanced with the need to achieve the user's expectations. In an example, this level can be user-defined and / or pre-defined.

[0044] For example, before applying the function to the layout of a 3D scene, the usage method can include, for example, the step of determining the 3D scene layout during user interaction (e.g., performed by the user, such as when currently designing a 3D scene). The determination of the layout can include determining the boundaries of the 3D scene and a set of bounding boxes representing the objects in the 3D scene.

[0045] Now discuss the determination of the set of bounding boxes. Determining the set of bounding boxes can include, for each bounding box in the set, the steps of determining the dimensions of the bounding box and positioning the bounding box with the determined dimensions within the determined boundaries of the 3D scene. The steps of determining the dimensions and positioning can be manually performed by the user. For example, the step of determining the dimensions can include the user inputting the width, depth, and height (e.g., through user interaction using a keyboard). The positioning step can include the user inputting the coordinates of points on the bounding box (e.g., corners or its center) and its orientation, or can include the user moving the bounding box to its position in the 3D scene (e.g., the bounding box can be displayed on the screen and can be moved by the user using a mouse). For one or more bounding boxes (e.g., all bounding boxes), the step of determining the dimensions can be semi-automatically performed. For example, the step of determining the dimensions can include the user selecting the category of the object represented by the bounding box and automatically suggesting (e.g., from a database storing the default dimensions of different categories of objects) the dimensions (e.g., width, depth, and height) of the bounding box representing the object. The suggested dimensions can be accepted by the user, or can be further refined by the user (e.g., manually as described above). For example, if the user wants to add a sofa to their 3D layout, the method can suggest a default bounding box that matches the category of the object input by the user while allowing the user to modify these dimension values.

[0046] Now discuss the determination of the boundaries of the 3D scene. Determining the boundaries of the 3D scene can include determining the corresponding set of points representing the boundaries. For example, the determination of the boundaries can include determining some of the points in the set (e.g., representing the corners of the 3D scene), and then sampling other points on the boundaries between these points representing the corners of the 3D scene (i.e., along the walls of the 3D scene).

[0047] Before applying this function to the layout of the 3D scene, the use method can also include the step of determining the viewpoint. For example, the determination of the viewpoint can be performed by the user, e.g., by inputting the coordinates and / or orientation of the viewpoint, or by selecting this information on the screen displaying the 3D scene. Alternatively, the determination of the viewpoint can be automatically performed, e.g., by another function predicting one or more relevant viewpoints of the 3D scene (considering its layout).

[0048] A typical example of a machine learning method and / or a computer implementation of the use method is to use a system suitable for this purpose to execute the machine learning method and / or the use method. The system can include a processor coupled to a memory and a graphical user interface (GUI), with a computer program recorded on the memory, the program including instructions for executing the machine learning method and / or the use method. The memory can also store a database. The memory is any hardware suitable for storage and may include several physically distinct parts (e.g., one for the program, possibly one for the database).

[0049] The data set used for training the function considered by the machine learning method can be stored in a database. A "database" refers to any collection of data (i.e., information) organized for searching and retrieval (such as a relational database, for example, based on a predefined structured language like SQL). When stored in memory, the database allows a computer to search and retrieve quickly. The structure of the database does facilitate the storage, retrieval, modification, and deletion of data, as well as various data processing operations. A database may consist of a file or a collection of files, and these files (collection of files) can be broken down into records, each record consisting of one or more fields. A field is the basic unit of data storage. A user can retrieve data mainly through queries. Using keywords and sorting commands, a user can quickly search, rearrange, group, and select fields in many records according to the rules of the database management system used to retrieve or create a report of a specific data aggregation.

[0050] The method generally manipulates modeling (3D) objects. A modeling object is any object defined by data stored, for example, in a database. By extension, the expression "modeling object" specifies the data itself. Depending on the type of system, a modeling object can be defined by different types of data. The system can actually be any combination of CAD systems, CAE systems, CAM systems, PDM systems, and / or PLM systems. In these different systems, the modeling object is defined by the corresponding data. It can be correspondingly called a CAD object, PLM object, PDM object, CAE object, CAM object, CAD data, PLM data, PDM data, CAM data, CAE data. However, these systems are not mutually exclusive because a modeling object can be defined by data corresponding to any combination of these systems. Thus, a system is very likely to be both a CAD and a PLM system at the same time.

[0051] A CAD system also means any system that is at least applicable to designing a modeling object based on a graphical representation of the modeling object, such as CATIA. In this case, the data defining the modeling object includes data that allows for the representation of the modeling object. A CAD system can provide a representation of a CAD modeling object, for example, using edges or lines (with faces or surfaces in some cases). The lines, edges, or surfaces can be represented in various ways, such as non-uniform rational B-splines (NURBS). Specifically, a CAD file contains specifications that can generate geometry, which in turn allows for the generation of a representation. The specifications of the modeling object can be stored in a single CAD file or multiple CAD files.

[0052] In an example, each 3D scene may represent a real space, such as a real indoor space. For example, the space represented by the 3D scene may be the space of a residence (e.g., a house or an apartment), such as the space of a kitchen, a bathroom, a bedroom, a living room, a garage, a laundry room, an attic, an office (e.g., a personal or shared one), a meeting room, a children's room, a nursery, a corridor, a dining room, and / or a library (this list may include other types of spaces). Alternatively, the space represented by the 3D scene may be other indoor spaces, such as a factory, a museum, and / or a theater. Alternatively, the space represented by the 3D scene may be an outdoor scene, such as a garden, a terrace, or an amusement park.

[0053] Each object in each 3D scene may represent the geometry of a real object located in the real space represented by the 3D scene. The real object may be manufactured in the real world after its virtual design is completed (e.g., using a CAD software solution or a CAD system). The 3D scene may include, for example, one or more furniture objects, such as one or more chairs, one or more lamps, one or more cabinets, one or more shelves, one or more sofas, one or more tables, one or more beds, one or more sideboards, one or more nightstands, one or more desks, and / or one or more wardrobes. Alternatively or additionally, the 3D scene may include one or more decorative objects, such as one or more accessories, one or more plants, one or more books, one or more frames, one or more kitchen accessories, one or more mats, one or more lamps, one or more curtains, one or more vases, one or more carpets, one or more mirrors, and / or one or more electronic objects (e.g., a refrigerator, a freezer, and / or a washing machine).

[0054] The usage method may include, during the process of space design (i.e., effective layout) in real life, which may include, after the usage method is executed, using the generated 2D images to show the space to be laid out. For example, such a showing may be for a user such as the owner of the house where the space is located. The user may use the generated 2D images to decide whether to acquire one or more objects within the 3D scene, and the display of one or more objects in the space may assist the user in making a choice. During the real space design process, the usage method may be repeated to determine multiple 2D images of the space. Repeating the usage method may be used to show the complete virtual interior of the space (i.e., including several 2D images of the space), and / or to obtain 2D images of the 3D scene with different styles and / or object appearances.

[0055] Alternatively or additionally, spatial design in real life can include performing similarity-based 3D object retrieval from a catalog to be placed at the bounding box location using the generated 2D images. Due to the authenticity of the generated 2D images, the trained function can achieve this. For example, spatial design in real life can include a user defining the layout of a given 3D scene by placing a 3D bounding box. Then, spatial design in real life can include using the trained function (e.g., by repeating the usage method discussed previously) to generate several 2D images of the given 3D scene. Then, spatial design in real life can include the user selecting one of the generated 2D images. For example, the user may particularly appreciate the appearance and / or style of one of the generated 2D images and wish to provide the most similar 3D object in the catalog for the given 3D scene (i.e., replace the bounding box with an actual 3D object). In this case, spatial design in real life may include, for each object in the generated 2D image, deriving the position of each object in the generated 2D image from the defined layout, cropping the object in the image, calculating the image embedding of the object (e.g., using a pre-trained language-image model such as the model described by Radford et al. in the paper "Learning transferable visual models from natural language supervision" at the International Conference on Machine Learning PMLR 2021, hereinafter referred to as CLIP), comparing the image embedding of the object with the image embeddings in the catalog to obtain the most similar object. Spatial design in real life can include replacing the bounding boxes in the 3D scene with the most similar objects in the catalog obtained for each object.

[0056] Alternatively or additionally (e.g., before being shown), the spatial design process in real life can include filling a 3D scene representing the space with one or more new objects by modifying the layout of the 3D scene (which may initially be empty, e.g., partially empty). The filling can include, for each new object, repeating the steps of determining the size and positioning of the bounding box representing the new object as described previously. Thus, the generated 2D images can include the new objects added to the 3D scene by modifying the layout. The spatial design process in real life allows for the creation of a richer and more enjoyable environment (for animation, advertising, and / or generating virtual environments, e.g., for simulation). The spatial design process in real life can be used to generate virtual environments. The spatial design process in real life can include being in a general process that can include repeating the spatial design process in real life for several 3D scenes, thereby showing several 3D scenes including objects.

[0057] Alternatively or additionally, the real-life spatial design process may include, after executing the method, physically arranging (i.e., in reality) the space such that the design of the space matches the 3D scene shown in the generated 2D image. For example, the space (without the objects represented by the input 3D scene) may already exist in the real world, and the real-life spatial design process may include positioning a real object (i.e., the object represented by a bounding box of the layout) represented by an object of the 3D scene within the already existing space (i.e., in the real world). The bounding box of the object may have been added by the user to the layout of the 3D scene. The real object may be positioned according to its position in the 3D scene of the bounding box. The real-life spatial design process may repeat this process in order to position different real objects within the existing space. Alternatively, at the time of executing the method, the space may not yet exist. In this case, the real-life spatial design process may include constructing the space (i.e., including filling this space with real objects) according to the 2D image of the generated 3D scene (i.e., by placing real objects at the positions of the bounding boxes representing them in the 3D scene layout). Since this method improves the positioning of 3D objects in the 3D scene, this method also improves the construction of the space corresponding to the 3D scene, thereby increasing the productivity of the real-life spatial design process.

[0058] Now, the obtaining of the data set is discussed.

[0059] The data set includes a plurality of 2D images of the 3D scene (e.g., more than 50,000 2D images, e.g., of the same type of space). For each 2D image, the data set also includes the 3D scene (e.g., only the layout of the 3D scene) imaged (e.g., in part) in the 2D image, and the viewpoint from which the 2D image was taken (e.g., the coordinates of the viewpoint within the 3D scene). Each 2D image in the data set may be taken for a corresponding (i.e., different) 3D scene. Alternatively, the data set may include 2D images of the same 3D scene, e.g., images taken from different viewpoints and / or under different illuminations. The data set may also include, for each 2D image, information indicating the layout of the 3D scene and its viewpoint (e.g., a table including lines, each line including a 2D image reference, a reference to the corresponding 3D scene layout, and the coordinates of the 2D image viewpoint).

[0060] The 2D images of the data set may be photorealistic 2D images of the 3D scene generated (e.g., by a designer) before executing the method. For example, these 2D images may include perspective, occlusion, and / or illumination factors. To achieve such rendering, the 2D images of the 3D scene in the data set may have been manually remade (at least in part, e.g., where it is difficult to render due to perspective, occlusion, and / or illumination factors) by the designer.

[0061] The dataset can be stored in a database. Obtaining the dataset can include retrieving the dataset from the database. Then, obtaining the dataset can include storing the retrieved dataset in a memory. After recording, a machine learning method can perform training of a function based on the recorded dataset. Alternatively, obtaining the dataset can include providing access to the dataset in the database. In this case, the machine learning method can use this access to perform training of the function.

[0062] The space of the 3D scene representation in the dataset may or can exist in the real world (already existing at the time of or after obtaining the dataset). For example, the space can be a real space in the real world (in terms of layout), and objects can be placed within these real spaces at positions specified in the layout of the 3D scene included in the dataset. The 3D scene can represent a space that has been designed (e.g., by an interior designer) and then realized in the real world (i.e., multiple 3D scenes correspond to spaces of virtual designs that have been or may be reproduced in people's homes). In an example, each space represented in the dataset is of the same type. For example, all the spaces represented by the 3D scene and in the dataset may be a kitchen, bathroom, bedroom, living room, garage, laundry room, attic, office (e.g., personal or shared), meeting room, children's room, nursery, corridor, dining room, or library (this list may include other types of spaces). In this case, the layout obtained during the execution of the usage method can be the layout of the 3D scene, which is of the same type as the layout type in the dataset. It allows for generating more realistic 2D images and improves the stability of the usage method. Alternatively, the dataset can include different types of spaces. In this case, the output domain of the generative image model is larger, and the number of spaces represented in the dataset may be higher. The training of the generative image model may also take longer.

[0063] In an example, the layout of each 3D scene can include a set of bounding boxes representing the objects in the 3D scene. Each bounding box can be rectangular in space and can encapsulate the outer contour of the object it represents. For each bounding box, the layout can include parameters representing the position, size, and orientation of the object represented by the bounding box in the 3D scene. For example, for each bounding box, the layout can include parameters representing the position of the bounding box (e.g., the coordinates of the corners or the center of the bounding box), parameters representing the size of the bounding box (such as the width, depth, and height of the bounding box), and parameters representing the orientation of the bounding box (e.g., the rotation with respect to each axis of the global reference system). Optionally, for each bounding box, the layout can include parameters representing the class of the object represented by the bounding box. The classes of the objects can be pre-determined and can each correspond to the type of the object it represents. The classes of the objects can be the types of decorative and functional objects discussed above.

[0064] The layout of each 3D scene may also include the boundaries of the 3D scene. For example, the boundaries of a 3D scene may be represented by a corresponding set of points, e.g., corresponding to the corners of the 3D scene or sampled along the walls of the 3D scene. The layout of the 3D scene may include the coordinates of the points in its corresponding set.

[0065] Training of the function may include training a scene encoder and a generative image model to generate 2D images of a dataset when they take as input the corresponding layouts and viewpoints included in the dataset (e.g., in a supervised manner). For example, the scene encoder and the generative image model may each include their respective parameters (e.g., weights), and supervised training may include determining the values of these respective parameters so that they best reproduce the 2D images of the dataset when taking as input the corresponding layouts and viewpoints included in the dataset. The supervised training of the function may include training the scene encoder and the generative image model together (i.e., their respective parameters may be determined simultaneously or in the same process).

[0066] In an example, the scene encoder may include a layout encoder configured to encode a set of bounding boxes. For each bounding box of the set (e.g., visible or invisible from the viewpoint), the layout encoder may take as input parameters representing the position, size, and orientation of the object represented by the bounding box in the 3D scene. Optionally, for each bounding box, the layout encoder may also take as input a parameter representing the class of the object represented by the bounding box. These parameters may be the parameters included in the 3D scene layout as described above. The layout encoder may infer these parameters from the layout taken as input by the function. The layout encoder may be configured to output a vector (hereinafter referred to as a "layout vector") embedding the parameters.

[0067] For example, the layout encoder may include a position encoding module that takes as input the parameters of the bounding box and outputs a layout vector. The position encoding module may be configured to deterministically increase the dimension of the scalar values of the parameters taken as input. For example, for each bounding box, the position encoding module may be configured to output a position vector representing the position and size of the bounding box and an orientation vector representing the orientation of the bounding box. Optionally, the layout encoder may also include a first multi-layer perceptron configured to increase the dimension of the orientation vector output by the position encoding module. When the layout encoder also takes as input the class of the object, the layout encoder may also include a second multi-layer perceptron that, for each bounding box, takes as input a parameter representing the class (or category) of the object represented by the bounding box and outputs a class vector. The layout encoder may also include a concatenation layer configured to concatenate the vectors output by the position encoding module, the first multi-layer perceptron, and / or the second multi-layer perceptron and output the aforementioned layout vector.

[0068] In an example, the scene encoder may further include a floor encoder configured to encode the boundaries of the 3D scene. As described above, the boundaries of the 3D scene may be represented by corresponding sets of points, and the floor encoder may take the corresponding sets of points as input and output a floor vector. For example, the floor encoder may include a PointNet model (e.g., as described in the paper "PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation" by Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas in CVPR 2017) and an optional multi-layer perceptron (MLP). The PointNet model is configured to encode the corresponding sets of points, and the multi-layer perceptron is configured to take the output of the PointNet model as input.

[0069] In an example, the scene encoder may further include a camera encoder configured to encode the viewpoint. For example, the viewpoint may include the parameters of the camera that generates the 2D image (e.g., camera position, field of view, and pitch angle). The scene encoder may be configured to take these camera parameters as input and output a camera vector. For example, the scene encoder may include a position encoder and an optional multi-layer perceptron, and the position encoder is configured to take the camera parameters as input.

[0070] In an example, for each given 2D image of a given 3D scene in the dataset, the dimensions and positions of the objects represented by the bounding boxes in the layout of the given 3D scene are defined in a coordinate system based on the position and orientation of the camera that captured the given 2D image. Thus, the camera position is already encoded in the layout vector output by the layout encoder. Therefore, the viewpoint may include only two scalar values representing the field of view and the camera pitch angle, respectively. This reduces the number of learning parameters, improves robustness, and thus promotes the convergence of the function.

[0071] In an example, the scene encoder may further include a Transformer encoder. The Transformer encoder takes as input the concatenation of the set of bounding boxes encoded by the layout encoder (i.e., the layout vector), the viewpoint encoded by the camera encoder (i.e., the camera vector), and the 3D scene boundary encoded by the floor encoder (i.e., the floor vector). The Transformer encoder outputs a scene encoding tensor. For example, the Transformer encoder may include a Transformer model configured to take vectors as input to form a sequence of tokens represented as a tensor (the scene encoding tensor). The vectors taken as input by the Transformer model may be the concatenation of all the previously discussed vectors (i.e., the layout vector, the camera vector, and the floor vector), optionally padded (or supplemented) with one or more "zero" tokens to form a vector of a fixed (e.g., predetermined) size.

[0072] The generative image model can be any generative image model capable of generating 2D images, conditioned on the output scene encoding tensor. The generative image model may be a deep neural network trained on a large image dataset to learn the latent distribution of the training images. By sampling from the learned distribution, such a model can be configured to generate new images with the characteristics in the training dataset. Examples of generative image models for generating 2D images conditioned on the output scene encoding tensor include Generative Adversarial Network (GAN), Variational Autoencoder (VAE), and diffusion models.

[0073] In an example, the generative image model can be a diffusion model. The diffusion model can be configured to generate the output 2D image by iteratively denoising an initial noise image based on the scene encoding tensor output by the scene encoder. Examples of diffusion models include cascade models or latent diffusion models. A cascade model is a model that includes several diffusion modules. For example, one for outputting an image conditioned on the scene encoding tensor, and then including a super-resolution model that upscales the image to a higher resolution. During inference, the diffusion model can generate the output 2D image by iteratively denoising the initial noise image. Each iteration of denoising may include determining a new version of the initial noise image with less noise than the previous version of the initial noise image determined during the previous iteration. The determination of the new version can be based on the prediction of the noise in the previous version.

[0074] Now, the training of the diffusion model is discussed. The training of the diffusion model can be based on the noisy versions generated from the 2D images included in the dataset. For example, the machine learning method can include generating the noisy versions of the 2D images in the dataset (by adding noise to these 2D images), and the training of the diffusion model can be based on the generated noisy versions of the 2D images. Given the scene encoding tensor output by the scene encoder, the diffusion model can be trained to remove the noise added to the 2D images in the dataset. In this case, the training of the diffusion model and the scene encoder can consider a training loss that penalizes the distance between the predicted noise in the generated noisy version and the actual noise in the generated noisy version.

[0075] In an example, the diffusion model can have an architecture that includes a denoiser. The denoiser can include several blocks configured to generate the generated 2D images. At least one of these blocks can be enhanced using the scene encoding tensor through cross-attention. The cross-attention mechanism can be an attention mechanism applied between elements from different sequences. Cross-attention can be applied between the representations returned by the transformer encoder for different tokens and the visual features computed within the denoiser. It improves the ability to learn the visual and spatial dependencies / relationships that exist between the scene features (encoded by the transformer encoder) and their visual representation in the image (generated by the denoiser). Therefore, this mechanism is particularly suitable for generating 2D images of 3D scenes, where the position / concept of the spatiality of objects in the 2D image is crucial.

[0076] In an example, the diffusion model can be configured to operate in the latent space (i.e., it can be a latent diffusion model). In this case, during training, the diffusion model can be trained to denoise the compressed latent representations of the 2D images in the dataset. In this case, the training of the function can include compressing the 2D images in the dataset in the latent space (e.g., a smaller dimension), thereby obtaining the compressed latent representations. The training can be performed based on these compressed latent representations (rather than directly based on the 2D images). During the inference process, the diffusion model can take as input a compressed initial noise tensor of the same dimension as the compressed latent representation (instead of the initial noise tensor). Given the scene encoding tensor output by the scene encoder, the diffusion model can iteratively remove the noise from this compressed initial noise tensor. After that, decompression can be applied to the result to obtain the generated 2D image. Examples of implementations for such compression / decompression include variational autoencoders (VAEs).

[0077] Now discuss the application of the function. The application of the function can initially include the step of forming the input of the diffusion model. When the diffusion model does not operate in the latent space, this step can include sampling an initial noise tensor, for example, having the shape of the 2D image to be generated. The diffusion model can take the initial noise tensor as input. When the diffusion model operates in the latent space, this step can include sampling a compressed initial noise tensor (i.e., having the same dimension as the compressed latent representation). The diffusion model can take this compressed initial noise tensor as input.

[0078] Then, the application of the function can include applying a scene encoder to the obtained layout, thereby outputting a scene encoding tensor. After that, the application of the function can include using the diffusion model conditioned on the output scene encoding tensor to generate a 2D image of the 3D scene. When not operating in the latent space, the diffusion model can iteratively remove noise from the sampled initial noise tensor, or from its compressed representation. When operating in the latent space, the application of the function can also include the step of decompressing the clean latent space (obtained by iteratively applying a denoiser) to return to the image space and obtain the generated 2D image.

[0079] Reference Figures 1 to 14 , discuss implementation examples of machine learning methods and usage methods.

[0080] The trained function is conditioned on the output scene encoding tensor (i.e., has 3D awareness). This allows leveraging the knowledge of the 3D structure and relationships of objects in the scene or environment, resulting in more accurate, natural, and immersive content due to improved consideration of factors such as perspective, occlusion, or lighting. In particular, the trained function only includes a scene encoder and a single conditional diffusion model, which are trained end-to-end specifically for this task. It does not utilize large general pre-trained image synthesis priors, does not include the training phase of several separately trained modules, and does not require training a neural volume renderer or NeRF (Neural Radiance Fields) for each generated scene. Due to the layout of the training samples, 3D awareness is incorporated, so it does not rely on a separate depth estimator that is usually defective and propagates errors, nor does it require a multi-view dataset for training. This usage method also enhances user interaction and controllability.

[0081] The machine learning method and the usage method solve the problem of generating high-quality, user-specified 2D views of a 3D environment that is not composed of prefabricated 3D models and textures, but rather consists of a high-level / abstract description of the environment, i.e., annotated 3D bounding boxes representing elements (or objects) in the scene.

[0082] To this end, the function is trained to generate an image of the scene, given as input a camera viewpoint (position and rotation), a set of annotated bounding boxes (position, size, orientation, and a label representing the object category) representing the objects in the scene, and the corners of the space represented by the 3D scene (i.e., boundary, shape, or floor plan).

[0083] Machine Learning Methods and Usage This technical problem was solved using a deep learning based approach. Its pipeline can be divided into two main stages.

[0084] In the first phase (offline phase), the machine learning method performs supervised training of the function. Given a dataset of pairs of 2D images and corresponding underlying scene annotations (camera viewpoint and field of view, annotated bounding boxes of objects visible or invisible in the scene, and floor corners), the machine learning method consists of training a function that includes a deep learning pipeline consisting of the following components:

[0085] A scene encoder (hereinafter also referred to as a scene layout encoder), comprising:

[0086] oLayout encoder that outputs a vector embedding for each object present in the scene.

[0087] oCamera encoder, which outputs a vector embedding that captures the camera information.

[0088] o Floor encoder that outputs a vector embedding that captures the shape information of the floor.

[0089] o Transform encoder, which takes as input the sequence of concatenated embeddings and outputs a new sequence of representations / embedded sequences.

[0090] A diffusion model (hereafter also referred to as a denoising diffusion model) that is conditioned on a noisy version of the provided 2D image and the scene embedding output by the scene encoder as input. The denoising diffusion model can operate directly in the image space or in the latent space. In this case, the denoising diffusion model may contain a variational autoencoder (VAE).

[0091] The goal of the supervised training phase is to endow the function with the ability to reproduce (or generate) 2D reference images, with the corresponding scene annotations (i.e., layouts) of the 3D scene as input.

[0092] In the second stage (inference stage or online stage), the trained function can be used to generate 2D images. On the one hand, by taking user-defined scene annotations as input, a scene embedding tensor is calculated using a layout encoder, a camera encoder, a floor encoder, and a transformer encoder. On the other hand, a random Gaussian noise image is sampled (either in the image space or the latent space, depending on the nature of the diffusion model). The scene embedding and the denoised image are iteratively fed into the trained denoising diffusion model, which outputs a 2D image corresponding to the desired scene at the desired viewpoint.

[0093] The main advantages of the machine learning method and its usage include:

[0094] · Time efficiency: Once trained, the function (or model) can create high-quality 2D images faster than existing rendering solutions in traditional computer graphics: Generation does not use expensive ray tracing or post-processing effects. Generating a 2D view takes no more than a few seconds on a GPU and less than a minute on a CPU.

[0095] · Creativity and user-driven:

[0096] o The trained function can output scenes containing objects that do not exist in the original dataset without the need to create new 3D models and textures.

[0097] o The trained function allows users to easily and quickly provide abstract 3D layouts without specifying and manually selecting the exact 3D items to be represented (which is usually both cumbersome and time-consuming), thus enabling a new creation workflow.

[0098] · Space efficiency:

[0099] o Once the deep learning model is trained, it occupies approximately a few GB on disk but can generate an infinite number of object variations. In contrast, storing thousands of 3D meshes and their textures is not very space-efficient.

[0100] o Since it can be specialized for relatively small datasets, the 3D-aware diffusion model can be smaller (i.e., have fewer training parameters) than the models used in other methods.

[0101] · Image fidelity: The trained function can output images that take into account the effects of unseen objects (e.g., for an interior scene, a window outside the camera frustum will still affect the output image by influencing the lighting).

[0102] Now, the definitions of certain terms are introduced.

[0103] Deep Neural Networks (DNN) is a powerful set of neural network learning techniques. It is a biologically inspired programming paradigm that enables computers to learn from observed data. In object recognition, the success of DNNs is attributed to their ability to learn rich intermediate media representations, rather than the hand-designed low-level features (such as Zernike moments, HOG, bag of words, SIFT, etc.) used in other methods (min-cut, SVM, Boosting, random forests, etc.). More specifically, DNNs focus on end-to-end learning based on raw data. In other words, by completing end-to-end optimization starting from raw features and ending with labels, it maximally gets rid of feature engineering.

[0104] Generative image models are a type of deep neural network that are trained on large image datasets to learn the latent distribution of the training images. By sampling from the learned distribution, the model can generate new images with the characteristics in the training dataset. GAN (Generative Adversarial Network), VAE (Variational Autoencoder), and diffusion models are widely regarded as the most popular generative image models, and diffusion models are currently considered the state-of-the-art method in this field.

[0105] Diffusion models are a type of deep learning model that can be used for image generation. Its goal is to learn the structure of the dataset by simulating how data points diffuse in the latent space. Diffusion models consist of three parts: the forward process, the reverse process, and the sampling stage. In the forward process, Gaussian noise is added to the training data through a Markov chain. The goal of training a diffusion model is to teach it how to gradually eliminate the addition of noise. This is done in the reverse process, where the diffusion model reverses the noise addition performed in the forward process, thus restoring the data. In the sampling stage, the image generation diffusion model starts from a random Gaussian noise image. After being trained to reverse the diffusion process of the images in the training dataset, the model can generate new images similar to the images in the dataset. It achieves this by reversing the diffusion process, starting from pure Gaussian noise until a clear image is obtained.

[0106] Autoencoders are a neural network architecture used for dimensionality reduction and data compression. It consists of an encoder that maps the input data to a low-dimensional representation and a decoder that reconstructs the original data from the encoded representation. By compressing and reconstructing the data, autoencoders extract meaningful features and perform tasks such as data compression. Variational Autoencoders (VAE) are a special type of autoencoder that combines probabilistic modeling. Instead of learning a deterministic mapping, VAE learns the parameters of the probability distribution on the latent space.

[0107] A Transformer is a deep neural network architecture that has a remarkable ability to perceive the relationships between elements in an input sequence. Due to a mechanism called self-attention, the Transformer enables the model to learn the correlations of each element with other elements and appropriately weigh the context information. The Transformer module takes a sequence as input and outputs a new vector representation of the input data, where the relationships in the input sequence are emphasized.

[0108] Cross-attention extends the self-attention mechanism by allowing the selection of correlations or context information between different sequences. The inputs for cross-attention are two different sequences of the same or different modalities (e.g., text or image). The model learns to process the relevant information from one sequence to another. Cross-attention is suitable when dealing with tasks that involve integrating information from other sources to enhance the model's capabilities.

[0109] In the context of generative AI models for image synthesis, conditioning refers to the process of injecting additional information into the image generation process to obtain results that match user-driven constraints. Conditioning can take various forms, including, for example, text (such as DALL-E 2, Midjourney, or Stable Diffusion) or images (such as ControlNet or semantic segmentation).

[0110] The (3D) bounding box of a three-dimensional (3D) object is the smallest cuboid that encloses the object. Its position, dimensions, and orientation characterize the 3D bounding box. Two opposite vertices are sufficient to fully describe a 3D bounding box.

[0111] "Viewpoint" refers to the perspective or "camera" from which a rendering is captured. It may include four parts: position, orientation, field of view, and pitch angle. The position and orientation of the viewpoint and the bounding box can be defined within a reference frame.

[0112] The term "3D abstract scene" refers to a list of labeled bounding boxes (the layout of a 3D scene), representing the objects in the scene (the labels corresponding to the classes of the objects), the viewpoint, and optionally other elements that can enrich the description of the environment (such as information about the spatial shape). The adjective "abstract" emphasizes that the objects in the scene have no visual representation and are defined without features beyond their bounding boxes and labels.

[0113] One-hot encoding is a technique for representing categorical variables as binary vectors. It involves creating a binary vector for each category, where only one coordinate is set to 1 and the remaining coordinates are set to 0. This representation allows the conversion of categorical data into a numerical format that can be processed by a deep neural network.

[0114] A scene encoder refers to a specialized deep neural network that learns to extract a comprehensive representation from a 3D scene, which may include spatially located objects, layouts, or viewpoints. The scene encoder receives different inputs according to its specific architecture and the user's requirements and generates a high-dimensional vector output. This encoded representation should capture the important features of the scene and serve as a valuable input for subsequent stages of deep learning models.

[0115] Figure 1 A flowchart showing examples of machine learning methods and usage methods.

[0116] The pipeline consists of a diffusion model 100 conditioned by a novel 3D scene encoder 200. Like other deep learning models, it has an offline stage S100 (by performing the machine learning method training function) and an online stage S200 (by performing the usage method to generate 2D images, also known as the inference stage).

[0117] Now, the offline training stage S100 will be discussed in more detail. The goal of this stage is to simultaneously (i) train the scene encoder 200 to produce a comprehensive mathematical representation that can be used for conditioning, and (ii) train the diffusion model 100 to generate an image from noise. The scene encoder 200 takes as input a set of elements that characterize a 3D scene (the layout of the 3D scene) and outputs a scene encoding tensor. The diffusion model 100 takes as input a noisy version of the image to be generated and the scene encoding tensor and outputs a denoised version of the input image. This training is end-to-end: a loss value is calculated and backpropagated to adjust the weights of the diffusion model and the scene encoder. Setting up the training stage may include the following subtasks:

[0118] · Data preprocessing step: Data sampling of the dataset to be processed, especially scene annotations (i.e., layouts), so as to pass them to the scene encoder.

[0119] · Architecture definition step: The scene encoder 200 can return a single fixed-size tensor embedding for the entire scene being rendered. The diffusion model 100 can take as input a noisy version of the image and the scene encoding vector.

[0120] It can return an estimate of the noise added to the image and use it to propose a low-noise version of the input image.

[0121] · Training loss definition step: The training loss function can measure the distance between the predicted noise in the input image and the true noise in the image added through the forward process.

[0122] · Training step: Training can be performed by iterating several times over the dataset (pairs of images and scene annotations).

[0123] Now, the image generation / inference phase S200 will be discussed in more detail. The goal of this phase is to output a rendering that matches the viewpoint, given an abstract 3D scene. For this phase, the methods used may include the following subtasks:

[0124] · Scene embedding vector determination step: Using the scene encoder 200 and the input abstract 3D scene, calculate the scene embedding vector. The abstract 3D scene does not necessarily have to be part of the database (for example, it can be user-created or generated using other techniques).

[0125] ● Random Gaussian noise image generation step: The Gaussian noise image 301 may have the dimensions of the desired final image when training in pixel space, or the dimensions of the VAE latent space when training a latent diffusion model.

[0126] · Iterative denoising step for the generated image: Using the diffusion model 100, first denoise the random Gaussian noise image 301, and then iteratively denoise the output of the diffusion model. The U-Net denoiser (the DNN backbone of the diffusion model) can condition on the scene embedding vector using cross-attention between the U-Net layers and the scene embedding vector. After a fixed number of denoising steps, the final clear image 302 is generated.

[0127] Now, an implementation example of the general framework described above will be discussed. This example focuses on generating indoor scenes.

[0128] Now, the details of the acquisition and content of the dataset used for training functions will be introduced. The data used can be extracted from HomeByMe renderings taken by users (i.e., 2D images created by real users). HomeByMe is a free indoor design application that allows users to model their homes in 3D by selecting and placing furniture from a rich catalog of objects and generate realistic 2D renderings of their spaces. Whenever a high-quality rendering is made in the application, an exhaustive annotation file is saved together with the image. The raw data from this annotation file contains information about the rendering (semantic segmentation maps and / or 2D bounding boxes of visible objects) and information about the 3D scene used for the rendering (3D bounding boxes of objects, the shape of the space, and / or the viewpoint). From this raw data, three elements can be extracted:

[0129] · 3D bounding boxes: For each object in the scene (not necessarily visible in the renderings taken by the user), the annotation file contains a list of other features that describe the object, particularly the following attributes: the class of the object and its 3D bounding box. The raw data retrieved from the file defines the 3D bounding box through two 3D points corresponding to two opposite vertices of the bounding box. There are a total of 174 possible classes in the HomeByMe dataset.

[0130] · Viewpoint: The rendering viewpoint of the user can be saved in the annotation file. In particular, the position, orientation, and field of view of the camera are retrieved and then used in the pipeline.

[0131] · Shape of the space: The shape of the space is stored in the annotation file as a list of 2D points representing the corners of the space.

[0132] The functional example has been trained on approximately 60,000 pairs of bedroom projects (3D annotations, HQ renderings). However, it can be extended to larger datasets with other types of spaces.

[0133] Figure 2 Shows Figure 1 An example of the scene encoder 200. The scene encoder 200 includes a layout encoder 210 configured to encode a set of bounding boxes 201. For each bounding box in the set 201, the layout encoder 210 takes as input parameters representing the position, size, and orientation of the object represented by the bounding box in the 3D scene, and optionally, parameters of the class of the object represented by the bounding box. The scene encoder 200 also includes a floor encoder 230 configured to encode the boundaries 203 of the 3D scene. The scene encoder 200 also includes a camera encoder 220 configured to encode the viewpoint 202. As Figure 1 Shown, the scene encoder 200 also includes a transformer encoder 240. The transformer encoder 240 takes as input the concatenation of the set of bounding boxes 201 encoded by the layout encoder 210, the viewpoint 202 encoded by the camera encoder 220, and the boundaries of the 3D scene 203 encoded by the floor encoder 230. The transformer encoder 240 outputs a scene encoding tensor.

[0134] Now, the offline training phase S100 will be discussed in more detail.

[0135] Before the training step, machine learning can include a data processing step for processing the layout and viewpoint of each 2D image in the dataset.

[0136] The data processing steps may include a first step: processing 3D bounding boxes in each 3D scene. The first step may include converting the original 3D bounding box from a representation based on two opposite vertices to a representation by its position (x, y, z), dimensions (width w, height h, depth d), and orientation. The objects present in the scene may have only one degree of rotational freedom: their rotation around the vertical axis. Thus, the machine learning method may use only one angle θ to define the orientation of the bounding box. In practice, the machine learning method may use different representations and encode the orientation of the 3D bounding box by a pair corresponding to (cos(θ), sin(θ)). This parameterization is mathematically equivalent to the single-value parameterization, but it makes the deep learning models continuous for θ = 0 and θ = 2π. This is beneficial for the convergence of the model. Thus, the processed 3D bounding box is defined by a list of 8 parameters (x, y, z, w, h, d, cos(θ), sin(θ)).

[0137] The data processing steps may include a second step: processing the classes of the objects. Each object in the HomeByMe dataset may be described by a class that provides a broad description (chair, table, or door). There may be a total of 174 classes in the HomeByMe dataset. For feeding into the deep learning model, the class of the object is converted to a one-hot encoded representation in {0, 1} 174 . This step may be done in other ways (e.g., using common techniques such as word2vec to perform vector representation / embedding of the textual class of the object, as described in the paper “Efficient estimation of word representations in vector space” by Mikolov, T., Chen, K., Corrado, G., and Dean, J. in 2013, arXiv preprint arXiv:1301.3781).

[0138] The data processing steps may include a third step: processing the layout boundaries. The third step may include increasing the dimension of the floor points. The original points in the data annotation are 2D points (x, y) because their Z coordinate is implicitly 0. By using 0 as the Z coordinate, the 2D points are converted to 3D points. This step is necessary so that the 3D points are affected by the transformations described later.

[0139] The data processing steps may include a fourth step, processing the coordinates of the bounding box, specifically from world coordinates to camera coordinates (as Figure 3As shown). The original positions and orientations in the annotation file use the world coordinates defined in HomeByMe. To reduce the number of learning parameters, increase robustness, and thus promote convergence, the data processing step can perform a basis transformation from the original world coordinates to a view-point-based coordinate system. In the new coordinate system, the origin of the world is set to the camera position, and the basis vectors are selected so that: the "Z" basis vector remains unchanged, the "Y" basis vector is the projection of the forward vector of the view point onto the plane orthogonal to the "Z" vector, and the "X" vector is orthogonal to the first two vectors. Due to this change of basis, the view point can be fully described by two scalar values: the field of view (FOV) and the pitch angle (the angle between its forward vector and the "Y" basis vector). This change of basis affects the positions and rotations of all objects and points in the scene. This basis transformation is optional but helps with the convergence of the function.

[0140] Now discuss the architecture of the scene encoder in more detail.

[0141] The scene encoder consists of four components: a layout encoder, a camera encoder, a floor encoder, and a transformer module (or transformer encoder) (see Figure 2 ).

[0142] Now discuss the layout encoder. The scalar values (x, y, z, w, h, d, cos(θ), sin(θ)) that describe each bounding box in the scene can be passed through a Positional Encoding (PE) module, which deterministically increases the dimension of the scalar values. In this example, the scalar values are represented by vectors in . Positional encoding is able to generate different representations of the same scalar value, allowing the deep learning model to capture more subtle information when necessary. The use of positional encoding helps improve the convergence of the deep neural network.

[0143] After the positional encoding module, the position and dimension of the bounding box, which were originally described by three scalar values respectively, are described by a 192-dimensional vector (3 × 64 = 192). On the other hand, the rotation, which was originally described by a pair of scalar values, is described by a 128-dimensional vector after positional encoding. To ensure that the model weights the position, dimension, and rotation similarly, the high-dimensional version of the rotation is passed to a multi-layer perceptron, which maps it from to This step improves the convergence of the model.

[0144] The one-hot encoded category is a vector from {0, 1} 174 . To ensure that the weight of the category is similar to the position, dimension, and rotation of the bounding box, the category vector is passed to a multi-layer perceptron, which maps it to a low-dimensional representation in .

[0145] All previously computed vectors are concatenated into a single vector in This vector is a token representing the 3D bounding box of the marker.

[0146] Now discuss the camera encoder. A camera or viewpoint is fully described by two scalar values: the field of view and the pitch angle. Both of these values are sent to a higher dimension using positional encoding and then fed into a multi-layer perceptron that maps them to This vector is a token representing the viewpoint in the scene.

[0147] Now discuss the floor encoder. The floor is represented only by an unordered set of 3D points corresponding to its corners. This representation is ambiguous and cannot be easily interpreted by a deep neural network. As an alternative, the data processing step may include densely sampling points along the walls of the space so that the boundaries of the space are represented by a 3D point cloud, thereby generating a set of points sampled along the boundary. This 3D point cloud is then fed into a PointNet module that outputs an embedding vector. This embedding itself is fed into a multi-layer perceptron that maps the vector to This final vector is a token representing the floor points. The floor encoder improves the quality of the generated images.

[0148] Now discuss the Transformer module. The tokens for the 3D bounding box, the camera, and the floor are all concatenated together to form a sequence of tokens. These tokens are independent of each other. To capture the relationships between different elements in this sequence, a Transformer module is used. Due to the inherent architecture of the Transformer module, its operation can be improved with a fixed input size. However, since the number of 3D bounding boxes in a scene may vary from scene to scene, the sequence constructed by concatenating the outputs of the layout encoder, the camera encoder, and the floor encoder may have a variable length. To be compatible with the Transformer architecture, the concatenation of the vector sequence can be padded with "zero" tokens to make the sequence have a fixed length. Consistent with the distribution of the number of 3D bounding boxes in the dataset, the data processing step can pad the concatenation of the vectors to a length of 50 tokens. Thus, the sequence can be represented as a tensor of. This tensor can be fed into a Transformer module that can output a final scene embedding vector.

[0149] Now discuss the architecture of the diffusion model. The functionality can include one of two versions of the diffusion model: one that acts directly on the image space and another that acts on the latent space of a pre-trained VAE to increase the final image dimension. In the first case, diffusion occurs directly on the pixels in the image, while in the second case, diffusion occurs on the latent version of the image, which is then decoded using a VAE decoder. There is no fundamental difference between these two methods, and not much change is required except for the introduction of the VAE.

[0150] The diffusion model has a conditional generation architecture, characterized by a U-Net backbone consisting of four down blocks and four up blocks. Notably, the last two down blocks and the first two up blocks can be enhanced by cross-attention using the scene embedding vector. The number of up / down blocks and the number of blocks enhanced by cross-attention can vary according to the user's needs and preferences. This configuration provides an optimal trade-off between image quality and training time.

[0151] Now, let's discuss the training loss used for training. Different losses / parameterizations can be used to train the diffusion model. The noise added to the input image during the forward process is known as it is added deterministically. During training, the diffusion model attempts to predict the noise added to the image. The loss used in this case is the mean squared error (MSE) between the true noise ∈ and the predicted noise ∈ θ (where the prediction is based on the model parameters θ).

[0152]

[0153] Alternatively, other commonly used diffusion training parameterizations / losses can be used interchangeably. For example, the v-prediction parameterization with a minimum signal-to-noise ratio weighting value of 5.0 can achieve a good trade-off between image quality / resolution / computation (e.g., as discussed by Tiankai Hang et al. in their paper "Efficient Diffusion Training via Min-SNR Weighting Strategy" in ICCV 2023).

[0154] Now, let's discuss the generation / inference phase S200 in more detail. During inference, the diffusion model can be configured using different techniques. For example, two different sampling processes can be used: the Denoising Diffusion Probabilistic Model (DDPM) and the Denoising Diffusion Implicit Model (DDIM). For example, DDIM can provide an optimal balance between inference speed and image quality. When generating an image of a given 3D abstract scene, it may take approximately 1 second on an NVIDIA RTX A6000 GPU.

[0155] Now refer to Figures 4 to 12 Discuss the example of the result.

[0156] Figure 4Shows a first example of the layout and viewpoint of a 3D scene. The figure shows the bounding boxes of the objects that encapsulate the space of the 3D scene representation. The bounding boxes of the objects can be colored according to their categories. These layouts and viewpoints are used as inputs to a trained function that generates Figure 5 the 2D image shown. In the generated 2D image, the conditional elements are well represented. Off-screen objects also have an impact on the generated content (e.g., the lighting effect on the bed in front of the space window). The 3D bedroom layouts for qualitative results and quantitative evaluation come from a separate set that was not used in the training distribution. This helps to ensure the robustness and generalization ability of the model.

[0157] Similarly, Figure 6 shows a second example of the layout and viewpoint of a 3D scene, Figure 7 showing the resulting 2D image generated by the trained function.

[0158] Figure 8 shows a third example of the layout and viewpoint of a 3D scene, Figure 9 showing the resulting 2D image generated by the trained function. After that, the user manipulates the input 3D layout, for example, by removing the lamp on the left bedside table, and the function can generate Figure 10 another image as shown.

[0159] Figure 11 shows a fourth example of the layout and viewpoint of a 3D scene, Figure 12 showing the resulting 2D image generated by the trained function.

[0160] The trained function undertakes the dual challenges of generating 2D images from 3D scenes. First, the diffusion model follows its 3D conditions to ensure that objects appear in the expected positions. In other words, when given a viewpoint and a set of objects in 3D space, the model outputs objects in positions similar to those of a traditional 3D renderer. Second, in addition to accurately placing objects, the trained function also generates recognizable or correct objects. Therefore, the trained function is able to achieve high local image fidelity, resulting in high-quality output images.

[0161] The results of the quantitative evaluation for assessing the trained function are now presented. Metrics for evaluating 3D conditioning and local image quality are used to evaluate the 2D images generated by the trained function. The purpose of this metric is to determine whether the object is correctly placed in the 2D image and is recognizable. It requires the ground truth image and the generated image corresponding to the same viewpoint. It makes use of the CLIP model (as discussed in the paper "Learning Transferable Visual Models From Natural Language Supervision" by Radford, Alec, et al. in the Proceedings of the International Conference on Machine Learning PMLR in 2021). Contrastive Language-Image Pre-Training (CLIP) is a multimodal foundation model for learning the joint latent space of text and images. It has been trained on hundreds of millions of pairs (text, image), and thus is able to relate complex visual concepts to natural language descriptions. The embedding computed by the text encoder of CLIP from a text prompt will have a high cosine similarity to the embedding computed by its image encoder from an image semantically close to the prompt. Due to its powerful zero-shot capabilities, CLIP can perform multiple tasks such as image classification or open-vocabulary semantic segmentation and has been widely adopted in recent computer vision research work. Its shared latent space allows for the interchangeable use of text and image modalities. The metric is computed by performing the following steps for each object in the ground truth image:

[0162] 1. Crop the region corresponding to the object.

[0163] 2. Compute the CLIP embedding of the crop.

[0164] 3. Use CLIP as a zero-shot classifier for top-k retrieval to check whether the true class of the cropped object is among the top few classes in the retrieval. The CLIP embedding of the parsed text object class is used to compute the classifier key. Cosine similarity is used as the value for comparing the cropped CLIP and the class CLIP.

[0165] 4. If the crop in the ground truth image is recognized, crop the generated image in the same region.

[0166] 5. Compute the CLIP embedding of the generated crop.

[0167] 6. Use CLIP as a zero-shot classifier for top-k retrieval to check whether the class of the generated object is correct.

[0168] 7. The accuracy value is the ratio of the correctly recognized generated crops to the correctly recognized ground truth crops.

[0169] Figure 13Shows the results obtained using this metric on an evaluation set of approximately 100 scenarios. Quantitative evaluation of this function shows that in the top 10 retrieval scenarios (out of 174 classes), 60% of the objects are correctly identified. In the top 25 retrievals, the accuracy rate increases to over 75%. These results indicate that the objects generated by the diffusion model of the trained function are not only accurately placed but also realistic enough in most cases to be recognized by the CLIP zero-shot classifier.

[0170] Figure 14 Shows an example of the system, where the system is a client computer system, such as the user's workstation.

[0171] The client computer in the example includes a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and a random-access memory (RAM) 1070 also connected to the bus. The client computer is also equipped with a graphical processing unit (GPU) 1110, which is associated with a video random-access memory 1100 connected to the bus. The video RAM 1100 is also known as a frame buffer in the art. A mass storage device controller 1020 manages access to a mass storage device, such as a hard disk drive 1030. The mass storage device (suitable for tangibly embodying computer program instructions and data) includes all forms of non-volatile memory, including, for example, semiconductor storage devices (such as EPROM, EEPROM, and flash devices); magnetic disks (such as internal hard disks and removable disks); and magneto-optical disks. Any of the above can be supplemented or incorporated by a specially designed application-specific integrated circuit (ASIC). A network adapter 1050 manages access to a network 1060. The client computer may also include a haptic device 1090, such as a cursor control device, a keyboard, etc. Using a cursor control device in the client computer allows the user to selectively position the cursor at any desired location on a display 1080. In addition, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes a plurality of signal generation devices for inputting control signals to the system. Generally, the cursor control device can be a mouse, and the buttons of the mouse are used to generate signals. Alternatively or additionally, the client computer system may include a touchpad and / or a touch screen.

[0172] A computer program may include instructions executable by a computer, which instructions include means for causing the above system to perform the method. The program may be recorded on any data storage medium, including the memory of the system. For example, the program may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations thereof. The program may be implemented as an apparatus, for example, a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. Method steps may be performed by a programmable processor that performs an instruction program by operating on input data and generating output to perform the functions of the method. Accordingly, the processor may be programmable and coupled to receive data and instructions from, and to send data and instructions to, a data storage system, at least one input device, and at least one output device. An application program may be implemented in a high-level procedural or object-oriented programming language, or, as required, in assembly language or machine language. In any case, the above languages may be compiled or interpreted languages. The program may be a full installation program or an update program. In any case, the application program on the system generates instructions for performing the method. Alternatively, the computer program may be stored and executed on a server in a cloud computing environment that communicates with one or more clients via a network. In such a case, the processing unit executes the instructions contained in the program, thereby causing the method to be performed on the cloud computing environment.

Claims

1. A computer-implemented machine learning method for training a function configured to generate 2D images of 3D scenes, the function including a scene encoder and a generative image model, the scene encoder taking as input the layout and viewpoint of the 3D scene and outputting a scene encoding tensor, the generative image model taking as input the scene encoding tensor output by the scene encoder and outputting a generated 2D image, the machine learning method comprising: - obtaining a data set including 2D images and corresponding layouts and viewpoints of 3D scenes; and - training the function based on the obtained data set.

2. The machine learning method according to claim 1, wherein The layout of each 3D scene includes: - a set of bounding boxes representing objects in the 3D scene; and - the boundary of the 3D scene.

3. The machine learning method according to claim 2, wherein, The scene encoder includes a layout encoder configured to encode the set of bounding boxes, and for each bounding box of the set, the layout encoder takes as input parameters representing the position, size, and orientation of the object represented by the bounding box in the 3D scene.

4. The machine learning method according to claim 3, wherein For each bounding box of the set, the layout encoder also takes as input parameters of the class of the object represented by the bounding box.

5. The machine learning method according to claim 3 or 4, wherein, The scene encoder further includes a floor encoder configured to encode the boundary of the 3D scene.

6. The machine learning method according to claim 5, wherein, The scene encoder further includes a camera encoder configured to encode the viewpoint.

7. The machine learning method according to claim 6, wherein, For each given 2D image of a given 3D scene in the data set, the size and position of the object represented by the bounding box in the layout of the given 3D scene are defined in a coordinate system based on the position and orientation of the camera that captured the given 2D image, and each viewpoint includes the field of view and pitch angle of the camera.

8. The machine learning method according to claim 6 or 7, wherein, The scene encoder further includes a transformer encoder that takes as input the concatenation of the set of bounding boxes encoded by the layout encoder, the viewpoint encoded by the camera encoder, and the boundary of the 3D scene encoded by the floor encoder, and the transformer encoder outputs the scene encoding tensor.

9. The machine learning method according to any one of claims 1 to 8, wherein, The generative image model is a diffusion model.

10. The machine learning method according to claim 8, wherein, The diffusion model has an architecture including a denoiser, and at least one of the blocks in the denoiser is enhanced by cross-attention using the scene encoding tensor.

11. The machine learning method according to claim 9 or 10, wherein, The diffusion model is configured to operate in a latent space and is trained to denoise the compressed latent representation of the 2D images of the data set.

12. A method of using the function after machine learning according to any one of claims 1 to 11, the method of use comprising: - obtaining the layout of a 3D scene; and - applying the function to the layout of the 3D scene to generate a 2D image of the 3D scene.

13. According to the usage method described in claim 12, wherein, The generative image model is a diffusion model, and the application of the function includes: - applying the scene encoder to the obtained layout to output a scene encoding tensor; and - using the diffusion model conditioned on the output scene encoding tensor to generate the 2D image of the 3D scene.

14. A computer program comprising instructions for performing the machine learning method of any one of claims 1 to 11 and / or the usage method of claim 12 or 13.

15. A computer-readable storage medium having recorded thereon the computer program according to claim 14.

16. A system comprising a processor coupled to a memory, the memory having recorded thereon the computer program according to claim 14.