Conditional 3D layout prediction

The method addresses the challenge of generating realistic and diverse 3D scenes by employing a pre-configured machine learning function with conditional dropout and iterative denoising, resulting in improved 3D layout prediction and generation.

JP2026067807APending Publication Date: 2026-04-21DASSAULT SYSTEMES SA +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DASSAULT SYSTEMES SA
Filing Date
2025-09-10
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing 3D scene generation methods using generative adversarial networks struggle to produce realistic and diverse scene arrangements that are semantically consistent with the floor plan and maintain physical meaningfulness, often resulting in unrealistic placements of objects.

Method used

A computer implementation method using a pre-configured machine learning function with conditional dropout, iterative sampling, and noise injection to predict 3D layouts, incorporating a denoising diffusion model for improved realism and diversity.

Benefits of technology

The method enables the generation of more realistic and diverse 3D scene arrangements by leveraging conditional dropout and iterative denoising, allowing for flexible and efficient prediction of plausible 3D layouts, even with a large number of objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026067807000001_ABST
    Figure 2026067807000001_ABST
Patent Text Reader

Abstract

This document provides a method, data structure, and system for predicting 3D layouts. [Solution] A computer implementation method that uses a pre-configured machine learning function with conditional dropout for at least one layout parameter to take an input 3D layout and a given noise level and predict an output 3D layout, comprising: obtaining a set of conditioned inputs; determining one or more conditioned candidate 3D layouts for each conditioned input; determining a plurality of perturbed conditioned candidate 3D layouts; and applying the pre-configured function to each perturbed conditioned candidate. The one layout parameter is dropped out, reconstruction errors are calculated, the reconstruction errors are averaged, and a score is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more specifically, to methods, data structures, and systems related to 3D layout prediction.

Background Art

[0002] Some available solutions for generating 3D scenes involve machine learning techniques such as generative adversarial networks (GANs).

[0003] Current prior art presents significant limitations when attempting to obtain realistic and diverse scene arrangements. A realistic scene arrangement means that the scene composition is semantically consistent between objects and between objects and the floor plan, and physically meaningful. In other words, realistic scene arrangements tend to promote groups of objects with strong semantic relationships between objects and with the floor plan, and tend to prefer to arrange objects in a consistent physical way. Thus, realistic scene arrangements tend to avoid, for example, placing a bed in the kitchen (where the object is semantically inconsistent with the floor plan), placing an oven next to the bed (where the objects are semantically inconsistent with each other), and placing an object without the necessary physical support (e.g., a teacup floating instead of being placed on a table).

[0004] In this context, there remains a need for improved solutions for predicting 3D layouts.

Summary of the Invention

[0005] Accordingly, a computer implementation method is provided that uses a machine learning function, pre-configured to take an input 3D layout and a given noise level, and hereinafter referred to as "Usage." Usage includes obtaining the machine learning function. The 3D layout has a set of layout parameters, including a floor plan, a 3D arrangement of one or more 3D bounding boxes, and a semantic category for each 3D bounding box. Each bounding box is defined in the 3D arrangement by a predetermined set of values ​​for one or more bounding box parameters. The input 3D layout includes a given floor plan, a first 3D arrangement of one or more given 3D bounding boxes, and a given semantic category for each given 3D bounding box. Each bounding box is defined in the first 3D arrangement by a first value for a predetermined set of one or more bounding box parameters. The function is also pre-configured to predict an output 3D layout. The output 3D layout includes a given floor plan, a second 3D arrangement of one or more given 3D bounding boxes, and a given semantic category for each given 3D bounding box. Each bounding box is defined in the second 3D arrangement by a second value of a predetermined set of one or more bounding box parameters. The function is configured to predict a second value of a predetermined set of one or more bounding box parameters that is different from a first value of a predetermined set of one or more bounding box parameters. The function is further pre-configured with conditional dropout for at least one layout parameter, the at least one layout parameter including the floor plan and / or the semantic category for each 3D bounding box. The usage further includes obtaining a set of conditional inputs. Each conditional input includes a value that is different for each conditional input for one of the at least one layout parameters, and a value that is common across the conditional inputs for each of the other layout parameters including the floor plan and the semantic category for each 3D bounding box.The method of use further includes determining one or more conditioned candidate 3D layouts for each conditioned input. Each conditioned candidate 3D layout is the result of iterative sampling using the pre-configured function. The method of use also includes determining multiple perturbed conditioned candidate 3D layouts for each conditioned input. Each perturbed conditioned candidate 3D layout is determined by adding the respective noise to each conditioned candidate 3D layout. The method of use further includes applying the pre-configured function to each perturbed conditioned candidate 3D layout for each conditioned input and each perturbed conditioned candidate 3D layout, where one layout parameter is dropped out, thereby obtaining the respective unconditional output. The method of use further includes calculating the reconstruction error between each conditioned candidate 3D layout and its respective unconditional output for each conditioned input and each perturbed conditioned candidate 3D layout. The method of use also includes averaging the reconstruction errors across the multiple perturbed conditioned candidate 3D layouts for each conditioned input, thereby obtaining a score.

[0006] The aforementioned method of use may include one or more of the following features: The iterative sampling using the aforementioned pre-configured function includes the following iteration: Inject noise into the aforementioned input 3D layout to obtain a perturbed input 3D layout; Applying the pre-configured function at least once to the perturbed input 3D layout to obtain the output 3D layout; and The output 3D layout is used as the input for the next iteration. Here, optionally, the noise has a level that decreases with depth in the iteration; Applying the aforementioned pre-defined function at least once means, in each iteration, including: Applying the pre-configured function to the perturbed input 3D layout to obtain a first output 3D layout; Obtain a first intermediate 3D layout by calculating the gradient step between the perturbed input 3D layout and the first output 3D layout; Applying the pre-configured function to the first intermediate 3D layout to obtain a second output 3D layout; and Obtaining a second intermediate 3D layout by calculating the gradient step between the perturbed input 3D layout and the second output 3D layout, thereby obtaining the final 3D layout; The one or more conditioned candidate 3D layouts include the final result of the iterative sampling; Adding the respective noises to the final result of the iterative sampling includes sampling the noise level and sampling the respective noises according to the sampled noise level; The one or more conditioned candidate 3D layouts include one or more intermediate results of the iterative sampling; Adding the respective noises to each intermediate result of the iterative sampling includes sampling the respective noises according to the noise level of the intermediate iteration of the iterative sampling corresponding to the intermediate result; The method further includes ranking the conditioned candidate 3D layouts based on their respective scores, starting from the lowest score; and / or The aforementioned pre-configured function is parameterized as follows:

number

number

Number

Number

Number

Number

Number

Number

Number

Number

Number

[0007] Furthermore, a method for machine learning machine learning functions used in such applications (hereinafter referred to as the "machine learning method") is provided. The machine learning method includes obtaining a dataset of ground truth 3D layouts. Each ground truth 3D layout represents a scene and includes its respective floor plan, the respective 3D arrangement of one or more 3D bounding boxes (each bounding box is defined by a predetermined set of values ​​of one or more bounding box parameters), and the respective semantic category for each 3D bounding box. The machine learning method also includes obtaining a probability distribution of noise levels. The machine learning method further includes, for each ground truth 3D layout, obtaining each perturbed 3D layout, which can be calculated by perturbing at least one bounding box parameter of at least one 3D bounding box of the ground truth 3D layout. The perturbation includes sampling each noise level based on the probability distribution. The perturbation also includes, for each of the at least one bounding box parameter, sampling each noise value based on the respective noise level, and applying the respective noise value to the respective bounding box parameter. The machine learning method further includes training the function across the dataset based on a loss that penalizes the dissimilarity metric between each ground truth 3D layout and each predicted 3D layout obtainable by applying the function to each perturbed 3D layout. The training is performed with conditional dropout for at least one layout parameter, the at least one layout parameter including the floor plan and / or the semantic categories of each 3D bounding box.

[0008] The aforementioned machine learning method may include one or more of the following features. The aforementioned dissimilarity metric is of the following type.

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0009] Furthermore, a computer program is provided which includes instructions for performing the aforementioned usage method and / or the aforementioned machine learning method, and / or a data structure which includes a machine learning function trained by the aforementioned machine learning method.

[0010] Furthermore, a device is provided that includes a data storage medium on which the aforementioned data structure is recorded.

[0011] The device may form or function as a non-temporary computer-readable medium on, for example, a SaaS (Software as a Service), another server, or a cloud-based platform. Alternatively, the device may comprise a processor coupled to memory, in which the data structure is recorded. Thus, the device may form all or part of a computer system (for example, the device is a subsystem of the entire system). The system may further comprise a graphical user interface coupled to the processor. Non-limiting examples will be described below with reference to the attached drawings. [Brief explanation of the drawing]

[0012] [Figure 1] A flowchart illustrating an example of how to use the service is shown. [Figure 2] A flowchart illustrating an example of a machine learning method is shown. [Figure 3] Here is an example of a machine learning function architecture. [Figure 4] A schematic representation of an example of the disclosed solution is shown below. [Figure 5] Examples of the solutions that will be disclosed are provided below. [Figure 6] Examples of the solutions that will be disclosed are provided below. [Figure 7]An example of the system is shown. [Modes for carrying out the invention]

[0013] Referring to the flowchart in Figure 1, a computer implementation method using machine learning functions is proposed. The method of use includes obtaining a pre-configured (i.e., pre-trained) machine learning function to take in an input 3D layout and a given noise level (S10). The 3D layout has a set of layout parameters, including a “floor plan” parameter, a “3D arrangement of one or more 3D bounding boxes” parameter, and a “semantic category of each 3D bounding box” parameter. Each bounding box is defined in the 3D arrangement by a predetermined set of values ​​of one or more bounding box parameters. The input 3D layout includes a given floor plan (i.e., a given value of the “floor plan” parameter), a first 3D arrangement of one or more given 3D bounding boxes (i.e., a given value of the “3D arrangement of one or more 3D bounding boxes” parameter), and a given semantic category for each given 3D bounding box (i.e., a given value of the “semantic category of each 3D bounding box” parameter). Each bounding box is defined in the first 3D placement by a first value of a predetermined set of one or more bounding box parameters.

[0014] The function is also pre-configured to predict the output 3D layout. The output 3D layout includes a given floor plan, a second 3D arrangement of one or more given 3D bounding boxes, and a given semantic category for each given 3D bounding box. Each bounding box is defined in the second 3D arrangement by a second value of a predetermined set of one or more bounding box parameters.

[0015] The function is configured to predict a second value for a given set of bounding box parameters, which is different from a first value for a given set of bounding box parameters. The floor plan and semantic category of each 3D bounding box are constant variables of the function (i.e., the function does not change their values). In other words, the output of the pre-configured function, which includes the (predicted) second values ​​for a given set of bounding box parameters, is conditional on a given floor plan (value) and a given semantic category (value) for each given 3D bounding box. That is, the output of the pre-configured function is a conditional output. In other words, the conditional output includes a (conditionally predicted) second 3D arrangement of one or more given 3D bounding boxes, where each bounding box is defined in the (conditionally predicted) second 3D arrangement by the (conditionally predicted) second value of a given set of bounding box parameters. The output 3D layout may include the conditional output of the pre-configured function. That is, the output 3D layout may include a given floor plan, a (conditionally predicted) second 3D arrangement of one or more given 3D bounding boxes (each bounding box being defined in the second 3D arrangement by a (conditionally predicted) second value of a predetermined set of one or more bounding box parameters), and a given semantic category for each given 3D bounding box. In this case, the output 3D layout is a conditional 3D layout.

[0016] The function is further pre-configured with conditional dropout for at least one layout parameter. In other words, the function is trained with conditional dropout for the at least one layout parameter, and then the function can be used while dropping out the at least one layout parameter (i.e., the function is applicable even if no value is provided for the at least one layout parameter or a null value is provided). The at least one layout parameter to which the function is pre-configured with conditional dropout includes semantic categories of floor plans and / or each 3D bounding box, for example, consisting of floor plans, or each 3D bounding box's semantic category, or both floor plans and each 3D bounding box's semantic category. The other value (if any) of the at least one layout parameter that is not dropped out is a constant variable of the function. In other words, the output of the pre-configured function, which includes a second (predicted) value for a given set of one or more bounding box parameters, is unconditioned with respect to the at least one parameter (value) that is dropped out. That is, the output of the pre-configured function is an unconditional output. In other words, the unconditional output includes a second (unconditionally predicted) 3D arrangement of one or more given 3D bounding boxes, each bounding box being defined in the second (unconditionally predicted) 3D arrangement by a second (unconditionally predicted) value of a given set of one or more bounding box parameters.

[0017] The method of use further includes obtaining a set of conditional inputs (S20). Each conditional input includes a different value for each conditional input for one of the at least one layout parameters (denoted P*) (i.e., exactly one layout parameter selected within a set called “at least one layout parameter”, which includes the “floor plan” parameter and / or the “semantic category of each 3D bounding box” parameter) and a value common to all the other layout parameters among the floor plan and semantic categories of each 3D bounding box (i.e., each of the “floor plan” parameter and “semantic category of each 3D bounding box” that were not selected).

[0018] The method of use further includes determining one or more candidate 3D layouts for each conditioned input (S30). Each candidate 3D layout is the result of (e.g., unique) iterative sampling using the pre-configured function.

[0019] Iterative sampling involves applying a predefined function one or more times to predict an output 3D layout, starting from an input 3D layout. In other words, iterative sampling with a predefined function involves iterating over the predefined function one or more times.

[0020] The method of use also includes determining a plurality of perturbed 3D layout candidates for each conditioned input (S40). Each perturbed 3D layout candidate is determined by adding its respective (sampled) noise to its respective 3D layout candidate.

[0021] The method of use further includes applying the pre-configured function to each perturbed conditioned candidate 3D layout for each conditioned input and each perturbed conditioned candidate (S50), where one (individual / selected) layout parameter P* (i.e., a layout parameter that is dropped out and varies among conditioned candidates, but is not constant) is dropped out, thereby obtaining each unconditioned output.

[0022] The method of use further includes calculating the reconstruction error between each conditioned candidate 3D layout and its respective unconditional output for each conditioned input and each perturbed conditioned candidate 3D layout (S60). The reconstruction error is the "difference" between each conditioned output and its respective unconditional output in the conditioned candidate 3D layout. In other words, the reconstruction error is calculated by comparing the conditionally predicted second values ​​of a given set of one or more bounding box parameters with the unconditionally predicted second values ​​of a given set of one or more bounding box parameters.

[0023] The method of use also includes averaging the reconstruction errors across the plurality of perturbed conditioned candidate 3D layouts for each conditioned input (S70) to obtain a score.

[0024] This type of usage forms an improved solution for predicting (and potentially ranking) 3D layouts.

[0025] In particular, the disclosed computer implementation uses take a set of pre-configured machine learning functions and conditioned inputs as input and output (i.e., assign) a score to each conditioned candidate 3D layout predicted (by the machine learning function) associated with each conditioned input. The disclosed method, equivalent to a score-based / diffusion method, provides an improved solution for ranking (i.e., classifying) conditioned inputs and each of their one or more conditioned candidate 3D layouts. In other words, the disclosed method makes it possible to determine which conditioned candidate 3D layout is optimal according to knowledge of a pre-configured machine learning function, further configured with conditional dropout on the conditioned inputs.

[0026] Furthermore, this method of use forms a Self-Score Evaluation (SSE) approach, which leverages knowledge of a pre-configured function to select a set of conditioned inputs relevant to the generation of a 3D layout. In fact, the proposed SSE approach enables the capabilities of a pre-configured function for 3D layout generation by efficiently selecting conditioned inputs that are in harmony with the capabilities of the pre-configured function, thus resulting in a more realistic and plausible 3D layout. Thus, the disclosed SSE makes it possible to select conditioned inputs that lead to the most realistic 3D layout using a single pre-configured function (i.e., a single trained model).

[0027] Furthermore, this usage method enables various ways of obtaining conditioned inputs, thus demonstrating improved flexibility and synergy for real-world user-driven applications. In one example, the set of conditioned inputs obtained in S20 may be provided by a user who wants to determine what the best set of conditioned inputs is for solving a particular real-world problem. For example, to determine the best set of objects to be placed in a given floor plan, and / or the best floor plan to place a given set of objects, and / or to determine what the most relevant object(s) should be inserted or removed from a given 3D layout to optimize a rearrangement task (i.e., to get a lower score). Obtaining a set of conditioned inputs (S20) may also involve using an external third-party source, such as a Large-Scale Language Model (LLM), thereby obtaining a set of conditioned inputs generated by the LLM. In other words, obtaining a set of conditional inputs (S20) may include generating a set of conditional inputs by the LLM, thereby allowing a pre-configured function to be combined with the LLM-generated set of conditional inputs (i.e., taken as input), which may optionally be selected via the SSE approach. Generating a set of conditional inputs by the LLM is particularly useful in use cases where a considerable number (e.g., at least 20) of conditional inputs should be provided (i.e., obtained), thus avoiding the tedious operation of inputting a set of conditional inputs for the user of the usage and improving the ergonomics of this usage.

[0028] In usage, the function acquired (and used) in S10 is pre-configured to be a denoiser. In other words, the pre-configured function (i.e., the denoiser) is conditioned on the noise level used to perturb the input 3D layout. Such noise conditioning (during training) gives the pre-configured function a remarkable ability to denoise the input and predict realistic and diverse 3D scene arrangements.

[0029] The functions obtained (and used) in S10 are pre-configured using machine learning methods. Details of these machine learning methods will be described later in the explanation.

[0030] The function configured in this way is applied one or more times. That is, the pre-configured function is employed in an iterative sampling process. Iterative sampling is the application of the pre-configured function one or more times to predict the output 3D layout, starting from the input 3D layout. Recall that iterative sampling with a pre-configured function involves iterating the pre-configured function one or more times. Since the pre-configured function is a denoiser, iterative sampling corresponds to an iterative denoising process, which allows the pre-configured (i.e., trained) function to improve the realism (e.g., natural appearance) and diversity of the predicted 3D layout. Thus, the pre-configured function may form a denoising diffusion model.

[0031] Furthermore, the denoise-based approach improves 3D placement of densely furnished scenes, such as real-life scene arrangements. That is, the trained function predicts more realistic and diverse 3D scenes containing a large number of objects (e.g., at least 20) (compared to autoregressive methods that predict arrangements where objects are inserted sequentially, i.e., one at a time). In one example, the proposed method has the advantage of generating plausible (e.g., realistic and diverse) 3D placements for a well-furnished scene containing at least 20 objects and being scalable up to at least 50 objects.

[0032] Furthermore, the disclosed method can be a time-efficient iterative sampling process with a trade-off between sampling time, which can be set by the user, and the expected quality of the 3D layout. In one example, the usage may support efficient batch processing techniques and / or parallelization capabilities (e.g., on a GPU) for generating 3D scene placements of multiple scenes and / or multiple placements of a single scene in a single iterative sampling process. The function thus configured is also flexible, meaning that different tasks can be performed by the same disclosed usage using a pre-configured function. Such tasks include, but are not limited to, generating and 3D rearranging (partial) 3D placements of a furnished scene containing, for example, at least 20 objects to be placed within a given floor plan.

[0033] For example, in 3D placement generation, the position of the 3D object is arbitrarily initialized to the center of the room, and the rotation and / or dimensions of the 3D object are randomly initialized. In such an example, iterative denoising starts at a sufficiently high noise level and is performed for at least 30 steps, thus producing a fair compromise between the quality of the predicted 3D placement and the sampling time.

[0034] In a partial 3D placement generation application, some 3D objects already have known position and / or dimension and / or rotation values. Therefore, these objects are initialized to their known values, but the objects to be placed have a position initialized to the center of the room and a randomly initialized rotation. At each denoising step, the model output for known 3D objects (i.e., 3D objects with known values ​​for position and rotation parameters) may be replaced with their original position and orientation (i.e., rotation) values. Alternatively, the model output for known 3D objects may be replaced with their perturbed position and rotation values, which are noised at a level corresponding to the current sampling step. In either case, these known objects ultimately converge to their initial values ​​throughout the entire sampling process.

[0035] In other applications, such as 3D relocation, the position and rotation of 3D objects are initialized to their noisy (i.e., perturbed) values. The denoising process may be performed starting from a lower noise level value than that used for the 3D placement generation task.

[0036] Additionally or alternatively, the method of use may include further arranging real-world rooms according to the predicted layout. That is, each 3D object ultimately has a corresponding real-world physical object positioned and oriented in the real-world room according to the predicted layout. Thus, the method of use may reproduce and rank (i.e., create and store a list of ordered scores) a variety of realistic 3D scene arrangements that are feasible in the user's real-world home / apartment. In other words, the method of use is user-driven; that is, the method of use facilitates real-life user interaction (e.g., in a design planner application) to generate and rank 3D layouts that resemble real-world 3D scenes.

[0037] The method of use includes averaging the reconstruction errors across multiple perturbed 3D layouts of conditioned candidates for each conditioned input (S70), thereby obtaining a score. The obtained scores may be real numbers. The method of use may also include comparing the obtained scores corresponding to each conditioned input. Comparing the obtained scores may further include ordering the obtained scores (i.e., ranking them according to a criterion, for example, from lowest to highest score), creating a (digital) list of the obtained ordered scores, and outputting the created list (i.e., rank). The output list may include a list of scores, each score associated with a conditioned input.

[0038] For example, a user of the usage method may obtain a set of conditional inputs (S20), each conditional input including a floor plan designed by the user (e.g., from their apartment) and / or a list of 3D objects listed by the user (e.g., their rooms, i.e., a list of 3D objects of real-world rooms), generate and rank several 3D layouts, and finally select an output 3D layout (e.g., the output with the lowest score, e.g., the output with the 3D arrangement best suited to the purpose of interior design).

[0039] In another example of usage, the user may obtain a set of conditional inputs in S20, each conditional input containing a list of different 3D objects (e.g., one or two 3D objects different from each other), and obtain a ranking regarding the selection of the best conditional input (and thus obtain feedback from usage).

[0040] The aforementioned pre-configured function is configured to take a 3D layout and a given noise level as input. The input 3D layout includes a given floor plan, a first 3D arrangement of one or more given 3D bounding boxes, and a given semantic category for each given 3D bounding box. Each bounding box is defined in the first 3D arrangement by a first value of a predetermined set of one or more bounding box parameters. That is, in a given floor plan, each 3D bounding box (labeled by a semantic category) may be defined by a first value of a spatial attribute that defines its position, dimensions, and orientation in the scene.

[0041] The function is configured to predict an output 3D layout. The output 3D layout includes a given floor plan, a second 3D arrangement of one or more given 3D bounding boxes, and a given semantic category for each given 3D bounding box. Each bounding box is defined in the second 3D arrangement by a second value of a predetermined set of one or more bounding box parameters.

[0042] The function is further configured to predict a second value for a given set of bounding box parameters, which is different from a first value for a given set of bounding box parameters. A given floor plan and a given semantic category for each given 3D bounding box may be constants of the function; that is, the function does not change their values. "Second 3D placement of one or more given 3D bounding boxes" means that one or more given 3D bounding boxes may be placed within the same given floor plan in such a way that the predicted second value of their spatial attributes is different from the first value of the spatial attributes (i.e., the input). In other words, the trained function predicts second values ​​for the position, dimensions, and orientation of 3D bounding boxes, and therefore predicts a second 3D placement of one or more 3D bounding boxes. Other variables of the function may remain constant; that is, the predicted 3D placements may be performed within the same given floor plan and using the same list of given semantic categories.

[0043] In other words, the aforementioned pre-configured function is trained to do nothing but rearrange the 3D bounding boxes (one or more) of the input 3D layout (using repositioning and / or resizing and / or re-orienting).

[0044] The pre-configured function is further configured with respect to the conditioned input using conditional dropout. Because the pre-configured function is trained using conditional dropout, it acquires a remarkable ability to predict 3D layouts both conditionally and unconditionally, that is, with and without the conditioned input provided to the pre-configured function (i.e., dropped out). The ability of the pre-configured function to predict 3D layouts conditionally and unconditionally enables the calculation of scores for conditioned candidate 3D layouts in the SSE approach.

[0045] The term "conditional dropout" refers to a technique used during the training phase to drop out (i.e., ignore) a set of inputs (e.g., a subset) of a machine learning function with a certain dropout probability p. In other words, in each training iteration, a conditional input (or a selected subset of conditional inputs) is dropped out with probability p and replaced with a general input, such as a null vector. Thus, a machine learning function constructed (i.e., trained) using conditional dropout can predict outputs both conditionally (i.e., the trained machine learning function takes conditional inputs into account to predict a conditional output) and unconditionally (i.e., the machine learning function ignores conditional inputs to predict an unconditional output).

[0046] The pre-configured machine learning function acquired in S10 has been previously trained using a computer-implemented machine learning method, which will be described later. The machine learning method features conditional dropout with respect to at least one layout parameter, where the at least one layout parameter includes a given floor plan and / or a semantic category for each bounding box. This means that the function is pre-configured to take an input 3D layout and a given noise level and predict an output 3D layout both conditionally and unconditionally.

[0047] Conditional means that the pre-configured function takes an input 3D layout and predicts a conditional output. The input 3D layout includes a given floor plan, a first 3D arrangement of one or more given 3D bounding boxes (each bounding box is defined in the first 3D arrangement by a first value of a predetermined set of one or more bounding box parameters), and a given semantic category for each given 3D bounding box.

[0048] Unconditional means that the pre-configured function takes an input 3D layout without at least one layout parameter and predicts an unconditional output, wherein the at least one layout parameter includes a floor plan and / or semantic category for each 3D bounding box in a first 3D arrangement of one or more given 3D bounding boxes.

[0049] The method of use also includes obtaining a set of conditioned inputs (S20). The conditioned inputs include individual values ​​for one of the at least one layout parameters and the same values ​​for each of the other layout parameters among the floor plan and the semantic categories of each 3D bounding box. In one example, the at least one layout parameter (i.e., a parameter(s) that can be dropped out) may be a floor plan and / or a semantic category of each 3D bounding box in a first 3D arrangement of one or more given 3D bounding boxes. The set of conditioned inputs may be, for example, a list of one or more semantic categories associated with each 3D bounding box in the first 3D arrangement. The set of conditioned inputs may be, for example, provided by the user and / or generated by the LLM and / or generated by a model individually trained for 3D layout generation.

[0050] The method of use further includes determining one or more candidate 3D layouts for each conditioned input (S30).

[0051] The method of use employs iterative sampling using the pre-configured function acquired in S10. It should be recalled that iterative sampling means applying the pre-configured function one or more times to predict the output 3D layout, starting from the input 3D layout; that is, iterative sampling using the pre-configured function involves iterating the pre-configured function one or more times. Since the pre-configured function is a denoiser, iterative sampling corresponds to an iterative denoising process, which allows the pre-configured (i.e., trained) function to improve the realism (e.g., natural appearance) and diversity of the predicted 3D layout. Thus, the pre-configured function may form a denoising diffusion model.

[0052] Each conditioned candidate 3D layout is the result of iterative sampling (e.g., unique) using the pre-configured function. The result of iterative sampling (e.g., unique) using the pre-configured function is the predicted 3D layout, which is the conditioned candidate 3D layout.

[0053] In one example, the one or more conditioned candidate 3D layouts include the final result of iterative (conditional) sampling, and / or the one or more conditioned candidate 3D layouts include one or more intermediate (e.g., consecutive) results (i.e., results of steps of iterative sampling) of the iterative (conditional) sampling. In other words, the result of (e.g., unique) iterative sampling using the pre-defined function includes the final result (i.e., 3D layout) of (e.g., unique) iterative sampling, and / or one or more intermediate results (i.e., one or more intermediate 3D layouts), where each intermediate result is an intermediate (e.g., consecutive, e.g., between step 10 and step 40) output of each respective application of the pre-defined function in (e.g., unique) iterative sampling. Thus, for each conditioned input, the usage includes determining one or more conditioned candidate 3D layouts using (e.g., unique) iterative sampling, which includes one or more applications of the pre-defined function. The iterative sampling may be iterative conditional sampling, i.e., iterative sampling with a conditioned input. The one or more conditioned candidate 3D layouts may be the final result of a complete (i.e., overall, completed) iterative sampling using the pre-defined function, and / or the one or more conditioned candidate 3D layouts may be intermediate (e.g., consecutive) results of each of one or more applications of the pre-defined function in the iterative sampling.

[0054] The method of use further includes determining a plurality of perturbed conditioned candidate 3D layouts for each conditioned input (S40). Each perturbed conditioned candidate 3D layout is determined by adding its respective (sampled) noise to its respective conditioned candidate 3D layout. Each perturbed conditioned candidate 3D layout corresponds to its respective conditioned candidate 3D layout, where a predetermined set of values ​​of one or more bounding box parameters are perturbed by adding their respective noises.

[0055] In one example, adding the respective noises to the (identical) final result of iterative sampling includes sampling the noise level and sampling the respective noises according to the sampled noise level. Additionally or alternatively, adding the respective noises to each (different) intermediate result of iterative sampling includes sampling the respective noises according to the noise level of the intermediate iteration of the iterative sampling corresponding to the intermediate result.

[0056] The method of use also includes applying the pre-configured function to each perturbed conditioned candidate 3D layout for each conditioned input and each perturbed conditioned candidate 3D layout (S50), where one layout parameter is dropped out, thereby obtaining the respective unconditional output. In other words, the pre-configured function takes as input the perturbed conditioned candidate 3D layout corresponding to each conditioned input (e.g., the perturbed final result of iterative sampling and / or the perturbed intermediate result of iterative sampling), drops out at least one layout parameter (e.g., a list of semantic categories for each 3D bounding box in a 3D arrangement of one or more 3D bounding boxes), and outputs the respective unconditional output. Thus, the method of use includes applying the pre-configured function with one step (i.e., once) of unconditional denoising, i.e., conditional dropout, for each conditioned input and each perturbed conditioned candidate 3D layout, thereby obtaining the respective unconditional output.

[0057] The method of use further includes calculating the reconstruction error between each conditioned candidate 3D layout and its respective unconditional output for each conditioned input and each perturbed conditioned candidate 3D layout (S60). The reconstruction error is the "difference" (e.g., the distance between the outputs of the pre-configured function) between each conditioned output and its respective unconditional output in the conditioned candidate 3D layout. In other words, the reconstruction error is calculated by comparing a conditionally predicted second value for a given set of one or more bounding box parameters with an unconditionally predicted second value for a given set of one or more bounding box parameters. In one example, such a difference may be calculated using a dissimilarity metric. In yet another example, the reconstruction error may be calculated either after or during iterative sampling, depending on whether the one or more conditioned candidate 3D layouts contain the final result of iterative sampling or contain one or more intermediate results of iterative sampling. In the first example, iterative conditional sampling is used to generate conditioned candidate 3D layouts and their respective unconditional outputs, and the reconstruction error is calculated using a dissimilarity metric between each conditional output and each unconditional output in the conditioned candidate 3D layouts. In the second example, instead, in each iteration of iterative conditional sampling (e.g., between step 10 and step 40), each (perturbed) intermediate conditioned candidate 3D layout is taken as input to a one-step unconditional denoise, and thus an intermediate unconditional output is generated. The reconstruction error is calculated in each iteration of iterative conditional sampling using a dissimilarity metric between the intermediate conditional output and each (intermediate) unconditional output in the conditioned candidate 3D layouts.

[0058] The method of use also includes averaging the reconstruction error across the multiple perturbed conditioned candidate 3D layouts for each conditioned input (S70), thereby obtaining a score (e.g., a real number).

[0059] In one example, the method of use may include ranking the conditioned candidate 3D layouts based on their respective scores, starting from the lowest score, i.e., the best conditioned candidate 3D layout is the one with the lowest score.

[0060] The following describes additional optional features of iterative sampling using the aforementioned pre-configured function.

[0061] Iterative sampling using the pre-configured function, i.e., one or more applications of the pre-configured function, may include injecting noise into an input 3D layout to obtain a perturbed input 3D layout, applying the pre-configured function at least once to the perturbed input 3D layout to obtain an output 3D layout (for example, (i) as a direct result of one application of the pre-configured function, or (ii) as a result obtained by processing the result obtained by one application of the pre-configured function, or (iii) as a result obtained by multiple applications of the pre-configured function, where each application starts from the direct result of the previous application of the pre-configured function, or from the result obtained by processing the result obtained by the previous application of the pre-configured function), and using the output 3D layout as input to the next iteration.

[0062] The iterative sampling using the pre-configured function may include noise whose level decreases with depth in the iteration. In particular, the injection of noise into the input 3D layout in each iteration may include noise level scheduling, where the injected noise has a level that can decrease with depth in the iteration. Such noise level scheduling allows for improved quality of the predicted 3D layout. The pre-configured function is optimized to denoise the input 3D layout perturbed with different noise levels during the inference phase. In other words, the pre-configured function is noise-conditional and has the ability to denoise the input 3D layout to generate (i.e., predict) a realistic 3D layout. Furthermore, applying the pre-configured function at least once may include the following optional steps in each iteration of the iterative sampling: Firstly, the pre-configured function may be applied to the perturbed input 3D layout to obtain a first output 3D layout. Secondly, a first intermediate (e.g., middle) 3D layout may be obtained by calculating a gradient step between the perturbed input 3D layout and the first output 3D layout. Next, the pre-configured function may be applied to the first intermediate 3D layout to obtain a second output 3D layout. Finally, a second intermediate (e.g., intermediate) 3D layout may be obtained by computing a gradient step between the perturbed input 3D layout and the second output 3D layout, thereby obtaining the final 3D layout. These optional steps are collectively called second-order sampling steps because they are performed in each iteration of the iterative sampling. Implementing second-order sampling steps improves the generation of accurate and natural-looking 3D scenes while reducing the number of computationally expensive neural evaluations (i.e., the application of the pre-configured function).A perturbed 3D object may be obtained by applying a noise step (i.e., applying noise) to a 3D object placed at an initial position. The perturbed 3D object is placed at each noisy position. Firstly, a function may be applied to the perturbed 3D object to obtain a first model prediction in which the 3D object is placed at a first predicted position. Secondly, a first intermediate position (e.g., intermediate) may be calculated by applying a gradient step between the noisy position and the first predicted position. Next, a trained function may be applied to the 3D object placed at the calculated intermediate position to obtain a second model prediction in which the 3D object is placed at a second predicted position. Finally, a second intermediate position (e.g., intermediate) may be calculated by applying a gradient step between the noisy position and the second predicted position to obtain the final predicted position.

[0063] The aforementioned pre-configured function may be parameterized by a noise-conditional denoiser. The parameterization may be of the following types:

number

number

number

number

number

number

number

number

number

number

number

[0064] Such parameterization facilitates function training and helps the function learn (i.e., capture) the relationship between perturbed and clean constructions.

[0065] Noise-conditional denoiser that parameterizes the aforementioned pre-set function

number

number

number

[0066] Parameterization of a noise-conditional denoiser is equivalent to a noise-conditional score network with a set of trainable parameters θ.

number

number

number

number

number

number

number

number

[0067] In the above formula,

number

[0068] Referring to the flowchart in Figure 2, a computer implementation method for machine learning a machine learning function is proposed, namely, a machine learning method for training a function using conditional dropout, and thus obtaining the pre-configured function in step S10 of the computer implementation method. The machine learning method includes obtaining a dataset of ground truth 3D layouts (S80). Each ground truth 3D layout represents a respective scene. Each ground truth 3D layout includes a respective floor plan, a respective 3D arrangement of one or more 3D bounding boxes, and a respective semantic category for each 3D bounding box. Each bounding box is defined by a predetermined set of values ​​of one or more bounding box parameters.

[0069] The machine learning method further includes obtaining a probability distribution of the noise level (S90). In one example, the probability distribution obtained in S90 may be a Gaussian distribution.

[0070] The machine learning method also includes obtaining each perturbed 3D layout for each ground truth 3D layout (S100). Each perturbed 3D layout is a 3D layout that is computed (e.g., computed, e.g., the machine learning method includes such computation) by perturbing at least one bounding box parameter of at least one (e.g., each) 3D bounding box of the ground truth 3D layout (e.g., the machine learning method includes such perturbation). In other words, the machine learning method may include computed at least one (e.g., each) each perturbed 3D layout, and / or retrieved (e.g., on local or remote memory) or received (e.g., from a remote third-party computer system) at least one each perturbed 3D layout retrieved or received, so that each perturbed 3D layout is pre-computed. The perturbation includes sampling each noise level based on the probability distribution (S100a). The perturbation also includes sampling each noise value for each of the at least one parameter based on the respective noise level (S100b), and applying the respective noise value to the respective bounding box parameter (S100c).

[0071] The machine learning method further includes training (and outputting) a function (S110). The function is configured (after training S110) to take an input 3D layout and a given noise level and predict (i.e., output or generate) an output 3D layout.

[0072] Training (S110) is performed across the dataset based on a loss that penalizes the dissimilarity metric between each ground truth 3D layout and each predicted 3D layout obtainable (i.e., obtainable) by applying a function to each of the perturbed 3D layouts.

[0073] Training (S110) is further performed with respect to at least one layout parameter, the at least one layout parameter including a floor plan and / or semantic categories of each 3D bounding box.

[0074] Therefore, the function is configured to be a denoiser, which can denoise the input 3D layout so that the input 3D layout is transformed into a more realistic predicted output 3D layout. A function configured in this way is also flexible, which means that, as mentioned above, the trained function can be used to perform different tasks.

[0075] Functions trained with S110 can also output predicted 3D layouts conditionally or unconditionally, thanks to dropout during training. Furthermore, machine learning / training with conditional dropout reduces overfitting of the training layout.

[0076] The neural network architecture of the aforementioned pre-configured function may be as described in European Patent Application No. EP24306557.0, filed on 23 September 2024, which is incorporated herein by reference. In particular, the neural network architecture may follow any example described in European Patent Application No. EP24306557.0.

[0077] For example, the pre-configured function may include an architecture that includes a (noise-recognizing) Transformer. The Transformer has representations of equal length (i.e., embeddings, e.g.)

number

[0078] The aforementioned pre-configured function may further include an encoder for generating representations (e.g., of equal length) that will be taken as input by the Transformer.

[0079] In one example, the function may include a noise encoder that generates a representation of a given noise level.

[0080] The pre-configured function may include a 3D object encoder that generates a first representation of each given 3D bounding box. The 3D object encoder may optionally be configured to generate a representation of each parameter and a representation of a semantic category, and to concatenate all the generated representations.

[0081] The pre-configured function may further include a floor encoder that generates a representation of a given floor plan. The floor encoder may include a sampling module for generating samples of the given floor plan. Sampling of a given floor plan may include, for example, sampling a fixed number of 3D points (e.g., 250 equally spaced 3D points) on the contour (i.e., perimeter) of the floor plan. Such sampling makes it possible to obtain a 3D cloud representation of the floor plan as input to a point-cloud encoder. The floor encoder may also include a point-cloud encoder for processing the samples.

[0082] The aforementioned pre-configured function may further include an MLP that takes a representation of the predicted 3D layout as input and outputs a third representation of each given 3D bounding box.

[0083] Further optional features of the neural network architecture and the pre-configured functions are provided on pages 20, line 3 to 24, line 30 of the specification at the time of filing of European Patent Application No. EP24306557.0.

[0084] Machine learning methods employ a data-driven approach that allows a function to learn placement patterns and relationships between objects and between objects and a constrained environment in order to predict realistic 3D layouts. In other words, a data-driven approach in machine learning methods allows a trained function to learn interactions (i.e., relationships) (i.e., semantic consistency) between 3D objects and interactions (i.e., spatial inference) between 3D objects and a constrained environment solely from the training dataset.

[0085] In one example, the training dataset acquired by S80 may include a variety of realistic ground truth 3D layouts, each representing a 3D scene. 3D scenes may be acquired from digital 3D scene datasets and / or real-world 3D scenes. An example of a digital 3D scene dataset may be the HomeByMe® dataset or any subset thereof, which may include at least 1000 scenes (e.g., 10,000 scenes) with densely arranged furniture (e.g., each containing at least 20 objects).

[0086] Furthermore, the machine learning method is trained on the dataset based on a loss that imposes a larger penalty the greater the dissimilarity between the ground truth 3D layout and each predicted 3D layout. In particular, such a loss may be invariant for permutations of the same 3D objects. Such an option facilitates training and avoids penalizing predicted 3D layouts where the same objects are replaced in the ground truth 3D layout, thus forcing diversity in 3D scene generation.

[0087] In addition, the machine learning method employs a denoising approach to generate predictive outputs. The denoising approach performs better than other classes of existing generative models, such as GAN models. The machine learning method involves injecting different noise levels based on an acquired (S90) probability distribution to perturb samples from the training dataset. Furthermore, the model (i.e., the denoiser) is "noise-conditional" in the sense that it is configured to be applied to input samples with a given value of the noise level (i.e., the noise level is given as input to the model as a "condition"). Such a noise-based approach allows the trained function to learn from a perturbed (i.e., noisy) dataset, and thus provides the trained function with a remarkable ability to denoise the input and predict realistic and diverse 3D scene arrangements. In other words, the machine learning method may train a function that best "denoses" an input 3D layout with an arbitrary (i.e., arbitrary) noise level to predict a realistic (i.e., natural-looking) 3D layout.

[0088] Furthermore, the denoise-based approach improves 3D placement of densely furnished scenes, such as real-life scene arrangements. Specifically, the trained function predicts more realistic and diverse 3D scenes containing a large number of objects (e.g., at least 20) (compared to autoregressive methods that predict arrangements where objects are inserted sequentially, i.e., one at a time). In one example, the proposed method generates plausible (e.g., realistic and diverse) 3D placements for a well-furnished scene containing at least 20 objects, and demonstrates the advantage of being scalable up to at least 50 objects. The dataset acquired in S80 may include ground truth 3D layouts containing at least 20 objects and / or ground truth 3D layouts containing at least 40 objects. The input 3D layouts acquired in S10 may contain at least 20 objects or at least 40 objects, respectively. The proposed solution actually achieves better (i.e., more accurate) results in terms of the physical consistency and realism of the predicted 3D layouts. This is because the denoising approach allows the trained function to simultaneously learn the relationships between 3D objects, that is, to acquire non-trivial interdependencies between 3D objects and between each 3D object and a given floor plan, for example, using a self-attention mechanism. Simultaneously (i.e., all at once) means that during the training of the function S100, a first value of a given set of one or more bounding box parameters for each 3D bounding box in the first 3D arrangement may be input simultaneously. In other words, the function may simultaneously take all 3D objects in the input 3D layout as input. In other words, the trained function captures all spatial and semantic relationships to obtain a realistic and diverse scene arrangement.Similarly, the second values ​​(i.e., the values ​​predicted by the function) of a given set of one or more parameters for each 3D bounding box in the second 3D arrangement may also be output simultaneously (for example, instead of outputting one object at a time, one after the other). Such simultaneous processing of one or more parameters for each 3D bounding box corresponds to better object grouping, i.e., the function's ability to identify objects that can be associated together in the predicted 3D arrangement.

[0089] The pre-configured function, trained according to the machine learning method described above, acquired in S10, and applied in the computer implementation method, takes a 3D layout and a given noise level as input.

[0090] Each 3D layout is a set of data containing a given floor plan, the 3D arrangement of one or more 3D bounding boxes, and the respective semantic category for each 3D bounding box. In other words, a 3D layout represents the arrangement of one or more 3D bounding boxes within a given floor plan. The 3D bounding box of a 3D object is the smallest rectangular cuboid that encloses the 3D object, with or without orientation constraints (such as the constraint that the cuboid must have faces parallel to the horizontal plane). Thus, a 3D bounding box is characterized by its spatial attributes (i.e., its position, its dimensions, and optionally its (unconstrained) orientation parameters) and its semantic category (i.e., the class of the object, e.g., those with the same function, e.g., books, chairs, etc.). A given set of one or more bounding box parameters may describe the spatial attributes of a 3D object. Each object spatial attribute may have a separate real-world interpretation. In one example, a given set of one or more bounding box parameters may include 3D position coordinates, three dimensions (i.e., height, depth, and length), and at least one parameter representing the object's orientation (e.g., cosine and sine of the angle around the vertical axis). Thus, a given set of one or more bounding box parameters may include or consist of eight parameters. The use of 3D bounding boxes captures the three-dimensional positioning of a 3D object. Therefore, thanks to the use of 3D bounding boxes, a trained function predicts accurate and realistic 3D positioning and sizing of a 3D object. In particular, the trained function, and the method of its subsequent use (i.e., the computer implementation method of obtaining such a pre-set function in S10), predicts a 3D layout that exhibits physically consistent positioning in three dimensions, and thus avoids subtle defects that impair the perceived validity of the entire scene, such as overlapping, floating or out-of-bounds objects, inaccessible areas, and inconsistent object positioning.

[0091] A floor plan is data that describes the plan view of a scene where 3D objects can be placed; that is, it represents the corners of a room. Therefore, the floor plan sets the boundaries of the 3D scene placement and conditions the 3D output layout. In machine learning methods, the floor plan may be obtained in S80 from an external 3D database. During training, the floor plan may be rotated at random angles along the vertical axis. In usage methods, the floor plan input in S10 may be imported from the real world by 3D scanning technology.

[0092] Similarly, 3D objects may be obtained in S80 from an external database and / or online catalog in the machine learning method. During inference, the 3D objects input in S10 may be imported from the real world. A given noise level is sampled from the obtained probability distribution in the machine learning method (S100a). In one example, the probability distribution obtained in S90 may be a Gaussian distribution.

[0093] During training, noise levels are introduced to perturb the dataset, i.e., for each ground truth 3D layout, each perturbed 3D layout is obtained and / or calculated. High levels of noise mean that the perturbed 3D layout is "far" from the ground truth 3D layout, and low levels of noise mean that the perturbed 3D layout is "close" to the ground truth 3D layout. In other words, the machine learning method includes obtaining each perturbed 3D layout (S100) by sampling different noise levels based on a probability distribution (S100a) for each ground truth 3D layout, and sampling different noise values ​​based on their respective noise levels (S100b) for each of the at least one parameter and applying their respective noise values ​​to each parameter of one or more 3D bounding boxes in the 3D arrangement (S100c). Thus, the function acquires the ability to predict 3D scenes perturbed at different noise levels.

[0094] Acquisition of at least one (e.g., each) perturbed 3D layout (S100) may include perturbing at least one parameter of at least one (e.g., each) 3D bounding box of the ground truth 3D layout, or retrieving the results of such perturbations (e.g., on local or remote memory) or receiving them (e.g., from a remote computer). Such perturbation includes sampling each noise level based on a probability distribution within a real interval (S100a). The noise level is a positive (e.g., real) number. The noise level may be the magnitude to which the parameter (i.e., spatial attribute) of the 3D bounding box is perturbed. The noise level may be the absolute value of a scalar extracted from the probability distribution. In one example, the noise level is a Gaussian distribution

number

number

number

number

number

[0095] The acquired ground truth scene dataset (S80) may be further augmented by random rotation of the scene along the vertical axis, and this random data augmentation may help improve training for predicting placement scenes where walls are not aligned with at least one coordinate axis. The function is trained over a training dataset of ground truth 3D layouts (e.g., the HomeByMe® dataset) and is based on a loss that penalizes the dissimilarity metric between each ground truth 3D layout and each predicted 3D layout obtainable by applying the function to each perturbed 3D layout. Each perturbed 3D layout may be supplied as input to the trained function to obtain each predicted 3D layout. Thus, the training loss may evaluate the "distance" between each ground truth 3D layout and each predicted 3D layout. The training loss may favor predicted 3D layouts that are "closer" to the ground truth 3D layouts.

[0096] In one example, the dissimilarity metric may be of the following type:

number

number

number

number

number

number

[0097] Such a dissimilarity metric is therefore equivalent to a chamfer distance, which measures the dissimilarity between sets of bounding boxes. In such an example, the dissimilarity metric is a set of one or more 3D bounding boxes in a ground truth 3D layout.

number

number

number

number

[0098] The dissimilarity metric is a computationally efficient differentiable distance.

number

number

number

number

number

number

number

number

number

[0099] In one example, for each pair of 3D bounding boxes (one in a set of one or more 3D bounding boxes in the ground truth 3D layout, and the other in a set of one or more 3D bounding boxes in the respective predicted 3D layout), the differentiable distance may calculate the Euclidean norm between the values ​​of their spatial parameters (e.g., spatial attributes such as position and orientation). The differentiable distance may also assess the dissimilarity between pairs of 3D bounding boxes in terms of dimensions and / or semantic categories. Thus, the dissimilarity distance may be named a “semantic-aware dissimilarity distance” (e.g., semantic-aware chamfer distance) because it is aware of (i.e., takes into account) the semantic categories associated with each 3D bounding box when assessing the dissimilarity between them. As a result, a penalty may be applied if pairs of 3D bounding boxes do not share the same spatial dimensions and the same semantic categories. The penalty parameter K may, in one example, be set higher than 10^4 or 10^6, for example, K = 10^8.

[0100] The underlying loss under which the function training S110 across the dataset is performed is of type

number

number

number

number

[0101] The implementation of the proposed solution is described below.

[0102] Figure 3 shows an example of a machine learning function architecture.

[0103] Referring to Figure 3, the machine learning function may feature a Transformer encoder denoiser network that takes as input a learned encoded representation of a noise level (i.e., magnitude) σ used to perturb the input scene (i.e., is conditioned on it) (therefore the denoiser is eligible as noise-aware). The 3D objects of the input scene are perturbed with some of its features (e.g., position and rotation attributes, or position, rotation, and bounding box dimensions, etc.), as are additional scene-level conditioning features such as a room floor plan / shape. It may output a predicted clean 3D object layout.

[0104] Referring to Figure 3, an example of a denoising architecture design is described below.

[0105] An example implementation of the deep architecture may consist of multiple trainable components: a noise encoder (i), a 3D object encoder (ii), a floor encoder (iii), a noise-recognizing Transformer encoder (iv), and / or a final MLP (v) that outputs predicted object position and rotation values.

[0106] The following describes the options for the noise encoder (i).

[0107] The scalar value of the sampled noise level σ is deterministically increased in dimensionality, for example,

number

number

number

number

number

[0108] The following describes the options for the 3D object encoder (ii).

[0109] A scalar value that describes each 3D bounding box in the scene.

number

number

[0110] After the PE module, the positions and dimensions of the bounding boxes, which were originally described by three scalar values, are now represented by a 192-dimensional vector.

number

number

number

[0111] Category C, when one-hot encoded,

number

number

number

number

[0112] All previously calculated vectors may be concatenated to a single vector in

Number

[0113] The options of the floor encoder (iii) will be described below.

[0114] Just to be safe, encoding the floor points of a room conditions 3D layout generation so that the resulting 3D object is located within the floor limits.

[0115] The 3D point cloud of the sampled floor points F may be supplied to a PointNet module, which, for example

Number

Number

[0116] The options of the noise recognition Transformer encoder (iv) will be described below.

[0117] Noise level tokens, 3D object tokens, and floor tokens may all be concatenated to form a sequence of tokens. These tokens may be independent of each other. A Transformer module may be used to capture the relationships between different elements of this sequence. Transformer modules require a fixed input size due to their inherent architecture. However, sequences constructed through the concatenation of outputs may have a variable length because the number of 3D bounding boxes in the scene differs for each scene sample. To be compatible with the Transformer architecture, the sequence may have a fixed length because it contains "zero" tokens (e.g.,

number

number

[0118] The following describes the options for the final MLP(v).

[0119] The new representations calculated for each 3D object by the Transformer can ultimately be passed to the MLP, which represents the predicted "clean" position of each 3D object.

number

Number

Number

Number

[0120] The resulting architecture has a total of 12.2 million trainable parameters.

[0121] The pre-set function is for noisy (i.e., perturbed) object 3D spatial attributes

Number

Number

Number

[0122] However, other use cases of the usage can also be achieved and implemented. For example, another use case of the usage may be the evaluation of a set of candidate floor plans using a given set of object semantic categories.

[0123] Training with conditional dropout may be performed on semantic categorical inputs, so that the set of category candidates can be evaluated in the inference phase using a computer implementation. More precisely, in each training iteration, probabilistic dropout is

number

number

number

[0124] The training phase may be performed on a loss that penalizes the dissimilarity metric between each ground truth 3D layout and each predicted 3D layout, which can be obtained by applying a function to each perturbed 3D layout. In the current implementation, the loss is the chamfer distance.

number

number

[0125] Refer to Figure 4 to explain the inference phase. In Figure 4, the machine learning function (obtained in S10, refer to Figure 1) is labeled DeBaRa.

[0126] The set of conditional inputs may also be a set of C semantic category candidates, each associated with the same given floor plan (S120).

[0127] In another example of implementing the proposed solution, the set of conditional inputs may be a set of floor plans, each with the same given semantic categories.

[0128] For each conditioned input, the method of use may determine one or more candidate conditioned 3D layouts. Each candidate conditioned 3D layout is the result of iterative (conditional) sampling using DeBaRa (S130). In other words, each candidate conditioned 3D layout is a 3D layout sampled from the learned conditional density. That is,

number

number

number

[0129] As a result of the inference phase (i.e., usage), each conditioned candidate 3D layout

number

number

[0130] The method of use may include ranking the conditioned candidate 3D layouts based on their respective scores, for example starting from the lowest score (S160), i.e., the best conditioned candidate 3D layout is the one with the lowest score. Thus, the optimal (i.e., best) conditioned candidate 3D layout

number

number

number

number

number

number

[0131] In other words, the optimal candidate is derived from the density estimate of its corresponding 3D spatial layout provided by the unconditional network.

[0132] Several examples illustrating practical implementations of SSE are described below.

[0133] a. Monte Carlo estimation A detailed pipeline for this first inference procedure is shown in Figure 4.

[0134] Such a first approach of SSE corresponds to conditionally obtaining one or more conditioned candidate 3D spatial layouts for each conditioned input (S120) obtained by iterative conditional sampling (S130) with 50 sampling steps. The next step in the pipeline is to apply the pre-configured function to the determined perturbed conditioned candidate 3D layouts using conditional dropout (thus obtaining their respective unconditional outputs), i.e., performing a one-step unconditional denoising (S140) at level σ (e.g., unbiased Monte Carlo estimation) for each conditioned candidate 3D layout. The one-step unconditional denoising at level σ may be performed multiple times, for example T denoising trials (i.e., evaluation iterations; e.g., T=50), and for each trial, a pair of noise level and noise value.

number

number

number

[0135] In other words, in this first approach, the conditional layout generation stage S130 (i.e., via iterative sampling using a conditional model, each conditional input)

number

number

number

[0136] An example of an algorithm for evaluating conditioned candidates (i.e., the SSE method of the disclosed solution) is detailed below. [Table 1] Also, the loss used to measure the optimal candidate during this inference stage (e.g., chamfer distance)

number

number

number

number

[0137] Furthermore, the chamfer distance loss is the predicted spatial feature

number

[0138] Furthermore, this approach differs from diffusion classifiers in the prior art, which assume a uniform prior probability for the conditioning probability. In other words, this approach assumes a uniform prior probability for the conditioning probability.

number

[0139] b. Score evaluation during sampling This second approach of SSE corresponds to conditionally obtaining one or more conditioned candidate 3D spatial layouts for each conditioned input (S120) at each sampling step of iterative conditional sampling. In other words, this second approach jointly executes the conditional layout generation stage S130 and the unconditional score evaluation stage S140.

[0140] At each sampling step, and at each conditional input

number

number

number

number

number

[0141] In this example of the proposed solution implementation, the reconstruction error is the intermediate conditional output.

number

number

[0142] Here, the conditioned candidate that minimizes the reconstruction error averaged over several sampling steps is the best candidate for ranking. In practice, during 50 steps of iterative sampling, the reconstruction error is calculated only between, for example, step 10 and step 40, because high and low noise levels tend to bias the estimation and significantly impair accuracy. The training object is the chamfer distance.

number

[0143] Additionally, it should be noted that diffusion loss may be calculated object-wise. In such an implementation of the usage, the usage may include identifying one or more elements of the acquired set of conditioned inputs that maximize and / or minimize the reconstruction error of the acquired set of conditioned inputs. One or more elements of the set of conditioned inputs may be floor plans and / or one or more semantic categories. For example, the usage may include identifying a 3D object of a conditioned candidate that maximizes the reconstruction error to be removed or replaced. Similarly, the usage may include identifying a 3D object of a conditioned candidate that minimizes the reconstruction error to be inserted into the current scene configuration.

[0144] Several quantitative and qualitative experiments on the disclosed methods are described below.

[0145] Such quantitative and qualitative experiments of the disclosed methods were performed on the publicly available 3D-FRONT® dataset. The following results were obtained by applying the Monte Carlo estimation method (a) for SSE described above.

[0146] <Binary Classification> The purpose of binary classification is to distinguish perturbed conditioning candidates from the correct ones.

[0147] For this purpose, the denoiser was trained on a 3D-FRONT® bedroom subset, with the aim of evaluating the efficiency of the method in a toy binary classification task. More precisely, binary classification tests the inference method's ability to distinguish a set of ground truth (good) object semantic categories from a randomly perturbed (broken or adversarial) set by replacing one or more object categories with random ones. This binary classification was performed on each of the 162 scenes in the test set, and accuracy percentages were calculated for several settings based on how much the adversarial set was perturbed. T=100 was set as the number of evaluation trials. For each setting, the experiment was repeated approximately 10 times.

[0148] Figure 5 shows the results of the binary classification experiment.

[0149] Binary classification tested six different settings. In the “None” setting, the adversarial set is the ground truth set, i.e., the object categories are not replaced with random ones. Thus, the “None” setting corresponds to classifying the ground truth set against itself. Clearly, the expected accuracy is (approximately) 50%. In the “Single” setting, only one object category is replaced with a random one in the perturbed adversarial set, while in the “All” setting, all object categories are replaced with random ones. The significant difference in accuracy percentage between the “None” and “Single” settings indicates that the disclosed solution for SSE can identify subtle changes in good conditioned inputs. As is evident, the accuracy percentage increases when the corrupted adversarial set contains more randomly replaced object categories compared to the ground truth set, thereby indicating that the disclosed solution for SSE can better identify more corrupted adversarial sets.

[0150] <3D Scene Compositing> A quantitative comparison in the context of 3D scene synthesis is described below. The comparison takes into account quantitative metrics obtained by combining usage methods with four different methods for obtaining a set of semantic categories (Layout GPT, Dataset Random, LLM, and LLM+SSE), as well as three other methods for 3D scene synthesis known from prior art (Layout GPT, ATISS, and DiffuScene). The objective is to employ usage methods to select conditioned candidates (i.e., conditioned inputs) obtained by a Large-Scale Language Model (LLM), each conditioned candidate associated with one or more 3D layouts generated by the usage method.

[0151] For this purpose, the denoiser was trained on the living and dining room subsets of 3D-FRONT®. Training of the 3D layout generation method was performed using conditional dropout against a set of semantic categories.

[0152] The method of use may use different sources to obtain a set of conditional inputs (in this particular case, a set of semantic categories) to generate a 3D layout. For quantitative comparison, the method of use may take the set of semantic categories generated by the following as input.

[0153] <Layout GPT> The method of use may involve taking a set of semantic categories generated by the Layout GPT method as input. The Layout GPT method is available at the following URL as of the priority date of this patent application: https: / / layoutgpt.github.io / .

[0154] <Dataset Random> Alternatively, to generate a test 3D layout, a set of semantic categories may be randomly selected from the training set.

[0155] <Large-Scale Language Models (LLMs)> The method of use may involve randomly selecting a set of semantic categories from a set generated by LLM. More precisely, the method of use may include a Layout GPT method to generate a set of semantic categories, and may involve randomly selecting one set from a set having the same number of objects as the ground truth test scene being considered.

[0156] <Large-scale language models and self-score evaluation (LLM+SSE)> This setting is similar to the LLM setting described above, but instead of randomly selecting a set of semantic categories as described above, the usage method (i.e., the self-score evaluation method) is used first to select the most appropriate set of semantic categories from those (candidate) sets that have the same number of objects as the ground truth scene. Therefore, this setting is directly comparable to the previous setting (LLM) and evaluates the effectiveness of the usage method (SSE).

[0157] The LLM-based method is implemented using the publicly available Llama-3-8B model (available at the following URL as of the priority date of this patent application: https: / / huggingface.co / meta-llama / Meta-Llama-3-8B).

[0158] Since LLM often generates out-of-distribution sets, using DeBaRa (i.e., the aforementioned pre-configured function) in conjunction with the disclosed SSE procedure on the generated semantic categories consistently improves the realism and relevance of the synthesized indoor scene.

[0159] Quantitative metrics for evaluating the realism and diversity of the generated 3D layouts may include the Frechet Inception Distance (FID), Kernel Inception Distance (KID * 1000), and Scene Classification Accuracy (SCA), all calculated on a top-down orthographic rendering. The spatial validity of the generation may be measured by the accumulated out-of-bounds object area (OBA m). 2The results may be further evaluated by reporting the following. All metrics may be calculated across each test subset. FID and KID compare the distribution of visual features extracted from the pre-trained convolutional neural network. SCA measures how well the convolutional neural network distinguishes the real scene (i.e., the ground truth test scene) from the generated one in a binary classification task. Therefore, an SCA score close to 50% is better, meaning that the generated scene is indistinguishable from the real scene.

[0160] The table below shows a quantitative comparison between three methods for 3D scene synthesis known from prior art (e.g., LayoutGPT, ATISS, and DiffuScene) and a pre-configured function DeBaRa combined with four different methods for obtaining a set of semantic categories as described above. [Table 2]

[0161] Figure 6 shows a top-down view of the 3D scenes generated by DeBaRa from several conditioned candidates generated by LLM, along with their associated SSE values. Result S170, with its lower score, clearly corresponds to a more natural-looking and realistic 3D layout compared to results S180 and S190, which correspond to higher SSE values.

[0162] For completeness, generation times are also provided, both with and without the disclosed conditioning evaluation method. Generation times are averaged across a 3D-FRONT® living room test subset. The results are as follows: • Generation of a single layout in 50 sampling steps: 0.488 seconds • Generation and evaluation of 16 candidates over 50 sampling steps and 100 evaluation trials: 0.894 seconds (with batch implementation) • Generating a single layout using DiffuScene: 32.796 seconds

[0163] The time was calculated using a single GPU (NVIDIA RTX A6000). Despite the additional network application steps induced by using the disclosed conditioned evaluation method, the proposed solution provides fast, real-time generation of 3D layouts in less than one second.

[0164] A learning method is a method of machine learning a model, which is a deep generative model. As is known by itself from the field of machine learning, the processing of input by a model involves applying an operation to the input, and the operation is defined by data containing weight values ​​or parameters. Learning a model (e.g., a neural network or a recurrencer) therefore involves determining the values ​​of weights / parameters based on a dataset configured for such learning, and such a dataset may be called a learning dataset or training dataset. For that purpose, a dataset contains data pieces, each forming its own training sample or training example. The training samples / examples represent the variety of situations in which the model will be used after it has been learned. Any training dataset in this specification may contain a number of training samples / examples greater than 1,000, 10,000, 1,000,000, or 1,000,000. In the context of this disclosure, “training is performed across a dataset” means that the dataset is the learning / training dataset of the model and that the weight / parameter values ​​are set based on it. In this disclosure, a training dataset is a dataset of training examples on which a deep generative model is learned / trained. In the implementation, the training dataset consists of hundreds of examples, each corresponding to a different HPP configuration.

[0165] As is well known from machine learning, a neural network may be defined by its architecture, parameters, and hyperparameters. The architecture consists of layers, beginning with an input layer where the number of neurons can be determined by the dimensions of the input data. Following this layer are several hidden layers, each having a given number of neurons and activation functions. These layers and neurons define the depth and width of the network, while the activation functions may introduce nonlinearity into the model. The output layer may have the same number of neurons as the variables in the output data. The interconnections between these layers define the topology of the neural network. The parameters of a neural network are learnable weights and biases, which are determined during the training process. In contrast, hyperparameters are predefined settings that are not learned from the training data. These include the number of hidden layers, the number of neurons per layer, etc. To train a neural network, at least two settings may be defined: firstly, a loss function, which is a metric that measures the error between the training data and the model's predictions, such as the mean squared error (MSE); and secondly, an optimizer that modifies the model's weights and biases during the training process to minimize the loss function. Each optimizer has its own set of hyperparameters.

[0166] The method is computer-implemented. This means that the steps (or substantially all steps) of the method are performed by at least one computer, or any similar system. Thus, the steps of the method are performed by the computer, possibly fully automatically or semi-automatically. In one example, at least some of the triggers for the steps of the method may be performed through user-computer interaction. The required level of user-computer interaction may depend on the anticipated level of automation and be balanced with the need to implement user preferences. In one example, this level may be user-defined and / or predefined.

[0167] A typical computer implementation of the method is to perform the method using a system adapted for this purpose. The system may comprise a processor coupled with memory and a graphical user interface (GUI), where memory stores a computer program containing instructions for performing the method. Memory may also store a database. The memory is any hardware adapted for such storage, and may comprise several physically separate parts (e.g., one for the program and possibly one for the database).

[0168] Figure 7 shows an example of a system, where the system is a client computer system, such as a user's workstation.

[0169] The example client computer comprises a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and random access memory (RAM) 1070 also connected to the bus. The client computer further comprises a graphical processing unit (GPU) 1110 associated with video random access memory 1100 connected to the bus. The video RAM 1100 is also known in the art as a frame buffer. A mass storage controller 1020 manages access to mass memory devices such as a hard drive 1030. Mass memory devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic disks such as internal hard disks and removable disks; and magneto-optical disks. Any of the above may be complemented by or incorporated into specially designed application-specific integrated circuits (ASICs). A network adapter 1050 manages access to a network 1060. The client computer may also include haptic devices 1090 such as a cursor control device and a keyboard. A cursor control device is used in a client computer to allow the user to selectively position the cursor at any desired location on the display 1080. In addition, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes a number of signal generating devices for input control signals to the system. Typically, the cursor control device may be a mouse, with the mouse buttons used to generate signals. Alternatively or additionally, the client computer system may include a sensitive pad and / or a sensitive screen.

[0170] The computer program may include instructions that can be executed by the computer, and such instructions include means for causing the system to perform the method. The program may be recordable on any data storage medium, including the system's memory. The program may be implemented, for example, in a digital electronic circuit, or in computer hardware, firmware, software, or a combination thereof. The program may be implemented as a device, for example, as a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. The steps of the method may be performed by a programmable processor that executes a program of instructions for performing the functions of the method by manipulating input data and producing outputs. Thus, the processor may be programmable and coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to transmit data and instructions to them. The application program may be implemented in a high-level procedural programming language or an object-oriented programming language, or, if necessary, in assembly language or machine code. In any case, the language may be a compiled language or an interpreted language. The program may be a full installation program or an update program. In any case, the application of the program to the system results in instructions for performing the method. Alternatively, a computer program may be stored and executed on a server in a cloud computing environment, which communicates with one or more clients over a network. In such a case, a processing unit executes instructions contained in the program, thereby causing the method to be executed on the cloud computing environment.

Claims

1. A computer implementation method, - Includes obtaining a machine learning function (S10), wherein the machine learning function is: The input 3D layout is pre-configured to capture an input 3D layout and a given noise level, wherein the 3D layout has a floor plan, a 3D arrangement of one or more 3D bounding boxes (each bounding box is defined by a predetermined set of values ​​of one or more bounding box parameters in the 3D arrangement), and a set of layout parameters including a semantic category for each 3D bounding box, and the input 3D layout is: Given floor plan, A first 3D arrangement of one or more given 3D bounding boxes, wherein each bounding box is defined in the first 3D arrangement by a first value of a predetermined set of the one or more bounding box parameters, and Each given 3D bounding box includes a given semantic category, and It is pre-configured to predict the output 3D layout, and the output 3D layout is: The aforementioned given floor plan, A second 3D arrangement of one or more given 3D bounding boxes, wherein each bounding box is defined in the second 3D arrangement by a second value of a predetermined set of the one or more bounding box parameters, and For each given 3D bounding box, including the given semantic category, The function is configured to predict a second value of a predetermined set of one or more bounding box parameters, which is different from the first value of the predetermined set of one or more bounding box parameters; and The function is further pre-configured with respect to at least one layout parameter using conditional dropout, wherein the at least one layout parameter includes the semantic category of the floor plan and / or each 3D bounding box; - Includes obtaining a set of conditional inputs (S20), each conditional input including a distinct value for one of the at least one layout parameters and the same value for each of the other layout parameters of the semantic categories of the floor plan and each 3D bounding box; - Regarding each conditional input: The process includes determining one or more conditioned candidate 3D layouts (S30), where each conditioned candidate 3D layout is the result of iterative sampling using the pre-set function; This includes determining a plurality of perturbed conditioned candidate 3D layouts (S40), each of which is determined by adding a corresponding noise to each conditioned candidate 3D layout; For each perturbed conditioned candidate 3D layout: This includes applying the pre-configured function to the perturbed conditioned candidate 3D layout (S50), where one of the layout parameters is dropped out, thereby obtaining its respective unconditional output; and This includes calculating the reconstruction error between each of the aforementioned conditioned candidate 3D layouts and each of the aforementioned unconditional outputs (S60); and This includes averaging the reconstruction error across the plurality of perturbed conditioned candidate 3D layouts (S70) to obtain a score. method.

2. The method according to claim 1, wherein the iterative sampling using the pre-configured function includes the following iteration: Injecting noise into the aforementioned input 3D layout to obtain a perturbed input 3D layout; Applying the pre-configured function at least once to the perturbed input 3D layout to obtain an output 3D layout; and The output 3D layout is used as the input for the next iteration. Here, the noise has a level that decreases with depth in the iteration.

3. Applying the aforementioned pre-set function at least once, in each iteration, includes the following according to claim 2: Applying the pre-configured function to the perturbed input 3D layout to obtain a first output 3D layout; Obtaining a first intermediate 3D layout by calculating a gradient step between the perturbed input 3D layout and the first output 3D layout; Applying the pre-configured function to the first intermediate 3D layout to obtain a second output 3D layout; and Obtain a second intermediate 3D layout by calculating a gradient step between the perturbed input 3D layout and the second output 3D layout, thereby obtaining the final 3D layout.

4. The method according to any one of claims 1 to 3, wherein the one or more conditioned candidate 3D layouts include the final result of the iterative sampling.

5. The method according to claim 4, wherein adding the respective noises to the final result of the repeated sampling includes sampling the noise level and sampling the respective noises according to the sampled noise level.

6. The method according to any one of claims 1 to 5, wherein the one or more conditioned candidate 3D layouts include one or more intermediate results of the iterative sampling.

7. The method according to claim 6, wherein adding the respective noises to each intermediate result of the iterative sampling includes sampling the respective noises according to the noise level of the intermediate iteration of the iterative sampling corresponding to the intermediate result.

8. The method according to any one of claims 1 to 7, further comprising ranking the conditioned candidate 3D layouts based on their respective scores, starting from the lowest score.

9. The method according to any one of claims 1 to 8, wherein the pre-configured function is parameterized as follows: 【Number 1】 Here, [Math 2] This represents a first 3D arrangement of one or more given 3D bounding boxes, [Math 3] This represents a given floor plan, [Math 4] This is a list of given semantic categories, σ is a given noise level, [Math 5] This is a noise-conditional score network having a set of trainable parameters θ, [Math 6] This is a noise-dependent preconditioning coefficient that modulates the predicted 3D layout, [Number 7] This is a noise dependency coefficient that conditions the noise level within the score network, [Number 8] and [Number 9] These are, [Number 10] and [Math 11] These are two noise-dependent coefficients that scale the result.

10. A computer implementation method for machine learning a machine learning function used in the method according to any one of claims 1 to 9, wherein the machine learning method is: - This includes obtaining a dataset of ground truth 3D layouts (S80), where each ground truth 3D layout represents a scene and includes: Each floor plan, A 3D arrangement of one or more 3D bounding boxes, where each bounding box is defined by a predetermined set of values ​​of one or more bounding box parameters, and The respective semantic categories for each 3D bounding box; - Including obtaining the probability distribution of the noise level (S90); - For each ground truth 3D layout, obtain a perturbed 3D layout that can be calculated by perturbing at least one bounding box parameter of at least one 3D bounding box of the ground truth 3D layout (S100), the perturbation including: Sampling each noise level based on the aforementioned probability distribution (S100a); and For each of the bounding box parameters of the at least one bounding box parameter mentioned above: S100b) sampling each noise value based on the respective noise levels; and Applying the respective noise values ​​to the respective bounding box parameters (S100c); and - Training the function across the dataset (S110) based on a loss that penalizes the dissimilarity metric between each ground truth 3D layout and each predicted 3D layout obtainable by applying the function to each perturbed 3D layout, The training is performed using conditional dropout with respect to at least one layout parameter, wherein the at least one layout parameter includes the semantic category of the floor plan and / or each 3D bounding box. method.

11. The aforementioned dissimilarity metric is of the following type: [Math 12] Here, [Number 13] This is a set of one or more 3D bounding boxes within the aforementioned correct 3D layout, [Number 14] This is a set of one or more 3D bounding boxes in the predicted 3D layout, N is [Number 15] and [Number 16] It is a common size, [Number 17] This is a differentiable distance. Here, the differentiable distance imposes a greater penalty when the dissimilarity between the 3D bounding boxes is greater in terms of dimensions and / or semantic categories, for example of the following types: [Number 18] Here, [Number 19] and [Number 20] These are, [Math 21] , and [Number 22] This is a vector of values ​​for the spatial bounding box parameter, [Number 23] This is the Euclidean norm, [Number 24] This is a penalty parameter, [Number 25] teeth, [Number 26] and [Number 27] This is an indicator function that is equal to 1 if the dimensions are the same, and equal to 0 otherwise. [Number 28] teeth, [Number 29] and [Number 30] This is an indicator function that is equal to 1 if they have the same semantic category, and equal to 0 otherwise. Method according to claim 10

12. The loss of the trained function is the dissimilarity metric. [Number 31] and noise-dependent weighting function [Number 32] The product of [Number 33] Expected value [Number 34] The method according to claim 10 or 11.

13. A data structure, A computer program comprising instructions for performing the method according to any one of claims 1 to 9 and / or the method according to any one of claims 10 to 12, and / or A machine learning function trained by the method described in any one of claims 10 to 12, Data structure.

14. A computer-readable storage medium on which the data structure described in claim 13 is recorded.

15. A system comprising a processor coupled to memory, wherein the memory records the data structure described in claim 13.