Scene generation method and device and storage medium

By combining the planning guidance function with the diffusion model, the problem of lack of physical constraints in scene generation in existing technologies is solved, and the physical rationality and interactivity of the generated scenes are achieved while maintaining authenticity and naturalness, meeting the needs of large-scale physical interactive scene generation.

CN120671484APending Publication Date: 2025-09-19BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410309028.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing scene generation technologies lack consideration of physical constraints, which makes the generated scenes easily violate physical constraints, and the quality assessment indicators fail to effectively evaluate the physical rationality and interactivity of the generated scenes, affecting the correctness of the simulation environment.

Method used

Combining the planning guidance function with the diffusion model, the scene probability distribution is generated through collision avoidance, scene layout constraints and reachability constraints, ensuring the physical rationality and interactivity of the generated scene while maintaining authenticity and naturalness.

Benefits of technology

Effectively generate physically interactive scenes to meet the needs of large-scale physical interactive scene generation, ensure the physical rationality and interactivity of the generated scenes, while maintaining authenticity and naturalness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671484A_ABST
    Figure CN120671484A_ABST
Patent Text Reader

Abstract

The invention provides a scene generation method and device and a storage medium. The method comprises the steps of obtaining a planning guide function, scene to-be-learned parameters and a diffusion model; determining scene probability distribution in a noise reduction iteration process according to the planning guide function, the scene to-be-learned parameters and the diffusion model; according to the method, the planning guide function and the diffusion model are combined, the physical rationality and interactivity of the generated scene are ensured in a simple and effective mode, meanwhile, the authenticity and naturalness of the generated scene are kept, and the scene capable of being physically interacted is effectively generated; and the large-scale physical interactive scene generation requirement can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of scene generation technology, and in particular to a scene generation method, device and storage medium. Background Art

[0002] The exploration of scene generation has been a continuous focus in the field of computer vision, initially with the goal of facilitating interior design by generating diverse 3D environments that appear realistic and natural. However, with the recent developments in EAI (Embedded Artificial Intelligence) research and the growing demand for embedded AI, the goal of scene generation has taken on new dimensions.

[0003] In related technologies, scenes are generated through diffusion models and scene generation datasets, and the quality of the generated scenes is evaluated using quality assessment indicators. Quality assessment indicators often use visual quality scores to test the performance of the model.

[0004] However, the scene generation dataset is designed with non-interactive objects in mind and lacks consideration of physical constraints. Therefore, it is easy to violate physical constraints, which poses a major challenge to the algorithm learning the physically reasonable arrangement of interactive objects. The quality assessment indicators only evaluate the naturalness and realism of the generated scenes, and do not consider the physical rationality and interactivity of the generated scenes, which affects the correctness of the generated scenes when applied to simulation environments. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art.

[0006] To this end, one purpose of the present invention is to propose a scene generation method, which combines a planning guidance function with a diffusion model to ensure the physical rationality and interactivity of the generated scene in a simple and effective way, while maintaining the authenticity and naturalness of the generated scene, effectively generating a physically interactive scene, and being able to meet the needs of large-scale physical interactive scene generation.

[0007] To this end, a second object of the present invention is to provide a scene generating device.

[0008] To this end, a third object of the present invention is to provide a non-transitory computer-readable storage medium.

[0009] In order to achieve the above-mentioned purpose, an embodiment of the first aspect of the present invention proposes a scene generation method, which includes: obtaining a planning guidance function, scene parameters to be learned and a diffusion model; determining the scene probability distribution of the denoising iterative process based on the planning guidance function, the scene parameters to be learned and the diffusion model; and generating a predicted scene based on the scene probability distribution.

[0010] According to the scene generation method of an embodiment of the present invention, by obtaining a planning guidance function, scene parameters to be learned and a diffusion model, physical constraints such as collision avoidance, scene layout constraints and reachability constraints are cleverly converted into a diffusion model, so that physical rationality and interactivity are used as conditions for scene generation, and the scene layout distribution is generated using the diffusion model and the scene parameters to be learned. The planning guidance function is applied to improve the physical rationality and interactivity of the generated scene to determine the scene probability distribution of the denoising iterative process, and the generated scene is iterated for multiple rounds in the order of adjacent time indexes. A predicted scene is generated based on the scene probability distribution after multiple rounds of iterations. By combining the planning guidance function with the diffusion model, the physical rationality and interactivity of the generated scene are ensured in a simple and effective way, while maintaining the authenticity and naturalness of the generated scene, and effectively generating a physically interactive scene, which can meet the needs of large-scale physical interactive scene generation.

[0011] In some embodiments, the scene probability distribution of the noise reduction iterative process is determined based on the planning guidance function, the scene parameters to be learned and the diffusion model, including: adjusting the scene parameters to be learned according to the planning guidance function to obtain scene constraint parameters to be learned; and determining the scene probability distribution of the noise reduction iterative process based on the scene constraint parameters to be learned and the diffusion model.

[0012] In some embodiments, determining the scene probability distribution of the noise reduction iterative process based on the scene constraint parameters to be learned and the diffusion model includes: inputting the scene constraint parameters to be learned into the diffusion model to determine the scene probability distribution of the noise reduction iterative process.

[0013] In some embodiments, the scene parameters to be learned are adjusted according to the planning guidance function to obtain the scene constraint parameters to be learned, including: determining the scene optimization index according to the planning guidance function; adjusting the scene parameters to be learned according to the scene optimization index to determine the optimal conditions; and determining the scene constraint parameters to be learned according to the optimal conditions.

[0014] In some embodiments, generating a predicted scene based on the scene probability distribution includes: calculating the product of the scene probability distribution under adjacent time indexes; determining the probability of the target scene distribution based on the product; and generating the predicted scene based on the probability of the target scene distribution.

[0015] In some embodiments, obtaining scene parameters to be learned includes: obtaining a noise target function; and determining the scene parameters to be learned based on the noise target function, wherein the scene parameters to be learned include: a mean learning parameter and a variance learning parameter.

[0016] In some embodiments, the noise objective function is:

[0017]

[0018] in, is a predefined function, x t is the parameter of scene distribution during the process of adding or denoising. is a plane graph, and x0 is the initial scene state parameter.

[0019] In some embodiments, the noise objective function is:

[0020]

[0021] Among them, O is the scene optimization index, Σ and λ are scaling factors.

[0022] In some embodiments, the scene probability distribution of the denoising iterative process is:

[0023]

[0024]

[0025] Among them, O is the scene optimization index, Σ and λ are scale factors, and the It is the planning guidance function.

[0026] In some embodiments, when adjusting the scenario parameters to be learned according to the scenario optimization index to determine the optimal condition, the optimal condition is:

[0027]

[0028]

[0029] in, g is x t =μ in The first-order gradient estimate at , C is a constant.

[0030] In order to achieve the above-mentioned purpose, an embodiment of the second aspect of the present invention proposes a scene generation device, which includes: an acquisition module for acquiring a planning guidance function, scene parameters to be learned and a diffusion model; a determination module for determining the scene probability distribution of the denoising iterative process based on the planning guidance function, the scene parameters to be learned and the diffusion model; and a generation module for generating a predicted scene based on the scene probability distribution.

[0031] According to the scene generation device of the embodiment of the present invention, by obtaining the planning guidance function, the scene parameters to be learned and the diffusion model, physical constraints such as collision avoidance, scene layout constraints and reachability constraints are cleverly converted into the diffusion model, so that physical rationality and interactivity are used as conditions for scene generation, the scene layout distribution is generated using the diffusion model and the scene parameters to be learned, and the planning guidance function is applied to improve the physical rationality and interactivity of the generated scene to determine the scene probability distribution of the noise reduction iterative process, and the generated scene is iterated for multiple rounds in the order of adjacent time indexes. The predicted scene is generated according to the scene probability distribution after multiple rounds of iterations. By combining the planning guidance function with the diffusion model, the physical rationality and interactivity of the generated scene are ensured in a simple and effective way, while maintaining the authenticity and naturalness of the generated scene, and effectively generating a physically interactive scene, which can meet the needs of large-scale physical interactive scene generation.

[0032] In order to achieve the above-mentioned purpose, an embodiment of the third aspect of the present invention proposes a non-temporary computer-readable storage medium, on which a scene generation program is stored. When the scene generation program is executed by a processor, the scene generation method described in the above-mentioned embodiment is implemented.

[0033] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0035] Figure 1 is a flow chart of a scene generation method according to one embodiment of the present invention;

[0036] Figure 2 is a flow chart of a scene generation method according to a specific embodiment of the present invention;

[0037] Figure 3 is a structural block diagram of a scene generation device according to an embodiment of the present invention.

[0038] Reference numerals:

[0039] Scene generating device 2;

[0040] Acquisition module 21; determination module 22; generation module 23. DETAILED DESCRIPTION

[0041] The embodiments of the present invention will be described in detail below. The embodiments described with reference to the accompanying drawings are exemplary. The embodiments of the present invention will be described in detail below.

[0042] Simulated environments now support a wide range of complex, concrete tasks, making scenario generation a crucial data source. This provides embedded AI with an unlimited supply of scenarios, enabling them to robustly learn skills such as navigation and manipulation. This trend highlights the growing importance of scenario generation in the context of embedded AI research.

[0043] Indoor scene generation is defined as a layout prediction problem, where each object in the indoor scene is usually represented by its 3D bounding box, semantic label, or shape feature, so that the corresponding model can be retrieved from the 3D asset library and placed at a specific location.

[0044] To correctly model the layout of objects in the training dataset, current methods such as ATISS and DiffuScene typically represent the arrangement of objects as a scene graph and use scene priors, such as the spatial relationship between objects and the frequency of co-occurrence of categories, to approximate the scene layout distribution. When generating new scenes, these methods use iterative sampling or optimization strategies to avoid scenes that violate the designed scene priors, thereby improving generation efficiency.

[0045] Diffusion models have demonstrated promising results in generative AI across various fields. Through an iterative denoising process, diffusion models are able to handle high-dimensional distributions without mode collapse. This iterative process also provides flexible ways to condition and guide the model, effectively influencing its reasoning. For example, using physically constrained objectives as conditional guidance for physically feasible planning and motion generation, or implementing a physics-based motion projection module to incorporate physical laws into the denoising diffusion process for motion generation, can be more efficient during inference. Compared to constrained sampling methods such as Markov Chain Monte Carlo, diffusion guidance is more efficient during inference.

[0046] When evaluating the quality of generated scenes, common metrics use visual quality scores to test model performance. However, these realism metrics fail to consider the physical plausibility and interactivity of the generated scenes, which are crucial for applying the scenes in simulation environments. In fact, a commonly used scene generation dataset, such as the 3D-FRONT dataset, often exhibits these physically unreasonable layouts. Furthermore, existing technologies have not conducted in-depth research on the interactivity of scenes and the interactivity of objects. For example, a procedural generation process for interactive scenes based on rule constraints and statistical scene priors has been used to eliminate these interactions. However, these generated scenes are limited by predefined priors, resulting in unrealistic scenes that are detrimental to embedded artificial intelligence learning.

[0047] Achieving a seamless transition from traditional scene generation algorithms to algorithms tailored for specific scenarios presents significant challenges in scene generation. Because many scene generation tasks involve physics simulation, generated scenes must adhere to physical constraints while achieving a high degree of interactivity between objects (such as interactive objects or fluids) and scene layout (such as object reachability) to enable embedded AI to learn skills such as navigation and manipulation. These stringent interactivity requirements introduce multiple obstacles to scene generation algorithms.

[0048] First, existing scene generation datasets are designed with non-interactive objects in mind, and existing scene generation models lack consideration of physical constraints and are not necessarily trained on data that conforms to physical constraints. This makes the generated scenes prone to violations of physical constraints, posing a significant challenge for algorithms learning physically reasonable arrangements of interactive objects. In addition to data-level obstacles, introducing scene interactivity (such as maintaining sufficient workspace and ensuring object accessibility and interactivity) also poses very difficult challenges when designing optimizable objectives that reflect these abstract concepts. These challenges highlight the need for an effective scene generation algorithm that can ensure the physical rationality and interactivity of the scene while maintaining the naturalness and realism of traditional generation algorithms.

[0049] The following combination Figure 1-Figure 2 The scene generation method according to the embodiment of the present invention is described with examples.

[0050] like Figure 1 As shown, the scene generation method of the embodiment of the present invention at least includes steps S1 to S3.

[0051] Step S1: Obtain a planning guidance function, scene parameters to be learned, and a diffusion model.

[0052] Among them, the planning guidance function is a function that improves the physical rationality and interactivity of the generated scene, such as the collision avoidance function, the floor plan guidance function and the reachability guidance function. By obtaining the planning guidance function, the scene distribution can be guided and optimized during the prediction process. Through multiple guided optimizations, object collisions, objects exceeding the floor plan, and embedded artificial intelligence unreachable situations in the prediction results can be avoided, making the generated scene physically interactive.

[0053] In an embodiment, the diffusion model and scene parameters to be learned are obtained, for example, including the mean learning parameter μ θ and variance learning parameter Σ θ , where the scene x consists of N objects, denoted as x = {O1, ..., O N}. Each object is represented by O i ={c i , s i , ri , t i , f i}, where the semantic category c i ∈R C There are C options, size s i ∈R 3 , direction r i =(cosθ i , sinθ i )∈R 2 , position t i ∈R 3 , and the shape feature f encoded according to the shape of the object i ∈R 32 It is worth noting that common methods for scene generation only use the size s i and category c i To retrieve objects, but considering that the objects in the available interactive object dataset are very different from those in the scene generation dataset, such a method cannot be used across asset libraries. Therefore, we use shape features f i As a key indicator for object retrieval.

[0054] Specifically, a variational autoencoder is used to embed object geometric features and transform each 3D furniture model into a potential shape feature f i . In order to generate scenes with interactive objects, object assets from two datasets are considered, namely 3D-FUTURE, a static scene dataset containing only rigid objects, which contains the CAD models used in 3D-FRONT, and GAPartNet, an interactive object dataset containing various interactive objects. During the prediction process, given the rigid bodies in 3D-FRONT, we use implicit features to find the best match for the interactive objects in GAPartNet, thereby generating scenes with interactive objects. By incorporating interactive objects into the generated scenes trained only with static objects, we break through the data-level barrier between object-level interactive object datasets and static scene datasets containing only rigid objects.

[0055] In order to ensure that the generated scenes are physically reasonable and interactive, collision avoidance, scene layout constraints, and embedded artificial intelligence interactivity are considered as three key constraints and converted into guidance functions, namely the collision avoidance function, for example, denoted as The scene layout constraint function is recorded as And the reachability guidance function is denoted as

[0056] Specifically, a collision avoidance module is designed to reduce collisions between objects in the generated scene. Instead of calculating collisions between objects, the predicted bounding boxes and object centers are used as estimates to approximately calculate the object collision score. The collision scores of each pair of objects in the scene are summed and the negative value of the sum is taken to penalize object collisions to obtain a collision avoidance function, which can be recorded as Collision avoidance function The calculation formula is as follows:

[0057]

[0058] Among them, b i ={t i , r i , s i}, representing object O i The 3D bounding box of i , direction r i and size s i , IoU 3D Represents the 3D bounding box loU between object bounding boxes.

[0059] Similarly, the scene layout constraint module is designed to penalize objects located outside the given room plane, ensuring that the embedded artificial intelligence can navigate and interact within the scene. Extract the polygons that define the room boundary; then determine a set of external obstacles that identify the boundary. Representing the bounding box of the wall W with infinite thickness, the room distribution constraint is performed using the 3D bounding box loU score between the object and the wall to obtain the scene layout constraint function, such as Scene layout constraint function The calculation formula is as follows:

[0060]

[0061] Similarly, the reachability guidance module is designed to adjust the position of the object that has the greatest impact on the connectivity between different areas in the scene, ensuring that the embedded artificial intelligence can reach any area in the scene and interact with all objects. First, the generated scene is mapped to a two-dimensional room mask and the bounding box b of the embedded artificial intelligence is used to determine the room size. agent Determine the size of the embedded AI and calculate the walkable area of ​​the embedded AI in the scene. Then, use Gaussian distribution for each placed object in the scene to form a cost map for traversing the scene. Intuitively, the closer the point is to the object, the higher the cost. Using the cost map, use the A* algorithm to plan the shortest path between the two largest connected areas. The resulting path indicates the minimum cost path between the two areas. Select L objects of the embedded AI bounding box on this shortest path. To apply the guided model prediction, we can obtain the reachability guided function, for example, Reachability guidance function The calculation formula is as follows:

[0062]

[0063] Get the collision avoidance function Scene layout constraint function and reachability guidance function Finally, the three are integrated together to determine the planning guidance function, for example, as a posteriori optimization during inference.

[0064] Step S2: determining the scene probability distribution of the denoising iterative process according to the planning guidance function, the scene parameters to be learned, and the diffusion model.

[0065] In an embodiment, obtaining a planning guidance function After the scene parameters and diffusion model are learned, the diffusion model first learns the prior layout knowledge from the dataset, then uses the diffusion model and the scene parameters to be learned to generate the scene layout distribution, and applies the planning guidance function To improve the physical plausibility and interactivity of the generated scenarios, the guided prediction is formalized as a probabilistic problem of optimizing constraint satisfaction to determine the scenario probability distribution of the denoising iterative process.

[0066] Step S3: Generate a predicted scenario based on the scenario probability distribution.

[0067] In an embodiment, after determining the scene probability distribution of the noise reduction iterative process, multiple rounds of noise addition and noise reduction iterations are performed on the generated scenes in the order of adjacent time indexes, and a predicted scene is generated based on the scene probability distribution after the multiple rounds of iterations.

[0068] According to the scene generation method of an embodiment of the present invention, by obtaining a planning guidance function, scene parameters to be learned and a diffusion model, physical constraints such as collision avoidance, scene layout constraints and reachability constraints are cleverly converted into a diffusion model, so that physical rationality and interactivity are used as conditions for scene generation, and the scene layout distribution is generated using the diffusion model and the scene parameters to be learned. The planning guidance function is applied to improve the physical rationality and interactivity of the generated scene to determine the scene probability distribution of the denoising iterative process, and the generated scene is iterated for multiple rounds in the order of adjacent time indexes. A predicted scene is generated based on the scene probability distribution after multiple rounds of iterations. By combining the planning guidance function with the diffusion model, the physical rationality and interactivity of the generated scene are ensured in a simple and effective way, while maintaining the authenticity and naturalness of the generated scene, and effectively generating a physically interactive scene, which can meet the needs of large-scale physical interactive scene generation.

[0069] In some embodiments, the scene probability distribution of the noise reduction iterative process is determined based on the planning guidance function, the scene parameters to be learned and the diffusion model, including: adjusting the scene parameters to be learned based on the planning guidance function to obtain the scene constraint parameters to be learned; and determining the scene probability distribution of the noise reduction iterative process based on the scene constraint parameters to be learned and the diffusion model.

[0070] In an embodiment, obtaining a planning guidance function After learning the parameters and diffusion model of the scene, the planning guidance function Adjust time t and check whether the result generated by the conditions at time t satisfies the planning guidance function The constraints in are used to obtain the scene constraint parameters to be learned, such as the scaling factors Σ and μ, and the scene constraint parameters to be learned are input into the diffusion model to determine the scene probability distribution of the denoising iterative process.

[0071] In some embodiments, the scene parameters to be learned are adjusted according to the planning guidance function to obtain the scene constraint parameters to be learned, including: determining the scene optimization index according to the planning guidance function; adjusting the scene parameters to be learned according to the scene optimization index to determine the optimal conditions; and determining the scene constraint parameters to be learned according to the optimal conditions.

[0072] In an embodiment, obtaining a planning guidance function Then, according to the planning guidance function Determine the scenario optimization index, for example, record it as O, and adjust the time t according to the scenario optimization index O, and check whether the result generated by the condition at time t satisfies the planning guidance function The constraints in the equation are similar to those in diffusion, and the optimal conditions are determined as follows: So according to the optimal conditions Determine the scene constraint parameters to be learned.

[0073] In some embodiments, generating a predicted scene based on the scene probability distribution includes: calculating the product of the scene probability distribution under adjacent time indexes; determining the probability of the target scene distribution based on the product; and generating a predicted scene based on the probability of the target scene distribution.

[0074] In the embodiment, the initial scene state parameter x0 is used to represent a scene distribution in the data set, Gaussian noise is gradually added to the initial scene state parameter x0, and multiple forward processes q(x t+1 |x t ) is converted into the parameter x of the scene distribution during the noise addition process t , and then use the reverse denoising process p θ (x t+1 |x t ) to obtain the parameters x of the scene distribution during the noise addition process with learnable parameters θ tIn the above example, we restore the initial scene state parameter x0 and use the plane map As a condition for scene layout constraints, in this case, the product of the scene probability distribution under adjacent time indexes is calculated, that is, in, According to the product Determine the probability of target scene distribution Probability of target scene distribution The calculation formula is as follows:

[0075]

[0076] in, Represents a given planar graph The probability that the scenario distribution is x0.

[0077] According to the probability of target scene distribution Generate prediction scenarios.

[0078] In some embodiments, obtaining the scene parameters to be learned includes: obtaining a noise target function; and determining the scene parameters to be learned according to the noise target function, wherein the scene parameters to be learned include: a mean learning parameter and a variance learning parameter.

[0079] In the embodiment, when obtaining the scene parameters to be learned, the noise target function is first obtained, for example, recorded as Therefore, according to the noise objective function Determine the scene parameters to be learned, such as the mean learning parameter μ θ and variance learning parameter Σ θ .

[0080] In some embodiments, the noise objective function is:

[0081]

[0082] in, is a predefined function, x t is the parameter of scene distribution during the process of adding or denoising. is a plane graph, and x0 is the initial scene state parameter.

[0083] In an embodiment, the probability of the target scene distribution is The maximization of can be equivalently simplified to the noise objective function Right now

[0084]

[0085] in, is a predefined function of time t in the noise forward process. In order to learn this conditional model, a UNet with an attention module is used to Modeling and adding time t and floor plan in each UNet layer As a condition.

[0086] In some embodiments, the noise objective function is:

[0087]

[0088] Among them, O is the scene optimization index, Σ and λ are scaling factors.

[0089] In an embodiment, the planning guidance function can also be Integrate it into the learning and prediction of the diffusion model, and use the planning guidance function according to the formula in the diffusion model Reconstruct the noise objective function, that is

[0090]

[0091] In scene generation, planning guidance function Jingchen needs the real scene distribution to calculate the constraints, so the planning guidance function Convert to in, It is the scene distribution predicted by the given initial scene state parameter x0, not the scene distribution x t Optimize to improve the accuracy of real scene generation.

[0092] In some embodiments, the scene probability distribution of the denoising iterative process is:

[0093]

[0094]

[0095] Among them, O is the scene optimization index, Σ and λ are scaling factors, It is the planning guidance function.

[0096] In an embodiment, to generate a scene that takes the constraints into account, we can modify the denoising process:

[0097]

[0098] in, and λ are scaling factors. It is worth noting that these formulas utilize the predefined planning guidance function as a tilt function on the original scene distribution with constraints.

[0099] In some embodiments, when adjusting the scene parameters to be learned based on the scene optimization index to determine the optimal condition, the optimal condition is:

[0100]

[0101]

[0102] in, g is x t =μ in The first-order gradient estimate at , C is a constant.

[0103] In the embodiment, after determining the scene optimization index O, the parameter x of the scene distribution in the denoising process of the output generated by the condition in the denoising time t is checked. t Whether the planning guidance function is satisfied The constraint in , similar to diffusion, is around x at time t t =μ first-order Taylor expansion to estimate the optimal condition, that is

[0104]

[0105]

[0106] Reference below Figure 2 The scene generation method according to the embodiment of the present invention is described with examples.

[0107] like Figure 2 As shown, the scene generation method of the embodiment of the present invention at least includes steps S11 to S21.

[0108] Step S11: Obtain a planning guidance function and a diffusion model.

[0109] Step S12: Obtain the noise target function.

[0110] Step S13 : determining scene parameters to be learned according to the noise target function, wherein the scene parameters to be learned include: a mean learning parameter and a variance learning parameter.

[0111] Step S14, determining a scene optimization index according to the planning function; adjusting the scene parameters to be learned according to the scene optimization index to determine the optimal conditions; and determining the scene constraint parameters to be learned according to the optimal conditions.

[0112] Step S15: Input the scene constraint parameters to be learned into the diffusion model to determine the scene probability distribution of the noise reduction iterative process.

[0113] Step S16: adjusting the scene parameters to be learned according to the scene optimization index to determine the optimal conditions.

[0114] Step S17: determining the scene constraint parameters to be learned according to the optimal conditions.

[0115] Step S18: Input the scene constraint parameters to be learned into the diffusion model to determine the scene probability distribution of the noise reduction iterative process.

[0116] Step S19: Calculate the product of the scene probability distributions under adjacent time indexes.

[0117] Step S20: determining the probability of target scene distribution according to the product.

[0118] Step S21: Generate a predicted scenario based on the probability of target scenario distribution.

[0119] According to the scene generation method of an embodiment of the present invention, by obtaining a planning guidance function, scene parameters to be learned and a diffusion model, physical constraints such as collision avoidance, scene layout constraints and reachability constraints are cleverly converted into a diffusion model, so that physical rationality and interactivity are used as conditions for scene generation, and the scene layout distribution is generated using the diffusion model and the scene parameters to be learned. The planning guidance function is applied to improve the physical rationality and interactivity of the generated scene to determine the scene probability distribution of the denoising iterative process, and the generated scene is iterated for multiple rounds in the order of adjacent time indexes. A predicted scene is generated based on the scene probability distribution after multiple rounds of iterations. By combining the planning guidance function with the diffusion model, the physical rationality and interactivity of the generated scene are ensured in a simple and effective way, while maintaining the authenticity and naturalness of the generated scene, and effectively generating a physically interactive scene, which can meet the needs of large-scale physical interactive scene generation.

[0120] The following combination Figure 3 The scene generating device 2 according to the embodiment of the present invention is described.

[0121] like Figure 3 As shown, the scene generation device 2 of the embodiment of the present invention includes: an acquisition module 21, a determination module 22 and a generation module 23, wherein:

[0122] The acquisition module 21 is used to obtain the planning guidance function, the scene parameters to be learned and the diffusion model; the determination module 22 is used to determine the scene probability distribution of the denoising iterative process based on the planning guidance function, the scene parameters to be learned and the diffusion model; the generation module 23 is used to generate a predicted scene based on the scene probability distribution.

[0123] In the embodiment, the acquisition module 21 acquires the diffusion model and the scene parameters to be learned, for example, including the mean learning parameter μ θ and variance learning parameter Σ θAmong them, the planning guidance function is a function that improves the physical rationality and interactivity of the generated scene, such as the collision avoidance function, the floor plan guidance function and the reachability guidance function. By obtaining the planning guidance function, the scene distribution can be guided and optimized during the prediction process. Through multiple guided optimizations, object collisions, objects exceeding the floor plan, and embedded artificial intelligence unreachable situations in the prediction results can be avoided, making the generated scene physically interactive.

[0124] The scene x consists of N objects, denoted as x = {O1, ..., O N}. Each object is represented by O i ={c i , s i , r i , t i , f i}, where the semantic category c i ∈R C There are C options, size s i ∈R 3 , direction r i =(cosθ i , sinθ i )∈R 2 , position t i ∈R 3 , and the shape feature f encoded according to the shape of the object i ∈R 32 It is worth noting that common methods for scene generation only use the size s i and category c i To retrieve objects, but considering that the objects in the available interactive object dataset are very different from those in the scene generation dataset, such a method cannot be used across asset libraries. Therefore, we use shape features f i As a key indicator for object retrieval.

[0125] Specifically, a variational autoencoder is used to embed object geometric features and transform each 3D furniture model into a potential shape feature f i. In order to generate scenes with interactive objects, object assets from two datasets are considered, namely 3D-FUTURE, a static scene dataset containing only rigid objects, which contains the CAD models used in 3D-FRONT, and GAPartNet, an interactive object dataset containing various interactive objects. During the prediction process, given the rigid bodies in 3D-FRONT, we use implicit features to find the best match for the interactive objects in GAPartNet, thereby generating scenes with interactive objects. By incorporating interactive objects into the generated scenes trained only with static objects, we break through the data-level barrier between object-level interactive object datasets and static scene datasets containing only rigid objects.

[0126] In order to ensure that the generated scenes are physically reasonable and interactive, collision avoidance, scene layout constraints, and embedded artificial intelligence interactivity are considered as three key constraints and converted into guidance functions, namely the collision avoidance function, for example, denoted as The scene layout constraint function is recorded as And the reachability guidance function is denoted as

[0127] Specifically, a collision avoidance module is designed to reduce collisions between objects in the generated scene. Instead of calculating collisions between objects, the predicted bounding boxes and object centers are used as estimates to approximately calculate the object collision score. The collision scores of each pair of objects in the scene are summed and the negative value of the sum is taken to penalize object collisions to obtain a collision avoidance function, which can be recorded as Collision avoidance function The calculation formula is as follows:

[0128]

[0129] Among them, b i ={t i , r i , s i}, representing object O i The 3D bounding box of i , direction r i and size s i , IoU 3D Represents the 3D bounding box loU between object bounding boxes.

[0130] Similarly, the scene layout constraint module is designed to penalize objects located outside the given room plane, ensuring that the embedded artificial intelligence can navigate and interact within the scene. Extract the polygons that define the room boundary; then determine a set of external obstacles that identify the boundary. Representing the bounding box of the wall W with infinite thickness, the room distribution constraint is performed using the 3D bounding box loU score between the object and the wall to obtain the scene layout constraint function, such as Scene layout constraint function The calculation formula is as follows:

[0131]

[0132] Similarly, the reachability guidance module is designed to adjust the position of the object that has the greatest impact on the connectivity between different areas in the scene, ensuring that the embedded artificial intelligence can reach any area in the scene and interact with all objects. First, the generated scene is mapped to a two-dimensional room mask and the bounding box b of the embedded artificial intelligence is used to determine the room size. agent Determine the size of the embedded AI and calculate the walkable area of ​​the embedded AI in the scene. Then, use Gaussian distribution for each placed object in the scene to form a cost map for traversing the scene. Intuitively, the closer the point is to the object, the higher the cost. Using the cost map, use the A* algorithm to plan the shortest path between the two largest connected areas. The resulting path indicates the minimum cost path between the two areas. Select L objects of the embedded AI bounding box on this shortest path. To apply the guided model prediction, we can obtain the reachability guided function, for example, Reachability guidance function The calculation formula is as follows:

[0133]

[0134] Get the collision avoidance function Scene layout constraint function and reachability guidance function Finally, the three are integrated together to determine the planning guidance function, for example, as a posteriori optimization during inference.

[0135] Determine module 22 to obtain planning guidance function After the scene parameters and diffusion model are learned, the diffusion model first learns the prior layout knowledge from the dataset, then uses the diffusion model and the scene parameters to be learned to generate the scene layout distribution, and applies the planning guidance function To improve the physical plausibility and interactivity of the generated scenarios, the guided prediction is formalized as a probabilistic problem of optimizing constraint satisfaction to determine the scenario probability distribution of the denoising iterative process.

[0136] After determining the scene probability distribution of the denoising iterative process, the generating module 23 performs multiple rounds of denoising and denoising iterations on the generated scenes in the order of adjacent time indexes, and generates a predicted scene based on the scene probability distribution after the multiple rounds of iterations.

[0137] According to the scene generation device 2 of the embodiment of the present invention, by obtaining the planning guidance function, the scene parameters to be learned and the diffusion model, physical constraints such as collision avoidance, scene layout constraints and reachability constraints are cleverly converted into the diffusion model, so that physical rationality and interactivity are used as conditions for scene generation, and the scene layout distribution is generated by using the diffusion model and the scene parameters to be learned, and the planning guidance function is applied to improve the physical rationality and interactivity of the generated scene to determine the scene probability distribution of the noise reduction iterative process, and the generated scene is iterated for multiple rounds in the order of adjacent time indexes, and a predicted scene is generated according to the scene probability distribution after multiple rounds of iterations. By combining the planning guidance function with the diffusion model, the physical rationality and interactivity of the generated scene are ensured in a simple and effective way, while maintaining the authenticity and naturalness of the generated scene, and effectively generating a physically interactive scene, which can meet the needs of large-scale physical interactive scene generation.

[0138] The following describes a non-transitory computer-readable storage medium according to an embodiment of the present invention.

[0139] The non-transitory computer-readable storage medium of the embodiment of the present invention stores a scene generation program, and when the scene generation program is executed by a processor, the scene generation method of the above embodiment is implemented.

[0140] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "example," "specific example," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example.

[0141] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A scene generation method, characterized in that: include: Obtain planning guidance function, scene parameters to be learned, and diffusion model; Determining a scene probability distribution of a noise reduction iterative process according to the planning guidance function, the scene parameters to be learned, and the diffusion model; A predicted scenario is generated according to the scenario probability distribution.

2. The scene generation method according to claim 1, characterized in that: Determining a scene probability distribution of a noise reduction iterative process according to the planning guidance function, the scene parameters to be learned, and the diffusion model includes: Adjust the scenario parameters to be learned according to the planning guidance function to obtain scenario constraint parameters to be learned; The scene probability distribution of the noise reduction iterative process is determined according to the scene constraint parameters to be learned and the diffusion model.

3. The scene generation method according to claim 2, characterized in that: Determining a scene probability distribution of a noise reduction iterative process according to the scene constraint parameters to be learned and the diffusion model includes: The scene constraint parameters to be learned are input into the diffusion model to determine the scene probability distribution of the noise reduction iterative process.

4. The scene generation method according to claim 2, characterized in that: Adjusting the scenario parameters to be learned according to the planning guidance function to obtain scenario constraint parameters to be learned includes: Determining a scenario optimization index according to the planning guidance function; Adjust the parameters to be learned in the scenario according to the scenario optimization index to determine the optimal conditions; The scenario constraint parameters to be learned are determined according to the optimal conditions.

5. The scene generation method according to claim 1, characterized in that: Generating a prediction scenario according to the scenario probability distribution includes: Calculate the product of the scene probability distribution under adjacent time indexes; determining a probability of a target scene distribution based on the product; The predicted scenario is generated according to the probability of the target scenario distribution.

6. The scene generation method according to claim 1, characterized in that: Get the scene parameters to be learned, including: Get the noise objective function; The scene parameters to be learned are determined according to the noise target function, wherein the scene parameters to be learned include: a mean learning parameter and a variance learning parameter.

7. The scene generation method according to claim 6, characterized in that: The noise objective function is: in, is a predefined function, x t is the parameter of scene distribution during the process of adding or denoising. is a plane graph, and x0 is the initial scene state parameter.

8. The scene generation method according to claim 6, characterized in that: The noise objective function is: Among them, O is the scene optimization index, Σ and λ are scaling factors.

9. The scene generation method according to claim 1, characterized in that: The scene probability distribution of the denoising iterative process is: Among them, O is the scene optimization index, Σ and λ are scale factors, and the It is the planning guidance function.

10. The scene generation method according to claim 4, characterized in that: When adjusting the parameters to be learned in the scenario according to the scenario optimization index to determine the optimal condition, the optimal condition is: in, g is x t =μ in The first-order gradient estimate at , C is a constant.

11. A scene generation device, characterized in that: include: The acquisition module is used to obtain the planning guidance function, the scene parameters to be learned, and the diffusion model; a determination module, configured to determine a scene probability distribution of a noise reduction iterative process according to the planning guidance function, the scene parameters to be learned, and the diffusion model; A generation module is used to generate a prediction scenario based on the scenario probability distribution.

12. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores a scene generation program, and when the scene generation program is executed by a processor, the scene generation method according to any one of claims 1 to 10 is implemented.