Scene model generation method and related apparatus
By dicing and diffusion sub-model processing of the two-dimensional scene graph, large-scale three-dimensional scenes are generated, which solves the accuracy and accuracy of three-dimensional scene generation in the prior art, and is suitable for multiple application fields.
Patent Information
- Application Number
- PCT/CN2024/114603
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-08-26
- Publication Date
- 2025-07-03
AI Technical Summary
When generating large-scale scenes, existing three-dimensional scene generation technology has problems such as object shape distortion, low accuracy and high error rate of three-dimensional scene models. It is difficult to ensure the consistency of multiple depth images, resulting in inaccurate three-dimensional scenes.
By dicing the two-dimensional scene graph, a three-dimensional scene is generated using the coding network and diffusion sub-model, the inconsistency of the overlapping parts of the sub-blocks is eliminated, and a three-dimensional scene sub-block is merged using a parameter sharing diffusion method to generate a large-scale three-dimensional scene.
It improves the accuracy and efficiency of three-dimensional scene generation, ensures the geometric accuracy and fidelity of large-scale scenes, and is suitable for open-world games, extended reality, robot training scenarios and autonomous driving.
Smart Images

Figure CN2024114603_03072025_PF_FP_ABST
Abstract
Description
A method for generating a scene model and related device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 14, 2023, application number 202311525690.8, and application name “A method for generating a scene model and related devices”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of three-dimensional technology, and in particular to a technology for generating a scene model. Background Art
[0003] Three-dimensional (3D) scene model generation is a technique that uses computer graphics and other related technologies to create realistic three-dimensional virtual environments. Currently, 3D scene model generation technology is widely used in games, film, animation, architectural design, industrial design, virtual reality, and augmented reality. In these fields, 3D scene model generation technology can help producers and designers create realistic virtual environments, improving the quality and effectiveness of their works.
[0004] Currently, 3D scene generation technology primarily relies on depth estimation models to obtain depth images (RGB-D images), and then fuses multiple depth images to create a 3D scene. However, the multiple depth images generated by this technology cannot be guaranteed to be consistent with each other, resulting in erroneous 3D scenes.
[0005] Summary of the Invention
[0006] An embodiment of the present application provides a method for generating a three-dimensional scene model and a related device. By dividing the two-dimensional scene graph of a large scene into corresponding two-dimensional scene subgraphs, the three-dimensional scene sub-blocks are merged to eliminate the inconsistency of the overlapping parts of the sub-blocks, and a three-dimensional scene of a large-scale scene corresponding to the two-dimensional scene graph of the large scene is generated, thereby ensuring the accuracy of the three-dimensional training scene.
[0007] One aspect of the present application provides a method for generating a three-dimensional scene model, the method being executed by a computer device, comprising:
[0008] Obtaining a two-dimensional scene graph, wherein the two-dimensional scene graph is used to display layout information of the three-dimensional scene to be generated;
[0009] Cut the two-dimensional scene graph into M two-dimensional scene subgraphs, wherein there is at least one overlapping area between the M two-dimensional scene subgraphs, and M is an integer greater than 1;
[0010] The M two-dimensional scene subgraphs are encoded separately through the encoding network in the three-dimensional scene generation model to generate corresponding M two-dimensional scene information;
[0011] M two-dimensional scene information are input into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional scene, wherein the diffusion sub-model is used to generate M three-dimensional scene sub-blocks according to the M two-dimensional scene information, and the M three-dimensional scene sub-blocks correspond to M three-plane feature information. The diffusion sub-model is also used to merge the M three-dimensional scene sub-blocks into a three-dimensional scene by fusing the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with overlapping areas, and the three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes.
[0012] Another aspect of the present application provides a device for generating a three-dimensional scene model, which is deployed on a computer device and includes: a two-dimensional scene graph acquisition module, a two-dimensional scene graph slicing module, a two-dimensional scene information generation module, and a three-dimensional scene generation module; specifically:
[0013] A two-dimensional scene graph acquisition module is used to obtain a two-dimensional scene graph, wherein the two-dimensional scene graph is used to display the layout information of the three-dimensional scene to be generated;
[0014] A two-dimensional scene graph slicing module, configured to slice the two-dimensional scene graph into M two-dimensional scene subgraphs, wherein there is at least one overlapping region between the M two-dimensional scene subgraphs, and M is an integer greater than 1;
[0015] A two-dimensional scene information generation module is used to encode M two-dimensional scene subgraphs respectively through the encoding network in the three-dimensional scene generation model to generate corresponding M two-dimensional scene information;
[0016] A three-dimensional scene generation module is used to input M two-dimensional scene information into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional scene, wherein the diffusion sub-model is used to generate M three-dimensional scene sub-blocks according to the M two-dimensional scene information, and the M three-dimensional scene sub-blocks correspond to M three-plane feature information. The diffusion sub-model is also used to merge the M three-dimensional scene sub-blocks into a three-dimensional scene by fusing the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with overlapping areas. The three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes.
[0017] Another aspect of the present application provides a computer device, comprising:
[0018] memories, transceivers, processors, and bus systems;
[0019] Wherein, the memory is used to store programs;
[0020] The processor is used to execute the program in the memory, including executing the above-mentioned methods;
[0021] The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.
[0022] Another aspect of the present application provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is run on a computer, the computer is enabled to execute the above-mentioned methods.
[0023] Another aspect of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, so that the computer device performs the methods provided in the above aspects.
[0024] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0025] The present application provides a method for generating a three-dimensional scene model and a related device, the method comprising: obtaining a two-dimensional scene graph, wherein the two-dimensional scene graph is used to display layout information of a three-dimensional scene to be generated; slicing the two-dimensional scene graph to obtain M two-dimensional scene subgraphs, wherein there is at least one overlapping area between the M two-dimensional scene subgraphs, and M is an integer greater than 1; encoding the M two-dimensional scene subgraphs through an encoding network in a three-dimensional scene generation model to generate M corresponding two-dimensional scene information; inputting the M two-dimensional scene information into a diffusion submodel in the three-dimensional scene generation model to generate a three-dimensional scene, wherein the diffusion submodel is used to generate M three-dimensional scene sub-blocks according to the M two-dimensional scene information, the M three-dimensional scene sub-blocks corresponding to M three-dimensional plane feature information, and merging the M three-dimensional scene sub-blocks into a three-dimensional scene by fusing the M three-plane feature information corresponding to the M two-dimensional scene subgraphs with overlapping areas, and the three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes. The method provided in the embodiment of the present application cuts the two-dimensional scene graph into blocks, inputs the two-dimensional scene information corresponding to the M two-dimensional scene sub-graphs obtained by cutting into the diffusion sub-model, converts the M two-dimensional scene information into corresponding M three-dimensional scene sub-blocks through the diffusion sub-model, merges the three-dimensional scene sub-blocks through a parameter-sharing diffusion method to generate a three-dimensional scene, and generates corresponding two-dimensional scene sub-graphs by cutting the two-dimensional scene graph of a large scene, and merges the three-dimensional scene sub-blocks to generate a three-dimensional scene of a large-scale scene corresponding to the two-dimensional scene graph of the large scene, thereby eliminating the inconsistency of the overlapping parts of the sub-blocks and ensuring the accuracy of the three-dimensional training scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG1 is a schematic diagram of the architecture of a system for generating a three-dimensional scene model according to an embodiment of the present application;
[0027] FIG2 is a flow chart of a method for generating a three-dimensional scene model according to an embodiment of the present application;
[0028] FIG3 is a schematic diagram of input and output of a generation process of a three-dimensional scene model provided in one embodiment of the present application;
[0029] FIG4 is a schematic diagram of a slicing process in generating a three-dimensional scene model according to an embodiment of the present application;
[0030] FIG5 is a flowchart of a method for generating a three-dimensional scene model provided in another embodiment of the present application;
[0031] FIG6 is a flowchart of a method for generating a three-dimensional scene model provided in another embodiment of the present application;
[0032] FIG7 is a schematic diagram of a two-dimensional target scene segmentation process provided by an embodiment of the present application;
[0033] FIG8 is a flowchart of a method for generating a three-dimensional scene model provided in another embodiment of the present application;
[0034] FIG9 is a flowchart of a method for generating a three-dimensional scene model provided in another embodiment of the present application;
[0035] FIG10 is a schematic diagram of weighted processing of coincident vertices provided by an embodiment of the present application;
[0036] FIG11 is a flowchart of a method for generating a three-dimensional scene model provided in another embodiment of the present application;
[0037] FIG12 is a schematic diagram of denoising using a diffusion network according to an embodiment of the present application;
[0038] FIG13 is a flowchart of a method for generating a three-dimensional scene model according to another embodiment of the present application;
[0039] FIG14 is a schematic diagram of slicing a three-dimensional training scene according to an embodiment of the present application;
[0040] FIG15 is a schematic diagram of a process of converting three-plane feature information into an SDF according to an embodiment of the present application;
[0041] FIG16 is a schematic diagram of the structure of a device for generating a three-dimensional scene model according to an embodiment of the present application;
[0042] FIG17 is a schematic structural diagram of a device for generating a three-dimensional scene model according to another embodiment of the present application;
[0043] FIG18 is a schematic diagram of a server structure provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] An embodiment of the present application provides a method for generating a three-dimensional scene model, which divides a two-dimensional scene graph of a large scene into corresponding two-dimensional scene subgraphs, merges the three-dimensional scene sub-blocks, eliminates the inconsistency of overlapping sub-blocks, and generates a three-dimensional scene of a large-scale scene corresponding to the two-dimensional scene graph of the large scene, thereby ensuring the accuracy of the three-dimensional training scene and solving the problems of object shape distortion, low precision and high error rate of the three-dimensional scene model in the three-dimensional scene generation process in the prior art.
[0045] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the numbers used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0046] The embodiments of the present application can automatically generate a three-dimensional scene model through artificial intelligence (AI) technology.
[0047] Three-dimensional (3D) scene model generation is a technique that uses computer graphics and other related technologies to create realistic three-dimensional virtual environments. Currently, 3D scene model generation technology is widely used in games, film, animation, architectural design, industrial design, virtual reality, and augmented reality. In these fields, 3D scene model generation technology can help producers and designers create realistic virtual environments, improving the quality and effectiveness of their works.
[0048] The generation of 3D scene models is an important research direction in computer graphics, computer vision, and artificial intelligence. Its significance lies in its ability to present real-world objects and scenes in 3D, thus providing people with a more realistic, intuitive, and rich visual experience.
[0049] The generation of 3D scene models can be applied in many fields, such as virtual reality, augmented reality, game development, film production, architectural design, and industrial design. In virtual reality and augmented reality, the generation of 3D scene models can provide users with a more realistic immersive experience, allowing them to feel as if they are actually there. In game development and film production, the generation of 3D scene models can provide more realistic scenes and characters in games and films, improving the quality and visual experience of these games and films. In architectural design and industrial design, the generation of 3D scene models can help designers better present their design solutions, improving design efficiency and quality.
[0050] The generation of 3D scene models can also be applied in fields such as education and healthcare. In education, the generation of 3D scene models can provide students with a more intuitive and vivid teaching experience, improving their learning interest and effectiveness. In healthcare, the generation of 3D scene models can provide doctors with more accurate and intuitive diagnosis and surgical planning, improving medical standards and treatment effectiveness.
[0051] The generation of 3D scene models has important significance and application value. It can provide people with a more realistic, intuitive and rich visual experience, and can also be applied in many fields, providing strong support and assistance for the development and progress of various fields.
[0052] Currently, 3D scene generation technology has been widely used, with Text2Room being a common technique. Text2Room generates 3D scene models based on text descriptions. It first uses a pre-trained 2D diffusion model to generate a 2D image, then uses a pre-trained depth estimation model to obtain a depth image (RGB-D image). The camera position is then shifted slightly from the previous frame and the previous RGB-D image is extrapolated. By continuously shifting the camera position, multiple RGB-D images are generated, and finally, these multiple RGB-D images are fused to generate the 3D scene.
[0053] Although Text2Room technology has achieved certain results in 3D scene generation, it still has some shortcomings:
[0054] 1) Over-reliance on depth estimation models, which are not very accurate: they often lead to distortion of object shapes in 3D scenes and low precision; and due to occlusion problems, 3D scenes are often incomplete.
[0055] 2) Text2Room expands the scene by moving the camera. However, considering the problem that the field of view is blocked by the scene model when the camera moves, this technology is difficult to apply to large scene models outside a single room.
[0056] 3) The multi-frame RGB-D images generated by Text2Room cannot be guaranteed to be consistent with each other, which often leads to incorrect 3D scenes.
[0057] The existing three-dimensional scene generation methods have the disadvantages of low geometric accuracy and difficulty in expanding to large-scale scenes. To this end, the present application embodiment proposes a method for generating a three-dimensional scene model that can be used to generate large-scale three-dimensional scenes. Specifically:
[0058] First, a two-dimensional scene graph is obtained, wherein the two-dimensional scene graph is used to display the layout information (configuration information) of the three-dimensional scene to be generated; then, the two-dimensional scene graph is cut into blocks to obtain M two-dimensional scene sub-graphs, wherein there is at least one overlapping area in the M two-dimensional scene sub-graphs, and M is an integer greater than 1; then, the M two-dimensional scene sub-graphs are respectively encoded through the encoding network in the three-dimensional scene generation model to generate M two-dimensional scene information; finally, the M two-dimensional scene information is input into the diffusion sub-model in the three-dimensional scene generation model, and the M two-dimensional scene information is processed by the diffusion sub-model to generate a three-dimensional scene, wherein the diffusion sub-model is used to generate M three-dimensional scene sub-blocks according to the M two-dimensional scene information, the M three-dimensional scene sub-blocks correspond to M three-plane feature information, and the M three-dimensional scene sub-blocks are merged into a three-dimensional scene by performing parameter-sharing diffusion processing on the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with overlapping areas, and the three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes.
[0059] The method provided in the embodiment of the present application cuts the two-dimensional scene graph into blocks, inputs the two-dimensional scene information corresponding to the M two-dimensional scene sub-graphs obtained by cutting into the diffusion sub-model, converts the M two-dimensional scene information into corresponding M three-dimensional scene sub-blocks through the diffusion sub-model, merges the three-dimensional scene sub-blocks through a parameter-sharing diffusion method to generate a three-dimensional scene, and generates corresponding two-dimensional scene sub-graphs by cutting the two-dimensional scene graph of a large scene, and merges the three-dimensional scene sub-blocks to generate a three-dimensional scene of a large-scale scene corresponding to the two-dimensional scene graph of the large scene, thereby eliminating the inconsistency of the overlapping parts of the sub-blocks and ensuring the accuracy of the three-dimensional training scene.
[0060] The application scenarios of the method provided in the embodiments of the present application include: scene generation of large-scale three-dimensional maps in open-world games, scene generation of extended reality (XR), scene generation of robot training, and scene generation in the field of autonomous driving.
[0061] For ease of understanding, please refer to Figure 1, which is an application environment diagram of the method for generating a three-dimensional scene model in an embodiment of the present application. As shown in Figure 1, the method for generating a three-dimensional scene model in an embodiment of the present application is applied to a system for generating a three-dimensional scene model. The system for generating a three-dimensional scene model includes: a server and a terminal device; wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and the embodiment of the present application does not limit this.
[0062] The server first obtains a two-dimensional scene graph, where the two-dimensional scene graph is used to display the layout information of the three-dimensional scene to be generated; then, the server cuts the two-dimensional scene graph into blocks to obtain M two-dimensional scene sub-graphs, where there is at least one overlapping area in the M two-dimensional scene sub-graphs, and M is an integer greater than 1; then, the server encodes the M two-dimensional scene sub-graphs separately through the encoding network in the three-dimensional scene generation model to generate M two-dimensional scene information; finally, the server inputs the M two-dimensional scene information into the diffusion sub-model in the three-dimensional scene generation model, processes the M two-dimensional scene information through the diffusion sub-model, and generates a three-dimensional scene, where the diffusion sub-model is used to generate M three-dimensional scene sub-blocks based on the M two-dimensional scene information, the M three-dimensional scene sub-blocks corresponding to M three-plane feature information, and performs parameter-sharing diffusion processing on the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with overlapping areas to merge the M three-dimensional scene sub-blocks into a three-dimensional scene, and the three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes.
[0063] Generating 3D scene models has widespread application in the gaming industry. The methods provided in the embodiments of this application allow for the generation of large-scale, high-quality, and diverse 3D scenes based on user input. For example, they can be used to automatically generate large-scale 3D maps in open-world games.
[0064] The following will introduce the method for generating a 3D scene model in this application from the perspective of the server. Please refer to Figure 2 for the generation of the 3D scene model. The method for generating a 3D scene model provided in the embodiment of this application includes: S110 to S140. Specifically:
[0065] S110: Obtain a two-dimensional scene graph.
[0066] The two-dimensional scene graph is used to display the layout information of the three-dimensional scene to be generated.
[0067] It can be understood that the two-dimensional scene graph is a layout diagram of a pre-generated three-dimensional scene. Preferably, it can be a top view of the three-dimensional scene. For example, as shown in Figure 3, Figure 3 is a schematic diagram of the generation of a three-dimensional scene model provided in an embodiment of the present application. (a) in Figure 3 is a two-dimensional scene graph input by the user, and (b) in Figure 3 is a three-dimensional scene generated based on the two-dimensional scene graph input by the user. The user can draw a two-dimensional scene graph of the scene through a simple interactive interface, and obtain the corresponding three-dimensional scene by inputting the two-dimensional scene graph.
[0068] S120 , cutting the two-dimensional scene graph into blocks to obtain M two-dimensional scene subgraphs.
[0069] There is at least one overlapping region between the M two-dimensional scene subgraphs, where M is an integer greater than 1. The overlapping region may refer to an overlapping region formed between at least two of the M two-dimensional scene subgraphs. For example, overlapping regions formed between any two or more two-dimensional scene subgraphs belong to the at least one overlapping region.
[0070] It's easy to understand that by slicing the input 2D scene graph, the target scene graph of a large scene can be divided into multiple smaller subgraphs for subsequent processing and 3D scene generation. This slicing method can better handle the complex layout and details of large scenes, improving the efficiency and accuracy of 3D scene generation.
[0071] In the step of slicing the two-dimensional scene graph, the input two-dimensional scene graph is divided into M two-dimensional scene subgraphs. These subgraphs may have overlapping areas, that is, they jointly cover part of the original two-dimensional scene graph. For ease of understanding, please refer to Figure 4, which is a schematic diagram of slicing the two-dimensional scene graph. (a) in Figure 4 is a schematic diagram of slicing the two-dimensional scene graph, and (b) in Figure 4 is a schematic diagram of the corresponding three-dimensional scene generated based on the two-dimensional scene subgraphs. It can be seen from Figure 4 that when slicing the two-dimensional scene graph, the sliced subgraphs will include overlapping parts, and the corresponding three-dimensional scene sub-blocks generated accordingly will be spliced with the three-dimensional scene sub-blocks with overlapping areas including the overlapping parts between the subgraphs to generate a three-dimensional scene.
[0072] S130 , respectively encode the M two-dimensional scene subgraphs through the encoding network in the three-dimensional scene generation model to generate corresponding M two-dimensional scene information.
[0073] It can be understood that these two-dimensional scene subgraphs are encoded through the encoding network to generate M corresponding two-dimensional scene information. Each two-dimensional scene information is used to represent the pixel features of all pixels in the two-dimensional scene subgraph. The encoding network is a key component of the three-dimensional scene generation model. Its function is to convert the two-dimensional scene subgraphs into information that can be subsequently processed and used to generate a three-dimensional scene. In this step, the encoding network processes each two-dimensional scene subgraph, extracts its features and information, and converts this information into a representation that can be subsequently processed and used to generate a three-dimensional scene. By encoding the M two-dimensional scene subgraphs, the corresponding M two-dimensional scene information can be obtained, which can be used for subsequent processing and generation of the three-dimensional scene. This information may include features such as objects, layout, and texture in the subgraphs, as well as information such as their position and posture in three-dimensional space. By processing and generating this two-dimensional scene information, a three-dimensional scene can be obtained. This scene is composed of the original two-dimensional scene graph and the subgraphs obtained by cutting.
[0074] The embodiment of the present application does not limit the implementation method of S130. In one possible implementation method, M two-dimensional scene information can be directly input into the encoding network, and the encoding network directly encodes the two-dimensional scene information to obtain corresponding two-dimensional scene information.
[0075] In another possible implementation, M raw pixel information corresponding to M 2D scene sub-images can be obtained first. This raw pixel information can be pixel information of the sub-images; specifically, it can be the color value, brightness value, etc. of each pixel. Then, the M 2D scene sub-images are input into an encoding network within a 3D scene generation model. The encoding network encodes the M raw pixel information corresponding to the M 2D scene sub-images to generate M 2D scene information. During this process, the encoding network extracts and encodes features from the input raw pixel information, generating feature vectors that represent the raw pixel information. These feature vectors, the result of the encoding network's feature extraction and encoding of the raw pixel information, can be used in the subsequent 3D scene generation process. The encoding network converts the raw pixel information into a form that can be understood and processed by the 3D scene generation model, facilitating subsequent 3D scene generation. Encoding the raw pixel information through the encoding network effectively reduces the dimensionality and complexity of the data, improving the efficiency and accuracy of 3D scene generation.
[0076] It’s important to note that in practical applications, the structure and parameters of the encoding network need to be designed and adjusted according to the specific task and data to achieve optimal performance and results. Furthermore, the output of the encoding network also requires subsequent processing and optimization to further improve the quality and accuracy of 3D scene generation.
[0077] S140 , inputting M two-dimensional scene information into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional scene.
[0078] The diffusion sub-model is used to generate M three-dimensional scene sub-blocks based on M two-dimensional scene information. The M three-dimensional scene sub-blocks correspond to M three-plane feature information. The diffusion sub-model is also used to fuse the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with overlapping areas to merge the M three-dimensional scene sub-blocks into a three-dimensional scene. The three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes. The three mutually perpendicular planes refer to three orthogonal planes (XY, XZ, YZ) with specific directions and uses. These three mutually perpendicular planes are basic orthogonal planes in three-dimensional space, each carrying the geometric and texture information of the 3D scene at different perspectives. Each plane corresponds to a scene representation at a perspective. Together, they constitute an efficient and scalable representation method for 3D scenes.
[0079] It should be noted that the M three-plane feature information is generated based on the M two-dimensional scene information. For example, the diffusion sub-model can generate corresponding three-plane feature information based on this two-dimensional scene information, that is, converting the two-dimensional scene information into three-plane feature information, thereby representing the two-dimensional scene information in three mutually perpendicular planes (tri-plane). The three-dimensional plane feature information can be obtained before or after the three-dimensional scene sub-block is generated, or it can be generated simultaneously with the three-dimensional scene sub-block, and this embodiment of the application is not limited to this.
[0080] It is understood that the M encoded 2D scene information is input into the diffusion sub-model within the 3D scene generation model, where it is processed to generate a 3D scene. The diffusion sub-model is a key component of the 3D scene generation model. Its function is to generate corresponding 3D scene sub-blocks based on the input 2D scene information and merge these sub-blocks into a single 3D scene. In this step, the diffusion sub-model processes each 2D scene information to generate a corresponding 3D scene sub-block. These 3D scene sub-blocks are generated from the original 2D scene information and include information such as their position and posture in 3D space. The diffusion sub-model then generates corresponding three-dimensional feature information based on this 2D scene information. This information represents the 2D information of the 2D scene information in three mutually perpendicular planes. This three-dimensional feature information can include features such as the layout and texture of the sub-blocks, as well as their position and posture in 3D space. By fusing this three-dimensional feature information, the three-dimensional feature information corresponding to sub-blocks with overlapping areas can be merged, thereby merging these sub-blocks into a single 3D scene. By processing the diffusion sub-model, a 3D scene can be generated, which is composed of the original 2D scene graph and the sub-graphs obtained by slicing. This method can effectively handle the complex layout and details of large scenes, improving the efficiency and accuracy of 3D scene generation.
[0081] It is understandable that, since there are overlapping areas in the M 3D scene sub-blocks, during the fusion process of the three-plane feature information, the overlapping areas can be subjected to parameter sharing diffusion processing. Parameter sharing means that the M 3D scene sub-blocks share the parameters of the overlapping areas, thereby avoiding repeated calculation or storage of the same information in different 3D scene sub-blocks, thereby saving computing resources and storage. Diffusion processing is a process of gradually propagating information. In this case, it can refer to propagating the parameters of the overlapping areas from one 3D scene sub-block to another 3D scene sub-block, thereby achieving parameter sharing and ensuring information continuity and consistency throughout the entire large-scale scene.
[0082] From the above introduction, it can be seen that the three-dimensional scene generation model provided in the embodiment of the present application is a model for generating a three-dimensional scene based on a two-dimensional scene graph, including a coding network and a diffusion sub-model. The coding network is responsible for extracting key information from the input two-dimensional scene graph and converting it into a form suitable for subsequent processing. The coding network may include multiple convolutional layers for local features and spatial relationships in the two-dimensional scene graph, such as features such as objects, layouts, textures, etc., as well as information such as their position and posture in three-dimensional space.
[0083] In a preferred embodiment, denoising diffusion probabilistic models (DDPMs) can be used as the diffusion submodel in the 3D scene generation model. DDPMs are generative models based on a diffusion process that generate new samples by learning the underlying distribution of data. The core concept of DDPMs is to view the data generation process as a diffusion process, that is, a process that gradually evolves from an initial state to a final state. By learning the probability density function of the diffusion process, DDPMs can generate samples of the final state given an initial state. DDPM training consists of two phases: forward diffusion and backward diffusion. In the forward diffusion phase, DDPMs start from a random initial state and gradually evolve to the final state through a series of diffusion steps. In each diffusion step, DDPMs add some noise to the current state to simulate the randomness of the diffusion process. In the backward diffusion phase, DDPMs start from the final state and gradually restore it to the initial state through a series of denoising steps. In each denoising step, DDPMs attempt to remove the noise in the current state to restore the original sample. The generation process of DDPM is to gradually evolve to the final state through the forward diffusion process given an initial state, and then gradually restore to the initial state through the reverse diffusion process, and finally obtain the generated sample.
[0084] It is understandable that in order to ensure the accuracy of the 3D generated scene, it is necessary to smoothly transition the connections between the sub-blocks. In the embodiment of the present application, a 3D large scene is obtained by combining the diffusion generation processes of multiple small scenes.
[0085] The method provided in an embodiment of the present application seamlessly combines the geometric shapes of multiple overlapping 3D blocks by sharing the parameters of the overlapping regions of multiple diffusion generation processes. Specifically, the projected representations of the overlapping points on two three-plane feature information are determined and weighted as input to the next denoising generation iteration, ultimately resulting in a generated 3D scene. In this method, the projected representations of the overlapping points on the two three-plane feature information are first determined. These projected representations can be obtained by projecting the overlapping points onto the two three-planes and calculating their coordinates on the three planes. These projected representations are then weighted to account for their relative importance in the two three-planes. This weighting can be accomplished by calculating a weight for each projected representation and summing them. The weighted projected representations are then used as input to the next denoising generation iteration. In each iteration, a new scene is generated using a diffusion generation process. This process involves denoising the scene and generating new details to improve the quality and accuracy of the scene. In each iteration, the weighted projected representations are used as input to ensure that the geometric shapes of the multiple overlapping 3D blocks can be seamlessly combined. In this way, the diffusion generation processes of multiple small scenes can be merged into a large 3D scene while ensuring the accuracy and quality of the scene.
[0086] The method for generating a three-dimensional scene model provided in an embodiment of the present application cuts a two-dimensional scene graph into blocks, inputs the two-dimensional scene information corresponding to the M two-dimensional scene sub-graphs obtained by cutting into a diffusion sub-model, converts the M two-dimensional scene information into corresponding M three-dimensional scene sub-blocks through the diffusion sub-model, merges the three-dimensional scene sub-blocks through a parameter-sharing diffusion method to generate a three-dimensional scene, cuts the two-dimensional scene graph of a large scene into corresponding two-dimensional scene sub-graphs, merges the three-dimensional scene sub-blocks to generate a three-dimensional scene of a large-scale scene corresponding to the two-dimensional scene graph of the large scene, eliminates the inconsistency of the overlapping parts of the sub-blocks, and ensures the accuracy of the three-dimensional scene.
[0087] The embodiments of the present application provide a simple and practical high-quality, large-scale three-dimensional scene generation tool. The embodiments of the present application can give rise to a large number of new and diverse XR (Extended Reality), robotics, autonomous driving and other fields, and have broad application prospects, including generating large-scale user-defined three-dimensional scene content required for the metaverse, generating training scenes required for robots / autonomous driving cars, etc. The embodiments of the present application are widely used in the gaming industry, including as a large-scale 3D scene editor for open world games, and as a game development auxiliary tool to improve the efficiency of designing and developing game 3D scene maps.
[0088] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 2 of the present application, please refer to FIG. 5 , S140 further includes S141 to S142. Specifically:
[0089] S141. Input the M two-dimensional scene information into the multilayer perceptron network in the diffusion sub-model, encode the M two-dimensional scene information through the multilayer perceptron network, and obtain M three-plane feature information.
[0090] As you can understand, the basic structure of a multilayer perceptron consists of an input layer, hidden layers, and an output layer. The input layer receives input data, the hidden layer contains multiple neurons that extract and transform features from the input data, and the output layer outputs the final classification or regression results. In a multilayer perceptron, information is transmitted between neurons through connections. Each neuron has an activation function that performs a nonlinear transformation on the weighted sum of its inputs. The choice of activation function has a significant impact on the performance and generalization ability of the multilayer perceptron. The multilayer perceptron achieves different tasks by adjusting the number of neurons and the connection structure in the hidden layer. During training, the backpropagation algorithm adjusts the connection weights between neurons to minimize the loss function or maximize classification accuracy.
[0091] A multilayer perceptron network is a common neural network structure that extracts and encodes features from input information, generating a feature vector that represents the input information. In this step, a pre-trained multilayer perceptron network processes each two-dimensional scene, extracting features and information from it and encoding them into a three-dimensional feature vector. Specifically, the multilayer perceptron network maps M two-dimensional scene information onto three mutually perpendicular planes, generating M three-dimensional feature vectors corresponding to each of the three perpendicular planes.
[0092] More specifically, a multilayer perceptron network is used to map each pixel in each two-dimensional scene image onto three mutually perpendicular planes, obtaining three-dimensional feature information corresponding to each of the three perpendicular planes. For example, if the two-dimensional scene image is (x1, y1, z1), the three-dimensional position information corresponding to each pixel in the two-dimensional scene image is first mapped onto three mutually perpendicular planes (X1, Y1, Z1), obtaining three-dimensional feature information corresponding to each of the three perpendicular planes. The plane feature information in the X1 plane is (x1, y1); the plane feature information in the Y1 plane is (y1, z1); and the plane feature information in the Z1 plane is (x1, z1). Next, the three-dimensional feature information corresponding to the three mutually perpendicular planes is integrated to obtain the original three-dimensional feature information corresponding to the two-dimensional scene image: [(x1, y1), (y1, z1), (x1, z1)]. Finally, the original three-plane feature information is integrated through the encoder to obtain the three-plane feature information set [(x2, y2), (y2, z2), (x2, z2)] corresponding to the two-dimensional scene information, where x2 = a×x1+b×y1, y2 = b×y1+c×z1, z2 = a×x1+c×z1, and a, b, and c are the relative position parameters between the plane (X1, Y1, Z1) and the plane (X2, Y2, Z2).
[0093] In another preferred embodiment, each two-dimensional scene information is input into a multilayer perceptron network, and the multilayer perceptron network predicts the signed distance function information corresponding to the two-dimensional scene information based on the two-dimensional scene information, wherein each signed distance function information set includes at least one signed distance function information, and the signed distance function information is used to characterize the three-dimensional position information of the generated pixel point. The value of the signed distance function (SDF) is the sign of the distance to the surface of the object, that is, if the point is outside the object, the value of the SDF is positive; if the point is inside the object, the value of the SDF is negative; if the point is on the surface of the object, the value of the SDF is zero. Therefore, the SDF can be used to represent the surface shape and position of the object. The SDF can be calculated in many ways, one of the common methods is to use the distance field (Distance Field) to represent it. The distance field is a three-dimensional array, in which each element represents the distance from a point in space to the surface of the object. By calculating the distance field, the SDF can be easily calculated.
[0094] S142 , inputting the M three-plane feature information into the diffusion network in the diffusion sub-model, and performing denoising processing on the three-plane feature information corresponding to the two-dimensional scene sub-images with overlapping areas through the diffusion network to generate a three-dimensional scene.
[0095] It can be understood that the M three-plane feature information is input into the diffusion network in the diffusion sub-model. The diffusion network is a deep learning model based on the diffusion model that can perform denoising and generative processing on the input information. In this step, the diffusion network performs parameter-sharing denoising on the three-plane feature information corresponding to the 2D scene sub-graphs with overlapping areas. Specifically, for each 2D scene sub-graph, the diffusion network learns a corresponding parameter vector to represent the feature information of that sub-graph. When processing sub-graphs with overlapping areas, the diffusion network shares the parameter vectors corresponding to these sub-graphs to reduce duplicate calculations and the impact of noise. Through parameter sharing, the diffusion network can better process sub-graphs with overlapping areas, improving the accuracy and quality of the generated 3D scene. During the denoising process, the diffusion network optimizes and adjusts the input three-plane feature information to reduce the impact of noise and errors. Through the processing of the diffusion network, a more accurate and clear 3D scene can be obtained. This scene is composed of the original 2D scene graph and the sub-graphs obtained by cutting. It is important to note that the parameter sharing and denoising of the diffusion network are based on the training data and model structure. In practical applications, adjustments and optimizations need to be made according to specific scenarios and data to obtain the best processing effect.
[0096] In the method for generating a three-dimensional scene model provided in an embodiment of the present application, a diffusion network performs parameter-sharing denoising on the three-plane feature information corresponding to two-dimensional scene sub-images with overlapping areas. That is, the three-plane feature information corresponding to these sub-images is denoised and optimized so that they can better represent the original two-dimensional scene information. This can effectively handle the complex layout and details of large scenes, thereby improving the efficiency and accuracy of three-dimensional scene generation.
[0097] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG5 of the present application, please refer to FIG6 , S142 further includes S1421 to S1424. Specifically:
[0098] S1421 . Group the M two-dimensional scene sub-graphs according to their overlapping areas to obtain N two-dimensional scene sub-graph groups.
[0099] Each two-dimensional scene sub-image group includes at least two two-dimensional scene sub-images with the same overlapping area, and N is an integer greater than or equal to 1.
[0100] It can be understood that two-dimensional scene subgraphs with the same overlapping area are grouped together. For example, as shown in Figure 7, the two-dimensional scene graph is cut into blocks to obtain four two-dimensional scene subgraphs A, B, C, and D, among which subgraph A and subgraph B have the same overlapping area S1, subgraph B and subgraph C have the same overlapping area S2, subgraph B and subgraph D have the same overlapping area S3, subgraph C and subgraph D have the same overlapping area S4, and subgraph B, subgraph C, and subgraph D have the same overlapping area S5. Therefore, the two-dimensional scene subgraphs A, B, C, and D are grouped according to having the same overlapping area, and five two-dimensional scene subgraph groups [(A, B), (B, C), (B, D), (C, D), (B, C, D)] are obtained.
[0101] S1422 : Group the M three-plane feature information according to the N two-dimensional scene sub-image groups to obtain N three-plane feature information groups.
[0102] Each three-plane feature information group includes three-plane feature information corresponding to at least two two-dimensional scene sub-images with the same overlapping area.
[0103] It can be understood that, according to the N two-dimensional scene sub-image groups, the N three-plane feature information groups corresponding to the M three-plane feature information are determined, that is, the three-plane feature information corresponding to the sub-images belonging to the same group are grouped into the same three-plane feature information group. For example, according to sub-step S1421, the two-dimensional scene sub-images A, B, C, and D are grouped according to having the same overlapping area to obtain 5 two-dimensional scene sub-image groups [(A, B), (B, C), (B, D), (C, D), (B, C, D)], and the three-plane feature information T corresponding to the two-dimensional scene sub-image A is determined according to the 5 two-dimensional scene sub-image groups. A , the three-plane feature information T corresponding to the two-dimensional scene sub-graph B B , the three-plane feature information T corresponding to the two-dimensional scene sub-graph C C And the three-plane feature information T corresponding to the two-dimensional scene subgraph D D Divided into 5 three-plane feature information groups [(T A ,T B )、(T B ,T C )、(T B ,T D )、(T C ,T D )、(T B ,T C ,T D )].
[0104] S1423 , inputting the N three-plane feature information groups into the diffusion network in the diffusion sub-model, and performing denoising processing on the three-plane feature information in each three-plane feature information group through the diffusion network to obtain N three-plane denoised feature groups.
[0105] Among them, N three-plane feature information groups correspond to N diffusion parameters.
[0106] It can be understood that, by adopting the shared parameter method, N three-plane feature information groups are denoised through the diffusion network to obtain N three-plane denoised feature groups. Therefore, the N three-plane feature information groups correspond to N diffusion parameters. For example, according to the 5 three-plane feature information groups [(T A ,T B )、(T B ,T C )、(T B ,T D )、(T C ,T D )、(T B ,T C ,T D )] is diffused, the parameters for each group of diffusion processing are the same, that is, the three-plane feature information group (T A ,T B ) when diffusion treatment is performed A With T B The diffusion parameters are the same, for the three-plane feature information group (T B ,T C ) when diffusion treatment is performed B With T C The diffusion parameters are the same, for the three-plane feature information group (T B ,T D ) when diffusion treatment is performed B With T D The diffusion parameters are the same, for the three-plane feature information group (T C ,T D ) when diffusion treatment is performed C With T D The diffusion parameters are the same, for the three-plane feature information group (T B ,T C ,T D ) when diffusion treatment is performed B 、T C With T D The diffusion parameters are the same.
[0107] S1424 , performing three-dimensional merging mapping on the three-plane denoising feature information in each of the N three-plane denoising feature groups to generate a three-dimensional scene.
[0108] It is understood that the three-plane denoising feature information in each of the N three-plane denoising feature groups is 3D merged and mapped, that is, the three-plane denoising feature information in each group is fused to generate a three-dimensional scene. 3D merging and mapping refers to the process of merging and mapping multiple three-dimensional data or features to generate a new three-dimensional scene or model.
[0109] Furthermore, the 3D scene generation sub-blocks can be merged based on these three-plane denoising feature information. Specifically, a clustering algorithm can be used to cluster these 3D scene generation sub-blocks based on the similarity of the three-plane denoising feature information. For example, the K-Means algorithm can be used for clustering. First, K initial centroids are selected, and then each 3D scene generation sub-block S is clustered. i Each block is assigned to the cluster with the centroid closest to its weighted result. The centroid of each cluster is then recalculated, and the above steps are repeated until the cluster distribution no longer changes significantly. Finally, based on the clustering results, the 3D scene generation sub-blocks are merged to generate the 3D generated scene. The merging method can be selected based on the specific application scenario. For example, sub-blocks within the same cluster can be merged, or adjacent clusters can be merged.
[0110] The method for generating a 3D scene model provided in an embodiment of the present application groups 2D scene subgraphs with the same overlapping area into a group, and performs denoising and 3D merging mapping on the three-plane feature information in each group to generate a 3D scene. This reduces the impact of noise on 3D scene generation and improves the accuracy of 3D scene generation. Furthermore, the grouping process reduces computational complexity and improves processing efficiency.
[0111] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG6 of the present application, please refer to FIG8 , S1423 further includes S4231 to S4233. Specifically:
[0112] S4231. Determine coincidence points from each two-dimensional scene sub-image group.
[0113] The coincidence points are points in the coincidence area of all the two-dimensional scene sub-graphs in the two-dimensional scene sub-graph group.
[0114] It can be understood that the overlapping points are determined from the overlapping areas of all two-dimensional scene sub-images in each two-dimensional scene sub-image group. This application does not limit the number of the overlapping points. Preferably, the overlapping points can be pixel points in the overlapping area.
[0115] S4232. Determine, according to each three-plane feature information group, the first three-plane feature information and the second three-plane feature information corresponding to the coincidence point.
[0116] Among them, the first three-plane feature information is the three-plane feature information corresponding to the overlapping point in the first two-dimensional scene sub-image, the second three-plane feature information is the three-plane feature information corresponding to the overlapping point in the second two-dimensional scene sub-image, and the first two-dimensional scene sub-image and the second two-dimensional scene sub-image are any two two-dimensional scene sub-images in the two-dimensional scene sub-image group.
[0117] It can be understood that, based on the three-plane feature information group, two three-plane feature information (first three-plane feature information and second three-plane feature information) of the overlapping area of any two two-dimensional scene sub-images (first two-dimensional scene sub-image and second two-dimensional scene sub-image) in the two-dimensional scene sub-image group are determined.
[0118] Furthermore, suppose there are L coincident points, and these L coincident points are respectively recorded as P1, P2, ···, P L For each coincident point P i , find two mutually perpendicular planes A i ,B i , so that P i The projection is mapped onto these two planes. For each plane A i ,B i , the direction of the plane can be expressed by calculating its normal vector. Plane A i The normal vector is denoted as n Ai , plane B i The normal vector is denoted as n Bi For each coincident point P i , its coordinates on these two planes can be calculated (x Ai ,y Ai ), (x Bi ,y Bi These coordinates can be obtained by taking a weighted average of the coordinates of the projections of the coincident point Pi on the two planes. By obtaining the plane feature information of each coincident point in two mutually perpendicular planes, the generated coordinate information of each coincident point can be obtained, thereby providing accurate position information for subsequent merging operations.
[0119] S4233 , performing weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the coincident points in each three-plane feature information group to obtain N three-plane denoising feature groups.
[0120] It can be understood that weighted calculation is performed on the first three-plane feature information and the second three-plane feature information corresponding to each coincidence point, and the result obtained by the weighted calculation is used as the three-plane denoising feature group.
[0121] The method for generating a three-dimensional scene model provided in the embodiment of the present application obtains a more accurate three-plane denoising feature group by removing noise from the three-plane feature information, providing more reliable input data for subsequent three-dimensional generation, thereby improving the accuracy of three-dimensional generation.
[0122] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG8 of the present application, please refer to FIG9 , S4233 further includes the following S2331 to S2333. Specifically:
[0123] S2331. Obtain a first coordinate system corresponding to a first two-dimensional scene sub-image and a second coordinate system corresponding to a second two-dimensional scene sub-image in each two-dimensional scene sub-image group.
[0124] It is understood that a first coordinate system corresponding to each first two-dimensional scene sub-graph and a second coordinate system corresponding to each second two-dimensional scene sub-graph are obtained. A coordinate system is a mathematical system used to describe the position and orientation of points on a plane. Typically, a plane's normal vector and a reference point can be used to represent a plane's coordinate system.
[0125] S2332. Determine the horizontal coordinate distance difference between the first two-dimensional scene subgraph and the second two-dimensional scene subgraph, and the vertical coordinate distance difference between the first two-dimensional scene subgraph and the second two-dimensional scene subgraph according to the corresponding first coordinate system and second coordinate system in each two-dimensional scene subgraph group.
[0126] It is understood that, based on the first coordinate system corresponding to the first two-dimensional scene subgraph and the second coordinate system corresponding to the second two-dimensional scene subgraph, the distance difference between them in the horizontal and vertical directions is determined. These distance differences can be used to represent the relative positional relationship between the two planes in space.
[0127] S2333. Use the corresponding horizontal coordinate distance difference and vertical coordinate distance difference in each two-dimensional scene sub-image group as coefficients for weighted calculation, and perform weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the overlapping points in the three-plane feature information group corresponding to each two-dimensional scene sub-image group to obtain N three-plane denoising feature groups.
[0128] It can be understood that the horizontal coordinate distance difference and the vertical coordinate distance difference will be used as coefficients for weighted calculation, and the first three-plane feature information and the second three-plane feature information corresponding to the overlapping points in the three-plane feature information group corresponding to each two-dimensional scene sub-image group will be weighted calculated. Weighted calculation is a commonly used feature fusion method. By assigning different weights to different features, their impact on the final result is comprehensively considered. Using these two distance differences as weights, the first three-plane feature information and the second three-plane feature information are weighted calculated to obtain N three-plane denoising feature groups. The three-plane denoising feature information in these three-plane denoising feature groups is used in the subsequent three-dimensional generation process.
[0129] Furthermore, the first three-plane feature information includes first abscissa information and first ordinate information, and the second three-plane feature information includes second abscissa information and second ordinate information; S2333 further includes the following steps:
[0130] The horizontal coordinate distance difference corresponding to each two-dimensional scene sub-image group is used as the weighting coefficient of the first horizontal coordinate information and the second horizontal coordinate information corresponding to each two-dimensional scene sub-image group, and the vertical coordinate distance difference corresponding to each two-dimensional scene sub-image group is used as the weighting coefficient of the first vertical coordinate information and the second vertical coordinate information corresponding to each two-dimensional scene sub-image group. The first three-plane feature information and the second three-plane feature information corresponding to the overlapping points in the three-plane feature information group corresponding to each two-dimensional scene sub-image group are weightedly calculated to obtain N three-plane denoising feature groups.
[0131] For example, as shown in Figure 10, Figure 10 is a schematic diagram of weighted processing of overlapping points. For point P falling in the overlapping area, the projection features of the two three-plane feature information corresponding to it are (X1, Y1) and (X2, Y2) respectively. In order to make the geometry generated in the two blocks seamlessly connected, the weighted results of the two three-plane feature information are shared by linear weighting. The first area and the second area are two two-dimensional planes respectively, and the third area is the overlapping area of the first area and the second area. Assume that the projection features of the three-plane feature information of point P in the first area are (X1, Y1) and the projection features of the second area are (X2, Y2), then: Y1=(Y1+a×Y2) / (1+a); Y2=(Y2+a×Y1) / (1+a); X1=(X1+b×X2) / (1+b); X2=(X2+b×X1) / (1+b);
[0132] Wherein, a is the distance difference between the first area and the second area on the horizontal axis, and b is the distance difference between the first area and the second area on the vertical axis.
[0133] For each overlapping point P, the weighted result can be used as the new coordinate of the overlapping point. By weighting the three-plane feature information of the overlapping points, the three-plane feature mapping information of each overlapping point can be obtained, thereby providing more accurate position information for subsequent merging operations.
[0134] An embodiment of the present application provides a method for generating a three-dimensional scene model. By merging three-dimensional scene generation sub-blocks according to weighted results, a more accurate and complete three-dimensional generation scene can be obtained, thereby improving the quality of three-dimensional generation.
[0135] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 2 of the present application, please refer to FIG. 11 , S130 further includes S131 to S132. Specifically:
[0136] S131. Obtain M pieces of original pixel information corresponding to M two-dimensional scene sub-images.
[0137] It can be understood that the original pixel information of each two-dimensional scene sub-image in the M two-dimensional scene sub-images is obtained. Preferably, the original pixel information can be the pixel point information of the sub-image. Specifically, the original pixel information can be the color value, brightness value, etc. of each pixel point.
[0138] S132: Input the M two-dimensional scene sub-images into the encoding network in the three-dimensional scene generation model, and encode the M original pixel information corresponding to the M two-dimensional scene sub-images through the encoding network to generate corresponding M two-dimensional scene information.
[0139] It's easy to understand that the encoding network is a deep learning model that extracts and encodes features from the input 2D scene subgraphs, converting raw pixel information into a more abstract feature representation. The encoding network encodes the M raw pixel information corresponding to M 2D scene subgraphs, generating M 2D scene information. This 2D scene information, the result of the encoding network's feature extraction and encoding of the raw pixel information, can be used in the subsequent 3D scene generation process.
[0140] The present invention provides a method for generating a three-dimensional scene model. By encoding the raw pixel information of a two-dimensional scene subgraph to generate two-dimensional scene information, the method can better preserve the characteristics and details of the scene, thereby improving the accuracy of the three-dimensional scene generation. By preprocessing multiple two-dimensional scene subgraphs, the computational effort and complexity of the three-dimensional scene generation model can be reduced, thereby improving the efficiency of the three-dimensional scene generation.
[0141] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 5 of the present application, S141 further includes the following steps:
[0142] Through a multi-layer perceptron network, M two-dimensional scene information is mapped into three mutually perpendicular planes, and M three-plane feature information corresponding to the three mutually perpendicular planes are obtained.
[0143] It can be understood that the two-dimensional scene information is mapped into three mutually perpendicular planes through the Multi-Layer Perceptron (MLP) network. Specifically, the process can be divided into the following three steps:
[0144] 1) Input: Take M two-dimensional scene information as input and input it into the multi-layer perceptron network.
[0145] 2) Mapping: The input two-dimensional scene information is mapped into three mutually perpendicular planes through the hidden layers and output layers in the multi-layer perceptron network.
[0146] 3) Output: Obtain M three-plane feature information corresponding to three mutually perpendicular planes.
[0147] The purpose of this process is to convert 2D scene information into feature information in 3D space, which is then used to generate the 3D scene. By mapping the 2D scene information onto three mutually perpendicular planes, we can obtain three-dimensional feature information corresponding to each of these three planes. This feature information can be used to describe the scene's depth, height, and direction, thereby better representing the structure and characteristics of the 3D scene.
[0148] A multilayer perceptron network is an artificial neural network model that consists of multiple neurons connected to form different layers. In a multilayer perceptron network, the input layer receives input data, the hidden layer processes and transforms the input data, and the output layer outputs the processed results.
[0149] The hidden and output layers of a multilayer perceptron network play a key role in mapping two-dimensional scene information onto three perpendicular planes. The neurons in the hidden layers process and transform the input two-dimensional scene information, while the neurons in the output layer map the processed information onto three perpendicular planes.
[0150] In practical applications, the structure and parameters of a multilayer perceptron network need to be designed and adjusted based on the specific task and data. By adjusting parameters such as the number of neurons in the hidden layer, the connection structure, and the activation function, the network's processing power for input data and its output can be altered. Furthermore, by training the network, it can learn the mapping between input data and output, thereby improving its processing power and generalization capabilities for unknown data.
[0151] The embodiments of the present application provide a method for generating a three-dimensional scene model. By using a multilayer perceptron network for mapping, the complex features and relationships in the scene can be learned, thereby enhancing the scene representation capability. This can help the model better understand and generate three-dimensional scenes, improving the quality and accuracy of scene generation. Furthermore, by mapping two-dimensional scene information into three-dimensional space, the amount of computation and complexity can be reduced, computational efficiency can be improved, and the model can generate three-dimensional scenes more quickly, improving the real-time and response speed of scene generation.
[0152] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 5 of the present application, S142 further includes the following steps:
[0153] The M three-plane feature information is input into the diffusion network in the diffusion sub-model. The diffusion network performs K rounds of denoising on the three-plane feature information corresponding to the two-dimensional scene sub-images with overlapping areas to generate a three-dimensional scene.
[0154] The input of each round of denoising is the output of the previous round of denoising, and K is an integer greater than 1.
[0155] Specifically, the diffusion network is a deep learning model based on a generative adversarial network (GAN). It uses a diffusion process to generate high-quality content such as images or audio. In the diffusion sub-model, the diffusion network is used to denoise three-dimensional feature information to generate a three-dimensional scene. In the diffusion network, the input of each denoising round is the output of the previous denoising round. The diffusion network learns the distribution characteristics of the input data and performs K rounds of parameter-sharing denoising on the three-dimensional feature information corresponding to two-dimensional scene sub-images with overlapping areas, thereby generating a three-dimensional scene. In practical applications, the value of K typically needs to be adjusted based on the specific task and data. A larger value of K increases the number of denoising rounds and potentially improves the quality of the generated 3D scene, but also increases computational cost and time. Therefore, a trade-off needs to be struck between generation quality and computational efficiency.
[0156] As can be understood, referring to Figure 12, the parameters of the initial three-plane feature information are random noise. After K rounds of denoising, three-plane feature information with 3D geometric meaning is obtained, and a three-dimensional scene can be generated based on this three-plane feature information. By inputting M three-plane feature information into the diffusion network within the diffusion sub-model, the three-plane feature information corresponding to the two-dimensional scene sub-graphs with overlapping regions is subjected to K rounds of parameter-sharing denoising, thereby generating a three-dimensional scene. This process effectively utilizes the parameter sharing and denoising capabilities of the diffusion sub-model, improving the quality and efficiency of 3D scene generation.
[0157] An embodiment of the present application provides a method for generating a three-dimensional scene model. By performing K-round denoising processing on three-plane feature information with parameter sharing, three-plane feature information with 3D geometric meaning can be generated, and then a three-dimensional scene can be obtained. This can effectively reduce information loss, improve computing efficiency, and help the model better understand and generate three-dimensional scenes.
[0158] In an optional embodiment of the method for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 2 of the present application, please refer to FIG. 13 , the method for generating a three-dimensional scene model further includes S210 to S250 . Specifically:
[0159] S210: Obtain a three-dimensional training scene model.
[0160] Among them, the three-dimensional training scene model carries three-dimensional position marking information.
[0161] It is understandable that the three-dimensional training scene in the training set is obtained and the three-dimensional geometric scene dataset is used for training.
[0162] S220 , dividing the three-dimensional training scene model into blocks to obtain M three-dimensional training scene sub-blocks.
[0163] There is at least one overlapping area between the M three-dimensional training scene sub-blocks, and M is an integer greater than 1.
[0164] It is understood that before slicing the 3D training scene, the mesh corresponding to the 3D training scene needs to be converted into a watertight mesh. Preferably, the voxel remesh method in the modeling software can be used to convert the mesh corresponding to the 3D training scene into a watertight mesh. The watertight mesh has a strictly defined interior and exterior.
[0165] After watertightening the 3D training scene, the watertightened 3D training scene is sliced to obtain M 3D training scene sub-blocks. Specifically, the watertightened 3D training scene can be randomly cut into 3D training scene sub-blocks of fixed size, each 3D training scene sub-block being a cube. See Figure 14 , which shows an example of slicing for an indoor scene. Figure 14(a) shows the resulting sub-blocks, and Figure 14(b) shows the 3D training scene model for the indoor scene.
[0166] S230 , inputting the M three-dimensional training scene sub-blocks into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional training scene.
[0167] Among them, the diffusion sub-model is used to generate M three-plane training feature information based on M three-dimensional training scenes, and to merge the M three-dimensional training scene sub-blocks into a three-dimensional training scene by performing the same processing on the M three-plane training feature information corresponding to the M three-dimensional training scene sub-blocks with overlapping areas. The three-plane training feature information is used to represent the two-dimensional information of the three-dimensional training scene in three mutually perpendicular planes.
[0168] Furthermore, S230 further includes the following steps:
[0169] 1) Input M 3D training scene sub-blocks into the multi-layer perceptron network in the diffusion sub-model, encode the M 3D training scene sub-blocks through the multi-layer perceptron network, and obtain M three-plane training feature information.
[0170] It can be understood that, as shown in Figure 15, each three-dimensional training scene sub-block is input into the multi-layer perceptron network, and the multi-layer perceptron network predicts the signed distance function information corresponding to each three-dimensional training scene sub-block. The signed distance function information is used to represent the three-dimensional position generation information of the pixel point.
[0171] The value of the Signed Distance Function (SDF) is the sign of the distance to the surface of an object. That is, if the point is outside the object, the value of the SDF is positive; if the point is inside the object, the value of the SDF is negative; if the point is on the surface of the object, the value of the SDF is zero. Therefore, the SDF can be used to represent the surface shape and position of an object. The SDF can be calculated in many ways, one of the common methods is to use the distance field. The distance field is a three-dimensional array in which each element represents the distance from a point in space to the surface of an object. By calculating the distance field, the SDF can be easily calculated.
[0172] 2) The M three-plane training feature information is input into the diffusion network in the diffusion sub-model. The diffusion network performs parameter-sharing denoising on the three-plane training feature information corresponding to the 3D training scene sub-blocks with overlapping areas to generate a 3D scene.
[0173] It is understood that the M 3D training scene sub-blocks obtained by slicing are input into the encoding network of the 3D scene generation model, where they are processed to generate the 3D training scene. In a preferred embodiment, a denoising diffusion probabilistic model (DDPM) can be used as the diffusion sub-model. DDPM is a generative model based on the diffusion process that generates new samples by learning the underlying distribution of data. The core concept of DDPM is to view the data generation process as a diffusion process, that is, a process that gradually evolves from an initial state to a final state. By learning the probability density function of the diffusion process, DDPM can generate samples of the final state given an initial state. The DDPM training process is divided into two phases: forward diffusion and backward diffusion. In the forward diffusion process, DDPM starts from a random initial state and gradually evolves to the final state through a series of diffusion steps. In each diffusion step, DDPM adds some noise to the current state to simulate the randomness of the diffusion process. In the backward diffusion process, DDPM starts from the final state and gradually recovers to the initial state through a series of denoising steps. In each denoising step, DDPM attempts to remove the noise in the current state to restore the original sample. The DDPM generation process is to gradually evolve to the final state through the forward diffusion process given an initial state, and then gradually restore it to the initial state through the backward diffusion process, finally obtaining the generated sample.
[0174] It is understandable that in order to ensure the accuracy of the 3D generated scene, it is necessary to smoothly transition the connections between the sub-blocks. In the embodiment of the present application, a 3D large scene is obtained by combining the diffusion generation processes of multiple small scenes.
[0175] The method provided in an embodiment of the present application seamlessly combines the geometric shapes of multiple overlapping 3D blocks by sharing the parameters of the overlapping regions of multiple diffusion generation processes. Specifically, the projected representations of the overlapping points on two three-plane feature information are determined and weighted as input to the next iteration of denoising generation, ultimately resulting in a generated 3D scene. This method first determines the projected representations of the overlapping points on the two three-plane feature information. These projected representations can be obtained by projecting the overlapping points onto the two three-planes and calculating their coordinates on the three planes. These projected representations are then weighted to account for their relative importance in the two three-planes. This weighting can be accomplished by calculating a weight for each projected representation and summing them. The weighted projected representations are then used as input to the next iteration of denoising generation. In each iteration, a new scene is generated using the diffusion generation process. This process involves denoising the scene and generating new details to improve the quality and accuracy of the scene. In each iteration, the weighted projected representations are used as input to ensure that the geometric shapes of the multiple overlapping 3D blocks can be seamlessly combined. In this way, the diffusion generation processes of multiple small scenes can be merged into a large 3D scene while ensuring the accuracy and quality of the scene.
[0176] S240: Determine a generation loss based on the three-dimensional position mark information and the three-plane training feature information.
[0177] It can be understood that generation loss is a metric used to measure the generation performance of a 3D scene generation model. It can be obtained by calculating the generation error for each pixel. The 3D position generation information is generated by the 3D scene generation model based on the three-plane feature information of the training pixels, while the 3D position label information is the actual position information of the training pixels in the 3D training scene. Therefore, the generation loss can be calculated by calculating the generation error for each training pixel. The generation error can be obtained by calculating the distance, angle, and other information between the 3D position generation information and the 3D position label information. The smaller the generation loss, the better the generation performance of the 3D scene generation model, and vice versa. By calculating the generation loss, the 3D scene generation model can be evaluated and optimized to improve its generation performance.
[0178] S250 , adjusting the parameters of the diffusion sub-model in the three-dimensional scene generation model according to the generation loss, to generate a diffusion sub-model in a trained three-dimensional scene generation model.
[0179] It is understood that after calculating the generation loss, a backpropagation algorithm can be used to calculate the gradients of various parameters in the 3D scene generation model and adjust the parameters based on the gradients. The goal of this adjustment is to minimize the generation loss, thereby improving the generation performance of the 3D scene generation model. By continuously adjusting the parameters, the 3D scene generation model can gradually adapt to the training data, thereby improving its generation performance. Ultimately, after multiple iterations and adjustments, a trained 3D scene generation model can be obtained and used to generate new 3D scenes.
[0180] The present application provides a method for generating a three-dimensional scene model. By determining a generation loss, the generation performance of the three-dimensional scene generation model can be measured, thereby providing guidance for model optimization. By adjusting the parameters of the three-dimensional scene generation model, the generation performance of the model can be improved, making the model more adaptable to the training data, thereby improving the generation accuracy and efficiency of the three-dimensional scene generation model, and providing better support for practical applications.
[0181] The following is a detailed description of the device for generating a three-dimensional scene model in the present application, with reference to Figure 16. Figure 16 shows the device for generating a three-dimensional scene model in an embodiment of the present application, comprising: a two-dimensional scene graph acquisition module 110, a two-dimensional scene graph slicing module 120, a two-dimensional scene information generation module 130, and a three-dimensional scene generation module 140; specifically:
[0182] A two-dimensional scene graph acquisition module 110 is used to acquire a two-dimensional scene graph, wherein the two-dimensional scene graph is used to display layout information of the three-dimensional scene to be generated;
[0183] A two-dimensional scene graph slicing module 120 is configured to slice the two-dimensional scene graph into M two-dimensional scene subgraphs, wherein there is at least one overlapping region between the M two-dimensional scene subgraphs, and M is an integer greater than 1;
[0184] A two-dimensional scene information generation module 130 is configured to encode the M two-dimensional scene subgraphs respectively through an encoding network in a three-dimensional scene generation model to generate corresponding M two-dimensional scene information;
[0185] The three-dimensional scene generation module 140 is used to input M two-dimensional scene information into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional scene, wherein the diffusion sub-model is used to generate M three-plane feature information according to the M two-dimensional scene information, and to fuse the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with overlapping areas. The three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes.
[0186] The device for generating a three-dimensional scene model provided in an embodiment of the present application cuts a two-dimensional scene graph into blocks, inputs two-dimensional scene information corresponding to M two-dimensional scene sub-graphs obtained by cutting into a diffusion sub-model, converts the M two-dimensional scene information into corresponding M three-dimensional scene sub-blocks through the diffusion sub-model, merges the three-dimensional scene sub-blocks through a parameter-sharing diffusion method to generate a three-dimensional scene, cuts the two-dimensional scene graph of a large scene into corresponding two-dimensional scene sub-graphs, merges the three-dimensional scene sub-blocks to generate a three-dimensional scene of a large-scale scene corresponding to the two-dimensional scene graph of the large scene, eliminates the inconsistency of overlapping sub-blocks, and ensures the accuracy of the three-dimensional training scene.
[0187] The embodiments of the present application provide a simple and practical high-quality, large-scale three-dimensional scene generation tool. The embodiments of the present application can give rise to a large number of new and diverse XR (Extended Reality), robotics, autonomous driving and other fields, and have broad application prospects, including generating large-scale user-defined 3D scene content required for the metaverse, generating training scenes required for robots / autonomous driving cars, etc. The embodiments of the present application are widely used in the gaming industry, including as a large-scale 3D scene editor for open world games, and as a game development auxiliary tool to improve the efficiency of designing and developing game 3D scene maps.
[0188] In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, the three-dimensional scene generation module 140 is further configured to:
[0189] Input M two-dimensional scene information into the multilayer perceptron network in the diffusion sub-model, encode the M two-dimensional scene information through the multilayer perceptron network, and obtain M three-plane feature information;
[0190] The M three-plane feature information is input into the diffusion network in the diffusion sub-model. The three-plane feature information corresponding to the two-dimensional scene sub-image with overlapping areas is denoised by the diffusion network to generate a three-dimensional scene.
[0191] In the device for generating a three-dimensional scene model provided in an embodiment of the present application, a diffusion network performs parameter-sharing denoising on the three-plane feature information corresponding to two-dimensional scene sub-images with overlapping areas, i.e., denoising and optimizing the three-plane feature information corresponding to these sub-images so that they can better represent the original two-dimensional scene information, thereby effectively processing the complex layout and details of large scenes and improving the efficiency and accuracy of three-dimensional scene generation.
[0192] In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, the three-dimensional scene generation module 140 is further configured to:
[0193] Grouping the M two-dimensional scene subgraphs according to their overlapping areas to obtain N two-dimensional scene subgraph groups, wherein each two-dimensional scene subgraph group includes at least two two-dimensional scene subgraphs having the same overlapping area, and N is an integer greater than or equal to 1;
[0194] Grouping the M three-plane feature information according to the N two-dimensional scene sub-image groups to obtain N three-plane feature information groups, wherein each three-plane feature information group includes three-plane feature information corresponding to at least two two-dimensional scene sub-images with the same overlapping area;
[0195] Inputting N three-plane feature information groups into the diffusion network in the diffusion sub-model, performing denoising on the three-plane feature information in each three-plane feature information group through the diffusion network to obtain N three-plane denoised feature groups, wherein the N three-plane feature information groups correspond to N diffusion parameters;
[0196] The three-plane denoising feature information in each of the N three-plane denoising feature groups is three-dimensionally merged and mapped to generate a three-dimensional scene.
[0197] The 3D scene model generation device provided in the embodiments of the present application groups 2D scene sub-images with the same overlapping area into groups, and performs denoising and 3D merge mapping on the three-plane feature information in each group to generate a 3D scene. This reduces the impact of noise on 3D scene generation and improves the accuracy of 3D scene generation. Furthermore, the grouping process reduces computational complexity and improves processing efficiency.
[0198] In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, the three-dimensional scene generation module 140 is further configured to:
[0199] Determine coincidence points from each two-dimensional scene sub-graph group, wherein the coincidence points are points in the coincidence area of all two-dimensional scene sub-graphs in the two-dimensional scene sub-graph group;
[0200] Determining, according to each three-plane feature information group, first three-plane feature information and second three-plane feature information corresponding to the coincident point, wherein the first three-plane feature information is the three-plane feature information corresponding to the coincident point in the first two-dimensional scene sub-image, and the second three-plane feature information is the three-plane feature information corresponding to the coincident point in the second two-dimensional scene sub-image, and the first two-dimensional scene sub-image and the second two-dimensional scene sub-image are any two two-dimensional scene sub-images in the two-dimensional scene sub-image group;
[0201] A weighted calculation is performed on the first three-plane feature information and the second three-plane feature information corresponding to the coincident point in each three-plane feature information group to obtain N three-plane denoising feature groups.
[0202] The three-dimensional scene model generation device provided in the embodiment of the present application obtains a more accurate three-plane denoising feature group by removing noise from the three-plane feature information, providing more reliable input data for subsequent three-dimensional generation, thereby improving the accuracy of three-dimensional generation.
[0203] In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, the three-dimensional scene generation module 140 is further configured to:
[0204] Obtaining a first coordinate system corresponding to a first two-dimensional scene sub-image and a second coordinate system corresponding to a second two-dimensional scene sub-image in each two-dimensional scene sub-image group;
[0205] Determining, according to the first coordinate system and the second coordinate system corresponding to each two-dimensional scene sub-graph group, a distance difference between the abscissa of the first two-dimensional scene sub-graph and the second two-dimensional scene sub-graph, and a distance difference between the ordinate of the first two-dimensional scene sub-graph and the second two-dimensional scene sub-graph;
[0206] The corresponding horizontal coordinate distance difference and vertical coordinate distance difference in each two-dimensional scene sub-image group are used as coefficients for weighted calculation. The first three-plane feature information and the second three-plane feature information corresponding to the overlapping points in the three-plane feature information group corresponding to each two-dimensional scene sub-image group are weighted calculated to obtain N three-plane denoising feature groups.
[0207] An embodiment of the present application provides a device for generating a three-dimensional scene model. By merging three-dimensional scene generation sub-blocks according to weighted results, a more accurate and complete three-dimensional generation scene can be obtained, thereby improving the quality of three-dimensional generation.
[0208] In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, the first three-plane feature information includes first abscissa information and first ordinate information, and the second three-plane feature information includes second abscissa information and second ordinate information;
[0209] The three-dimensional scene generation module 140 is further configured to use the horizontal coordinate distance difference corresponding to each two-dimensional scene sub-image group as a weighting coefficient between the first horizontal coordinate information and the second horizontal coordinate information corresponding to each two-dimensional scene sub-image group, and use the vertical coordinate distance difference corresponding to each two-dimensional scene sub-image group as a weighting coefficient between the first vertical coordinate information and the second vertical coordinate information corresponding to each two-dimensional scene sub-image group, and perform weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the overlapping points in the three-plane feature information group corresponding to each two-dimensional scene sub-image group to obtain N three-plane denoising feature groups.
[0210] An embodiment of the present application provides a device for generating a three-dimensional scene model. By merging three-dimensional scene generation sub-blocks according to weighted results, a more accurate and complete three-dimensional generation scene can be obtained, thereby improving the quality of three-dimensional generation.
[0211] In an optional embodiment of the apparatus for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, the two-dimensional scene information generating module 130 is further configured to:
[0212] Get M original pixel information corresponding to M two-dimensional scene sub-images;
[0213] The M two-dimensional scene sub-images are input into the encoding network in the three-dimensional scene generation model, and the M original pixel information corresponding to the M two-dimensional scene sub-images is encoded by the encoding network to generate the corresponding M two-dimensional scene information.
[0214] The present invention provides a device for generating a three-dimensional scene model. By encoding the raw pixel information of a two-dimensional scene subgraph to generate two-dimensional scene information, the device can better preserve the scene's features and details, thereby improving the accuracy of the three-dimensional scene generation. By preprocessing multiple two-dimensional scene subgraphs, the computational effort and complexity of the three-dimensional scene generation model can be reduced, thereby improving the efficiency of the three-dimensional scene generation.
[0215] In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application,
[0216] The 3D scene generation module 140 is further configured to:
[0217] The M two-dimensional scene information is mapped to three mutually perpendicular planes through the multilayer perceptron network to obtain M three-plane feature information corresponding to the three mutually perpendicular planes. In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, the three-dimensional scene generation module 140 is further configured to:
[0218] The M three-plane feature information is input into the diffusion network in the diffusion sub-model. The three-plane feature information corresponding to the two-dimensional scene sub-image with overlapping areas is subjected to K rounds of denoising through the diffusion network to generate a three-dimensional scene. The input of each round of denoising is the output of the previous round of denoising.
[0219] An embodiment of the present application provides a device for generating a three-dimensional scene model. By performing K-round denoising processing on three-plane feature information with parameter sharing, three-plane feature information with 3D geometric meaning can be generated, and then a three-dimensional scene can be obtained. This can effectively reduce information loss, improve computing efficiency, and help the model better understand and generate three-dimensional scenes.
[0220] In an optional embodiment of the device for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 16 of the present application, referring to FIG. 17 , the device for generating a three-dimensional scene model 10 further includes: a three-dimensional training scene model acquisition module 210, a three-dimensional training scene model slicing module 220, a three-dimensional training scene generation module 230, a generation loss determination module 240, and a three-dimensional scene generation model training module 250; specifically:
[0221] A three-dimensional training scene model acquisition module 210 is used to acquire a three-dimensional training scene model, wherein the three-dimensional training scene model carries three-dimensional position mark information;
[0222] a 3D training scene model slicing module 220 configured to slice the 3D training scene model into M 3D training scene sub-blocks, wherein at least one overlapping region exists between the M 3D training scene sub-blocks, and M is an integer greater than 1;
[0223] A 3D training scene generation module 230 is configured to input M 3D training scene sub-blocks into a diffusion sub-model within a 3D scene generation model to generate a 3D training scene, wherein the diffusion sub-model is configured to generate M three-plane training feature information corresponding to the M 3D training scenes, and to merge the M three-plane training feature information corresponding to the M 3D training scene sub-blocks having overlapping areas into a 3D training scene, wherein the three-plane training feature information is configured to represent two-dimensional information of the 3D training scene in three mutually perpendicular planes.
[0224] A generation loss determination module 240 is configured to determine a generation loss based on the three-dimensional position marker information and the three-plane training feature information;
[0225] The three-dimensional scene generation model training module 250 is used to adjust the parameters of the diffusion sub-model in the three-dimensional scene generation model according to the generation loss, and generate a diffusion sub-model in the trained three-dimensional scene generation model.
[0226] In an optional embodiment of the apparatus for generating a three-dimensional scene model provided in the embodiment corresponding to FIG. 17 of the present application, the three-dimensional training scene generation module 230 is further configured to:
[0227] Input M 3D training scene sub-blocks into the multi-layer perceptron network in the diffusion sub-model, encode the M 3D training scene sub-blocks through the multi-layer perceptron network, and obtain M three-plane training feature information;
[0228] The M three-plane training feature information is input into the diffusion network in the diffusion sub-model. The three-plane training feature information corresponding to the three-dimensional training scene sub-blocks with overlapping areas is denoised by the diffusion network to generate a three-dimensional scene.
[0229] The present invention provides a device for generating a three-dimensional scene model. By determining a generation loss, the generation effect of the three-dimensional scene generation model can be measured, thereby providing guidance for model optimization. By adjusting the parameters of the three-dimensional scene generation model, the generation effect of the model can be improved, making the model more adaptable to the training data, thereby improving the generation accuracy and efficiency of the three-dimensional scene generation model, and providing better support for practical applications.
[0230] Figure 18 is a schematic diagram of a server structure provided in an embodiment of the present application. The server 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 322 (for example, one or more processors) and memory 332, and one or more storage media 330 (for example, one or more massive storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage medium 330 can be short-term storage or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 322 can be configured to communicate with the storage medium 330 to execute a series of instruction operations in the storage medium 330 on the server 300.
[0231] The server 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input and output interfaces 358, and / or one or more operating systems 341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0232] The steps executed by the server in the above embodiment may be based on the server structure shown in FIG15 .
[0233] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0234] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0235] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0236] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0237] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0238] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating a three-dimensional scene model, which is executed by a computer device and includes: Obtain a two-dimensional scene graph, where the two-dimensional scene graph is used to display the layout information of the three-dimensional scene to be generated; Slice the two-dimensional scene graph to obtain M two-dimensional scene sub-graphs, where there is at least one overlapping area between the M two-dimensional scene sub-graphs, and M is an integer greater than 1; Encode the M two-dimensional scene sub-graphs respectively through an encoding network in the three-dimensional scene generation model to generate corresponding M two-dimensional scene information; Input the M two-dimensional scene information into a diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional scene, where the diffusion sub-model is used to generate M three-dimensional scene sub-blocks according to the M two-dimensional scene information, the M three-dimensional scene sub-blocks correspond to M three-plane feature information, and the diffusion sub-model is further used to perform a fusion process on the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with the overlapping area to merge the M three-dimensional scene sub-blocks into the three-dimensional scene, and the three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in three mutually perpendicular planes.
2. The method for generating a three-dimensional scene model according to claim 1, where inputting the M two-dimensional scene information into a diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional scene includes: Input the M two-dimensional scene information into a multi-layer perceptron network in the diffusion sub-model, and encode the M two-dimensional scene information through the multi-layer perceptron network to obtain M three-plane feature information; Input the M three-plane feature information into a diffusion network in the diffusion sub-model, and perform denoising processing on the three-plane feature information corresponding to the two-dimensional scene sub-graphs with overlapping areas through the diffusion network to generate a three-dimensional scene.
3. The method for generating a three-dimensional scene model according to claim 2, where inputting the M three-plane feature information into a diffusion network in the diffusion sub-model, and performing denoising processing on the three-plane feature information corresponding to the two-dimensional scene sub-graphs with overlapping areas through the diffusion network to generate a three-dimensional scene includes: Group the M two-dimensional scene sub-graphs according to the overlapping areas of the M two-dimensional scene sub-graphs to obtain N two-dimensional scene sub-graph groups, where each two-dimensional scene sub-graph group includes at least two two-dimensional scene sub-graphs with the same overlapping area, and N is an integer greater than or equal to 1; Group the M three-plane feature information according to the N two-dimensional scene sub-graph groups to obtain N three-plane feature information groups, where each three-plane feature information group includes the three-plane feature information corresponding to at least two two-dimensional scene sub-graphs with the same overlapping area; Input the N three-plane feature information groups into a diffusion network in the diffusion sub-model, and perform denoising processing on the three-plane feature information in each three-plane feature information group through the diffusion network to obtain N three-plane denoised feature groups, where the N three-plane feature information groups correspond to N diffusion parameters; Perform three-dimensional merging mapping on the three-plane denoising feature information in each of the N three-plane denoising feature groups to generate a three-dimensional scene.
4. The method for generating a three-dimensional scene model according to claim 3, wherein inputting the N three-plane feature information groups into the diffusion network in the diffusion sub-model, and denoising the three-plane feature information in each of the three-plane feature information groups through the diffusion network to obtain N three-plane denoising feature groups, includes: Determine coincidence points from each of the two-dimensional scene sub-groups, where the coincidence points are points in the coincidence region of all the two-dimensional scene sub-graphs in the two-dimensional scene sub-group; According to each of the three-plane feature information groups, determine the first three-plane feature information and the second three-plane feature information corresponding to the coincidence points, where the first three-plane feature information is the three-plane feature information corresponding to the coincidence points in the first two-dimensional scene sub-graph, and the second three-plane feature information is the three-plane feature information corresponding to the coincidence points in the second two-dimensional scene sub-graph, and the first two-dimensional scene sub-graph and the second two-dimensional scene sub-graph are any two two-dimensional scene sub-graphs in the two-dimensional scene sub-group; Perform weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the coincidence points in each of the three-plane feature information groups to obtain N three-plane denoising feature groups.
5. The method for generating a three-dimensional scene model according to claim 4, wherein performing weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the coincidence points in each of the three-plane feature information groups to obtain N three-plane denoising feature groups, includes: Obtain the first coordinate system corresponding to the first two-dimensional scene sub-graph and the second coordinate system corresponding to the second two-dimensional scene sub-graph in each of the two-dimensional scene sub-groups; According to the corresponding first coordinate system and the second coordinate system in each of the two-dimensional scene sub-groups, determine the abscissa distance difference between the first two-dimensional scene sub-graph and the second two-dimensional scene sub-graph, and the ordinate distance difference between the first two-dimensional scene sub-graph and the second two-dimensional scene sub-graph; Use the corresponding abscissa distance difference and the ordinate distance difference in each of the two-dimensional scene sub-groups as coefficients for weighted calculation, and perform weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the coincidence points in the three-plane feature information group corresponding to each of the two-dimensional scene sub-groups to obtain N three-plane denoising feature groups.
6. The method for generating a three-dimensional scene model according to claim 5, wherein the first three-plane feature information includes first abscissa information and first ordinate information, and the second three-plane feature information includes second abscissa information and second ordinate information; Using the corresponding abscissa distance difference and the ordinate distance difference in each of the two-dimensional scene sub-groups as coefficients for weighted calculation, and performing weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the coincidence points in the three-plane feature information group corresponding to each of the two-dimensional scene sub-groups to obtain N three-plane denoising feature groups, includes: Taking the corresponding abscissa distance difference in each of the two-dimensional scene sub-groups as the weighting coefficient of the first abscissa information and the second abscissa information corresponding to each of the two-dimensional scene sub-groups, and taking the corresponding ordinate distance difference in each of the two-dimensional scene sub-groups as the weighting coefficient of the first ordinate information and the second ordinate information corresponding to each of the two-dimensional scene sub-groups, perform weighted calculation on the first three-plane feature information and the second three-plane feature information corresponding to the coincident points in the three-plane feature information group corresponding to each of the two-dimensional scene sub-groups, to obtain N three-plane denoised feature groups.
7. The method for generating a three-dimensional scene model according to any one of claims 1-6, wherein encoding the M two-dimensional scene sub-graphs respectively through an encoding network in the three-dimensional scene generation model to generate corresponding M two-dimensional scene information, includes: Obtaining M pieces of original pixel information corresponding to the M two-dimensional scene sub-graphs; Inputting the M two-dimensional scene sub-graphs into the encoding network in the three-dimensional scene generation model, and encoding the M pieces of original pixel information corresponding to the M two-dimensional scene sub-graphs through the encoding network to generate corresponding M two-dimensional scene information.
8. The method for generating a three-dimensional scene model according to any one of claims 2-7, wherein inputting the M two-dimensional scene information into a multi-layer perceptron network in the diffusion sub-model, and encoding the M two-dimensional scene information through the multi-layer perceptron network to obtain M three-plane feature information, includes: Mapping the M two-dimensional scene information to three mutually perpendicular planes through the multi-layer perceptron network to obtain M three-plane feature information corresponding to the three mutually perpendicular planes respectively.
9. The method for generating a three-dimensional scene model according to any one of claims 2-8, wherein inputting the M three-plane feature information into a diffusion network in the diffusion sub-model, and denoising the three-plane feature information corresponding to the two-dimensional scene sub-graphs with overlapping regions through the diffusion network to generate a three-dimensional scene, includes: Inputting the M three-plane feature information into the diffusion network in the diffusion sub-model, and performing K rounds of denoising processing on the three-plane feature information corresponding to the two-dimensional scene sub-graphs with overlapping regions through the diffusion network to generate a three-dimensional scene, where the input of each round of denoising processing is the output of the previous round of denoising processing, and K is an integer greater than 1.
10. The method for generating a three-dimensional scene model according to any one of claims 1-9, the method further includes: Obtaining a three-dimensional training scene model, wherein the three-dimensional training scene model carries three-dimensional position marking information; Slicing the three-dimensional training scene model to obtain M three-dimensional training scene sub-blocks, wherein there is at least one overlapping region between the M three-dimensional training scene sub-blocks, and M is an integer greater than 1; Input the M three-dimensional training scene sub-blocks into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional training scene. Among them, the diffusion sub-model is used to correspondingly generate M three-plane training feature information based on the M three-dimensional training scenes, and through fusing the M three-plane training feature information corresponding to the M three-dimensional training scene sub-blocks with overlapping regions, to merge the M three-dimensional training scene sub-blocks into the three-dimensional training scene. The three-plane training feature information is used to represent the two-dimensional information of the three-dimensional training scene in the three mutually perpendicular planes; Determine the generation loss according to the three-dimensional position marking information and the three-plane training feature information; Adjust the parameters of the diffusion sub-model in the three-dimensional scene generation model according to the generation loss, and generate the diffusion sub-model in the trained three-dimensional scene generation model.
11. The method for generating a three-dimensional scene model according to claim 10, wherein inputting the M three-dimensional training scene sub-blocks into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional training scene includes: Input the M three-dimensional training scene sub-blocks into the multi-layer perceptron network in the diffusion sub-model, and encode the M three-dimensional training scene sub-blocks through the multi-layer perceptron network to obtain M three-plane training feature information; Input the M three-plane training feature information into the diffusion network in the diffusion sub-model, and perform denoising processing on the three-plane training feature information corresponding to the three-dimensional training scene sub-blocks with overlapping regions through the diffusion network to generate a three-dimensional scene.
12. A device for generating a three-dimensional scene model, which is deployed on a computer device and includes: A two-dimensional scene graph acquisition module, configured to acquire a two-dimensional scene graph, where the two-dimensional scene graph is used to display the layout information of the three-dimensional scene to be generated; A two-dimensional scene graph slicing module, configured to slice the two-dimensional scene graph to obtain M two-dimensional scene sub-graphs, where there is at least one overlapping region among the M two-dimensional scene sub-graphs, and M is an integer greater than 1; A two-dimensional scene information generation module, configured to respectively encode the M two-dimensional scene sub-graphs through an encoding network in the three-dimensional scene generation model to generate corresponding M two-dimensional scene information; A three-dimensional scene generation module, configured to input the M two-dimensional scene information into the diffusion sub-model in the three-dimensional scene generation model to generate a three-dimensional scene. Among them, the diffusion sub-model is used to correspondingly generate M three-dimensional scene sub-blocks based on the M two-dimensional scene information. The M three-dimensional scene sub-blocks correspond to M three-plane feature information. The diffusion sub-model is further used to fuse the M three-plane feature information corresponding to the M two-dimensional scene sub-graphs with overlapping regions to merge the M three-dimensional scene sub-blocks into the three-dimensional scene. The three-plane feature information is used to represent the two-dimensional information of the two-dimensional scene information in the three mutually perpendicular planes.
13. A computer device, comprising: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to execute the program in the memory, including executing the method for generating a three-dimensional scene model according to any one of claims 1 to 11; The bus system is configured to connect the memory and the processor, so that the memory and the processor can communicate with each other.
14. A computer-readable storage medium, comprising instructions which, when running on a computer, cause the computer to execute the method for generating a three-dimensional scene model according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program which, when executed by a processor, performs the method for generating a three-dimensional scene model according to any one of claims 1 to 11.