3D scene generation method and device, electronic equipment, storage medium and product
By gradually generating 3D scene maps using multiple network models, the problem of time-consuming and inefficient generation of 3D scenes in the prior art is solved, and efficient 3D scene generation and real-time user interaction are achieved.
Patent Information
- Application Number
- CN202510335833.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the prior art, the 3D scene generation method takes a long time, is inefficient, and is difficult to realize user interaction.
Through the first network model, the second network model and the third network model, a 3D scene map is gradually generated, including determining the completion of incomplete images and depth maps, optimizing depth information, and splicing local scene maps to achieve real-time user interaction.
It improves the speed and efficiency of 3D scene generation, realizes diverse and coherent 3D scene generation, and supports real-time user interaction.
Smart Images

Figure CN120182501A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular to a 3D scene generation method, apparatus, electronic device, storage medium, and product. Background Art
[0002] In the prior art, a variety of methods for offline generating 3D scenes from a single image have been disclosed. These methods generally include generating multiple views or panoramas of the scene and subsequently converting them into a 3D representation. For example, methods such as Text2Room and LucidDreamer start with an input image and a user's text description, generate multiple scene images, and then use 3D optimization to refine the scene and create a more accurate and consistent 3D representation. Methods such as GenEx and Dreamscene360 utilize a pre-trained text-to-panorama diffusion model to synthesize coherent panoramas, which are then elevated to 3D to finally produce an exploitable 3D world. However, the 3D scene generation methods in the prior art generally have the problems of long time consumption and low efficiency. Summary of the Invention
[0003] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a 3D scene generation method, apparatus, electronic device, storage medium, and product.
[0004] According to one aspect of the embodiments of the present disclosure, there is provided a 3D scene generation method, including:
[0005] Based on a first global 3D scene graph corresponding to a first pose, determining an incomplete image and an incomplete depth map corresponding to a second pose; the second pose is determined based on processing of the first pose;
[0006] Processing the incomplete image and the received text description information through a first network model to determine a complete image;
[0007] Processing the incomplete depth map and the complete image through a second network model to determine a complete depth map;
[0008] Processing the second pose, the complete depth map, and the complete image through a third network model to determine a second global 3D scene graph; the second global 3D scene graph at least includes the first global 3D scene graph.
[0009] Optionally, the determining an incomplete image and an incomplete depth map corresponding to a second pose based on a first global 3D scene graph corresponding to a first pose includes:
[0010] Obtaining the first global 3D scene graph corresponding to the first pose;
[0011] Process the first pose based on the received rotation and translation matrix to obtain the second pose;
[0012] Process the first global 3D scene graph according to the rotation and translation matrix to determine the incomplete image and the incomplete depth map corresponding to the second pose.
[0013] Optionally, obtaining the first global 3D scene graph corresponding to the first pose includes:
[0014] Process the received first pose, preset image, and preset depth map through a third network model to determine the first global 3D scene graph; or,
[0015] Use the second global 3D scene graph as the first global 3D scene graph.
[0016] Optionally, processing the second pose, the complete depth map, and the complete image through a third network model to determine the second global 3D scene graph includes:
[0017] Extract features from the complete image through the backbone sub-network in the third network model to obtain matching features and image features;
[0018] Optimize the complete depth map based on the matching features and the second pose to determine the optimized depth map;
[0019] Determine the second local 3D scene graph based on the image features and the optimized depth map;
[0020] Determine the second global 3D scene graph based on the second local 3D scene graph and the first global 3D scene graph.
[0021] Optionally, optimizing the complete depth map based on the matching features and the second pose to determine the optimized depth map includes:
[0022] Obtain the historical matching feature closest to the second pose from the feature memory bank based on the rotation and translation matrix corresponding to the second pose; multiple historical matching features corresponding to different poses are pre-stored in the feature memory bank;
[0023] Process the historical matching feature, the matching feature, and the complete depth map through the cost volume sub-network in the third network model to determine the optimized depth map.
[0024] Optionally, processing the historical matching feature, the matching feature, and the complete depth map through the cost volume sub-network in the third network model to determine the optimized depth map includes:
[0025] Generate a plurality of depth candidate values based on the complete depth map;
[0026] Calculate the similarity between the matching feature and the historical matching feature in the depth space corresponding to the plurality of depth candidate values, and determine the probability value of each depth candidate value among the plurality of depth candidate values;
[0027] Perform weighted summation on the plurality of depth candidate values with the probability value as the weight to determine the optimized depth map.
[0028] Optionally, the determining the second local 3D scene graph based on the image feature and the optimized depth map includes:
[0029] Process the target depth value and the image feature through the segmentation sub-network in the third network model to obtain a target depth map and a Gaussian model;
[0030] Project the Gaussian model into the 3D space based on the target depth map to obtain the second local 3D scene graph.
[0031] Optionally, the determining the second global 3D scene graph based on the second local 3D scene graph and the first global 3D scene graph includes:
[0032] Project the second local 3D scene graph and the first global 3D scene graph into the pixel coordinate system;
[0033] Filter out the 3D image information that does not meet the depth consistency constraint in the second local 3D scene graph and the first global 3D scene graph in the pixel coordinate system;
[0034] Perform stitching on the 3D image information that meets the depth consistency constraint after screening in the second local 3D scene graph and the first global 3D scene graph to obtain the second global 3D scene graph.
[0035] Optionally, before determining the complete image by processing the incomplete image and the received text description information through the first network model, it further includes:
[0036] Train the first initial network model through a training data set to obtain the trained first initial network model; the training data set includes 3D images of various scenes, and each 3D image has a known scene category;
[0037] Perform compression and acceleration processing on the first initial network model through model distillation technology to obtain the first network model.
[0038] Optionally, the processing of the incomplete depth map and the complete image by the second network model to determine the complete depth map includes:
[0039] Based on the incomplete depth map, determine a binary validity mask;
[0040] Input the incomplete depth map, the binary validity mask, and the complete image into the second network model; the second network model is obtained through training.
[0041] According to another aspect of the embodiments of the present disclosure, a 3D scene generation device is provided, including:
[0042] A preprocessing module, configured to determine an incomplete image and an incomplete depth map corresponding to a second pose based on a first global 3D scene map corresponding to a first pose; the second pose is determined by processing the first pose.
[0043] A first processing module, configured to process the incomplete image and the received text description information through a first network model to determine a complete image;
[0044] A second processing module, configured to process the incomplete depth map and the complete image through a second network model to determine a complete depth map;
[0045] A third processing module, configured to process the second pose, the complete depth map, and the complete image through a third network model to determine a second global 3D scene map; at least the first global 3D scene map is included in the second global 3D scene map.
[0046] Optionally, the preprocessing module is specifically configured to obtain the first global 3D scene map corresponding to the first pose; process the first pose based on the received rotation and translation matrix to obtain the second pose; and process the first global 3D scene map according to the rotation and translation matrix to determine the incomplete image and the incomplete depth map corresponding to the second pose.
[0047] Optionally, when obtaining the first global 3D scene map corresponding to the first pose, the preprocessing module is configured to process the received first pose, a preset image, and a preset depth map through a third network model to determine the first global 3D scene map; or use the second global 3D scene map as the first global 3D scene map.
[0048] Optionally, the third processing module includes:
[0049] A feature extraction unit, configured to extract features from the complete image through a backbone sub-network in the third network model to obtain matching features and image features;
[0050] A depth optimization unit for optimizing the complete depth map based on the matching feature and the second pose to determine an optimized depth map;
[0051] A local scene generation unit for determining a second local 3D scene map based on the image feature and the optimized depth map;
[0052] An incremental fusion unit for determining the second global 3D scene map based on the second local 3D scene map and the first global 3D scene map.
[0053] Optionally, the depth optimization unit is specifically configured to obtain the historical matching feature closest to the second pose from the feature memory bank based on the rotation and translation matrix corresponding to the second pose; a plurality of historical matching features corresponding to different poses are pre-stored in the feature memory bank; the cost volume sub-network in the third network model is used to process the historical matching feature, the matching feature, and the complete depth map to determine the optimized depth map.
[0054] Optionally, when the depth optimization unit processes the historical matching feature, the matching feature, and the complete depth map through the cost volume sub-network in the third network model to determine the optimized depth map, it is configured to generate a plurality of depth candidate values based on the complete depth map; calculate the similarity between the matching feature and the historical matching feature in the depth space corresponding to the plurality of depth candidate values to determine the probability value of each depth candidate value in the plurality of depth candidate values; perform weighted summation on the plurality of depth candidate values with the probability value as the weight to determine the optimized depth map.
[0055] Optionally, the local scene generation unit is specifically configured to process the target depth value and the image feature through the segmentation sub-network in the third network model to obtain a target depth map and a Gaussian model; project the Gaussian model into the 3D space based on the target depth map to obtain the second local 3D scene map.
[0056] Optionally, the incremental fusion unit is specifically configured to project the second local 3D scene map and the first global 3D scene map into the pixel coordinate system; screen the 3D image information that does not meet the depth consistency constraint in the second local 3D scene map and the first global 3D scene map in the pixel coordinate system; perform stitching on the 3D image information that meets the depth consistency constraint after screening in the second local 3D scene map and the first global 3D scene map to obtain the second global 3D scene map.
[0057] Optionally, the device further includes:
[0058] A model training module is used to train a first initial network model with a training dataset to obtain the trained first initial network model; the training dataset includes 3D images of multiple scenarios, and each 3D image has a known scenario category; the first initial network model is compressed and accelerated through model distillation technology to obtain the first network model.
[0059] Optionally, the second processing module is specifically configured to determine a binary validity mask based on the incomplete depth map; input the incomplete depth map, the binary validity mask, and the complete image into the second network model; the second network model is obtained through training.
[0060] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, including:
[0061] A memory for storing a computer program product;
[0062] A processor for executing the computer program product stored in the memory, and when the computer program product is executed, implementing the 3D scene generation method described in any of the above embodiments.
[0063] According to still another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, implementing the 3D scene generation method described in any of the above embodiments.
[0064] According to yet another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer program instructions, and when the computer program instructions are executed by a processor, implementing the 3D scene generation method described in any of the above embodiments.
[0065] Based on the 3D scene generation method, device, electronic device, storage medium, and product provided in the above embodiments of the present disclosure, based on the first global 3D scene graph corresponding to the first pose, an incomplete image and an incomplete depth map corresponding to the second pose are determined; the second pose is determined through processing based on the first pose; the incomplete image and the received text description information are processed by the first network model to determine a complete image; the incomplete depth map and the complete image are processed by the second network model to determine a complete depth map; the second pose, the complete depth map, and the complete image are processed by the third network model to determine a second global 3D scene graph; at least the first global 3D scene graph is included in the second global 3D scene graph; in this embodiment, the first network model, the second network model, and the third network model simplify the 3D scene generation process, improve the 3D scene generation speed, and furthermore, allow interaction with the user during the 3D scene generation process, realizing diverse and coherently connected 3D scene generation.
[0066] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0067] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.
[0068] With reference to the accompanying drawings, the present disclosure can be more clearly understood according to the following detailed description, where:
[0069] Figure 1 is a schematic flowchart of a 3D scene generation method provided by an exemplary embodiment of the present disclosure;
[0070] Figure 2 is the present disclosure Figure 1 a schematic flowchart of step 102 in the illustrated embodiment;
[0071] Figure 3 is the present disclosure Figure 1 a schematic flowchart of step 108 in the illustrated embodiment;
[0072] Figure 4 is a schematic structural diagram of a third network model in a 3D scene generation method provided by an exemplary embodiment of the present disclosure;
[0073] Figure 5 is a processing schematic diagram of an alternative example of a 3D scene generation method provided by another exemplary embodiment of the present disclosure;
[0074] Figure 6 is a schematic structural diagram of a 3D scene generation device provided by an exemplary embodiment of the present disclosure;
[0075] Figure 7 illustrates a block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0076] Hereinafter, example embodiments according to the present disclosure will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the example embodiments described herein.
[0077] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0078] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0079] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0080] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, in the absence of a clear limitation or contrary indication in the context, it is generally understood to be one or more.
[0081] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after. The data referred to in the present disclosure may include unstructured data such as text, images, videos, etc., or may also be structured data.
[0082] It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be described one by one.
[0083] At the same time, it should be understood that for the sake of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0084] The following description of at least one exemplary embodiment is actually only illustrative and in no way constitutes any limitation to the present disclosure and its application or use.
[0085] Techniques, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and devices should be regarded as part of the specification.
[0086] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0087] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0088] Terminal devices, computer systems, servers and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules may include routines, programs, target programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0089] Application Overview
[0090] In the process of implementing the present disclosure, the inventors found that the 3D scene generation methods in the prior art have at least the following problems: They usually work offline, which prevents user interaction during the generation process. For example, WonderJourney uses an LLM to generate scene descriptions, adopts a text-driven process to achieve coherent 3D scene generation, and uses a vision-language model for result verification. This process takes several minutes and is not suitable for interactive use. WonderWorld reduces the reconstruction time by using FastLayered Gaussian Surfels and generates geometrically consistent scenes through guided diffusion depth estimation, but each scene still takes a long time. To improve the efficiency of 3D scene generation, the inventors propose the following 3D scene generation method.
[0091] Exemplary Method
[0092] Figure 1 It is a schematic flowchart of a 3D scene generation method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 1 shown, and includes the following steps:
[0093] Step 102: Based on the first global 3D scene graph corresponding to the first pose, determine the incomplete image and the incomplete depth map corresponding to the second pose.
[0094] Among them, the second pose is determined based on the first pose through processing.
[0095] In this embodiment, the generation process of the 3D scene is gradually expanded. Each time, a local 3D scene graph is generated based on the simulated camera (viewpoint) at a pose. By continuously changing the simulated camera from one pose to another pose to expand the 3D scene range. Therefore, in this embodiment, the first pose can be understood as the previous pose, and the second pose can be understood as the next pose. When the current local scene graph is the i-th iteration, among them, the second global 3D scene graph obtained from the (i - 1)-th iteration is used as the first global 3D scene graph of the i-th iteration. Optionally, the second pose can be obtained by rotating and / or translating the first pose. The corresponding rotation and translation matrix can be preset (by presetting the rotation and translation matrix to achieve preset size expansion of the scene each time, and gradually generate a complete 3D scene graph) or can be input by the user. When receiving the rotation and translation matrix input by the user, real-time user interactive 3D scene graph generation is realized.
[0096] Step 104: Process the incomplete image and the received text description information through the first network model to determine the complete image.
[0097] In an embodiment, the text description information can be input by the user or randomly matched with the preset text description information; when receiving the text description information input by the user, real-time user interactive 3D scene graph generation is realized; in this embodiment, the incomplete image is a 2D image with only partial content corresponding to the second pose abstracted from the first global 3D scene graph. The first network model combines this partial 2D image and the text description information to generate a complete image that complements the 2D image corresponding to the text description content. For example, the incomplete image includes a local tree crown, and the text description information input by the user is a tree. The first network model will output a complete image of the tree with the content complemented.
[0098] Step 106: Process the incomplete depth map and the complete image through the second network model to determine the complete depth map.
[0099] In this embodiment, since the second pose is obtained by rotating and translating the first pose, therefore, in the case of only knowing the first global 3D scene graph, the depth map corresponding to the second pose is an incomplete depth map, that is, the depth of the part in the incomplete depth map that does not correspond to the first global 3D scene graph is unknown; this embodiment uses the second network model to combine the complete image to fill in the depth part in the incomplete depth map, which can effectively fill in the depth information and enhance the consistency and integrity of the 3D scene.
[0100] Step 108: Process the second pose, the complete depth map, and the complete image through a third network model to determine a second global 3D scene graph.
[0101] The second global 3D scene graph includes at least the first global 3D scene graph.
[0102] In this embodiment, after obtaining the complete image based on the first network model and the complete depth map based on the second network model, the local 3D scene graph corresponding to the second pose can be determined through the third network model in combination with the second pose, and the local 3D scene graphs are stitched and merged according to the rotation and translation relationship between the first pose and the second pose to obtain a second global 3D scene graph with an expanded viewing angle range.
[0103] The 3D scene generation method provided in the above embodiments of the present disclosure determines an incomplete image and an incomplete depth map corresponding to a second pose based on the first global 3D scene graph corresponding to the first pose; the second pose is determined through processing based on the first pose; the first network model processes the incomplete image and the received text description information to determine a complete image; the second network model processes the incomplete depth map and the complete image to determine a complete depth map; the third network model processes the second pose, the complete depth map, and the complete image to determine a second global 3D scene graph; the second global 3D scene graph includes at least the first global 3D scene graph; this embodiment simplifies the 3D scene generation process through the first network model, the second network model, and the third network model, improves the speed of 3D scene generation, and allows interaction with the user during the 3D scene generation process, realizing diverse and coherently connected 3D scene generation.
[0104] This embodiment realizes a new framework designed specifically for real-time interactive 3D scene generation. To address the key challenge of inference efficiency, optimizations are made in both geometric representation and appearance modeling.
[0105] As Figure 2 shown, based on the above Figure 1 shown embodiment, step 102 may include the following steps:
[0106] Step 1021: Obtain the first global 3D scene graph corresponding to the first pose.
[0107] Optionally, the first global 3D scene graph is determined by processing the received first pose, a preset image, and a preset depth map through a third network model; or, the second global 3D scene graph is used as the first global 3D scene graph.
[0108] In this embodiment, during the iterative process of scene expansion, the first global 3D scene graph is usually the second global 3D scene graph output in the previous iteration. Additionally, there are special scenarios. Before the iteration starts, when the first global 3D scene graph is obtained for the first time, based on the preset image of the received plane, the first pose (initial pose) input by the user, and the preset depth map (which can be determined according to the depth information input by the user or randomly generated), they are input into the third network model, and the first global 3D scene graph can be generated through the third network model.
[0109] Step 1022: Process the first pose based on the received rotation and translation matrix to obtain the second pose.
[0110] Optionally, in this embodiment, the rotation and translation matrix is an n*n matrix. For example, it is a 4*4 matrix. The rotation angle and translation distance to be executed from the first pose to the second pose are represented by the rotation and translation matrix. In this embodiment, the rotation and translation matrix input by the user can be received through a preset interface.
[0111] Step 1023: Process the first global 3D scene graph according to the rotation and translation matrix to determine the incomplete image and incomplete depth map corresponding to the second pose.
[0112] In this embodiment, the first global 3D scene graph corresponds to the first pose. To obtain the part of the first global 3D scene graph corresponding to the second pose, the content in the first global 3D scene graph can be screened through the rotation and translation matrix. For example, the selection frame (3D image display area) corresponding to the first global 3D scene graph is rotated and translated through the rotation and translation matrix, and only the part that is still displayed in the selection frame after being processed by the rotation and translation matrix is retained, so as to obtain an incomplete 3D scene graph. At this time, the incomplete 3D scene graph is decomposed to obtain the incomplete image and incomplete depth map. In this embodiment, the rotation and translation matrix is used to change the pose of the simulated camera, so as to obtain the incomplete image and incomplete depth map corresponding to the first global 3D scene graph when the simulated camera is in the second pose, which provides the basis for obtaining the second global 3D scene graph corresponding to the second pose subsequently.
[0113] As Figure 3 shown, based on the above Figure 1 shown embodiment, step 108 may include the following steps:
[0114] Step 1081: Extract features from the complete image through the backbone sub-network in the third network model to obtain matching features and image features.
[0115] Optionally, the third network model may be StepSplat, which at least includes a backbone sub-network (backbone, the network structure of which may include several convolutional layers). Feature extraction is performed on the complete image through the backbone sub-network to obtain matching features and image features.
[0116] Step 1082: Based on the matching features and the second pose, optimize the complete depth map to determine the optimized depth map.
[0117] Optionally, the optimization of the complete depth map may be implemented based on a Depth Guided Cost Volume. In this embodiment, since the complete image and the complete depth map respectively come from the outputs of different network models, there may be a non-corresponding situation. In this embodiment, the correspondence between the optimized depth map after optimization and the complete image is improved through optimization, solving the problem of the mismatch between the generated image and the depth in the 3D scene.
[0118] Step 1083: Based on the image features and the optimized depth map, determine the second local 3D scene graph.
[0119] In this embodiment, when the optimized depth map is determined, the image features can be projected into the 3D space through the depth information included in the optimized depth map to obtain the second local 3D scene graph.
[0120] Step 1084: Based on the second local 3D scene graph and the first global 3D scene graph, determine the second global 3D scene graph.
[0121] In this embodiment, the third network model solves the geometric optimization in 3D scene generation and can quickly achieve 3D scene expansion (for example, achieve 3D scene expansion within 0.26 seconds). Different from the traditional 3DGS methods that rely on iterative training to update 3D representations, StepSplat draws on the idea of recent feed-forward methods and directly performs 3DGS inference, greatly improving the efficiency of 3D scene expansion.
[0122] In some alternative embodiments, Step 1082 may include:
[0123] Obtain the historical matching features that are the nearest neighbors to the second pose from the feature memory based on the rotation and translation matrix corresponding to the second pose.
[0124] The Feature Memory pre-stores historical matching features corresponding to multiple different poses for being called by the cost volume sub-network when calculating other poses subsequently.
[0125] The optimized depth map is determined by processing the historical matching features, matching features, and complete depth map through the cost volume sub-network in the third network model.
[0126] Optionally, multiple depth candidates are generated based on the complete depth map; the similarity between the matching features and the historical matching features is calculated in the depth space corresponding to the multiple depth candidates, and the probability value of each depth candidate among the multiple depth candidates is determined; the weighted sum of the multiple depth candidates is performed with the probability value as the weight to determine the optimized depth map.
[0127] Optionally, in this embodiment, the cost volume sub-network (Depth Guided Cost Volume) is used to call the historical matching features corresponding to the historical pose (e.g., the first pose) that is closest to the second pose corresponding to the matching features from the feature memory bank, generate multiple depth candidates based on the input complete depth map, and form a cost volume. That is, by calculating the normalized dot product correlation (similarity value) between the matching features and the historical matching features in the space corresponding to different depth candidates, this similarity value reflects the probability values of different depth candidates, and the optimized depth map (cost volume) is determined by performing a weighted sum of the multiple depth candidates with this probability value, which is more accurate than randomly generated multiple depth candidates in the prior art.
[0128] In this embodiment, by maintaining a feature memory bank to store the matching features of each iterative pose, the third network model extends the feed-forward paradigm to interactive 3D geometric representation, ensuring consistency during dynamic viewpoint changes. This module adaptively constructs a cost volume as the viewpoint changes.
[0129] In terms of network structure, the third network model further includes a segmentation sub-network, which receives the target depth value output by the cost volume sub-network and the image features output by the backbone sub-network. Correspondingly, step 1083 may include:
[0130] The target depth value and the image features are processed through the segmentation sub-network in the third network model to obtain a target depth map and a Gaussian model.
[0131] The Gaussian model is projected into the 3D space based on the target depth map to obtain a second local 3D scene graph.
[0132] In this embodiment, the segmentation sub-network may be a 2D U-Net network. Through this segmentation sub-network, the Gaussian model can be projected into the 3D space based on the target depth map, and thus the second local 3D scene graph corresponding to the second pose can be obtained.
[0133] In some alternative embodiments, based on the second local 3D scene graph corresponding to the second pose obtained in the above steps, step 1084 may include:
[0134] Project the second partial 3D scene graph and the first global 3D scene graph into the pixel coordinate system;
[0135] Filter the 3D image information in the second partial 3D scene graph and the first global 3D scene graph that does not meet the depth consistency constraint in the pixel coordinate system;
[0136] Perform stitching on the 3D image information that meets the depth consistency constraint after screening in the second partial 3D scene graph and the first global 3D scene graph to obtain the second global 3D scene graph.
[0137] In this embodiment, both the first global 3D scene graph and the second partial 3D scene graph can be represented by a Gaussian model.
[0138] Optionally, the fusion of the partial 3D scene graph implemented in this embodiment can be implemented based on an incremental fusion module (Incremental Fusion). The incremental fusion module updates the local Gaussian model to the global Gaussian model through depth constraints (the depth constraint means calculating the depth difference between the corresponding parts in the second partial 3D scene graph and the first global 3D scene graph, and determining the part with the smaller depth difference as the redundant part. Optionally, the corresponding parts can be determined according to the relationship between the first pose and the second pose to determine the similar parts in the two 3D scene graphs, project the 3D scene graph of the similar parts into the 2D image under the second pose, determine the points that may be redundant, and then calculate the depth difference)), ensuring a continuous and consistent 3D scene representation. This process includes projecting all global Gaussian models into the current pixel coordinate system, screening the Gaussian models that violate the depth consistency constraint, and only merging the valid local Gaussian models.
[0139] Figure 4 It is a schematic structural diagram of the third network model in the 3D scene generation method provided by an exemplary embodiment of the present disclosure. As Figure 4 shown, the third network model includes a backbone sub-network, a cost volume sub-network, a feature memory bank, and a segmentation sub-network.
[0140] The backbone sub-network receives a complete image for feature extraction and outputs matching features and image features. The obtained matching features are input into the cost volume sub-network and the feature memory bank; the image features are input into the segmentation sub-network.
[0141] The feature memory bank receives and saves the matching features, receives the second pose, searches for the historical matching features corresponding to the pose closest to the second pose according to the second pose, and inputs the historical matching features into the cost volume sub-network.
[0142] The cost volume sub-network receives the matching features, historical matching features, and the complete depth map; obtains an optimized depth map and inputs the optimized depth map into the segmentation sub-network.
[0143] The segmentation sub-network receives the image features and the optimized depth map, and obtains a second local 3D scene graph.
[0144] In addition, the third network model may further include: an incremental fusion module. The incremental fusion module receives the second local 3D scene graph and the first global 3D scene graph, performs depth matching on the second local 3D scene graph and the first global 3D scene graph, filters out the parts in the second local 3D scene graph that violate the depth consistency constraint, and splices the other parts with the first global 3D scene graph to obtain a second global 3D scene graph.
[0145] The first network model, the second network model, and the third network model provided in the above embodiments of the present disclosure are pre-trained before application. To train the first network model and the second network model, the embodiments of the present disclosure provide a training set specifically for interactive 3D scene generation. This embodiment constructs a data set based on a variety of existing 3D scene generation methods, and uses this data set to train all network models (the first network model, the second network model, and the third network model). Optionally, use a variety of 3D scene generation methods (any 3D scene generation method in the prior art) to create 3D scene graphs that each method is good at, and use a VLM model to verify whether the generated data conforms to the defined scene. The finally obtained data set contains more than 6 million frames rendered by simulating interactive trajectories, including rotation paths, linear movements, and mixed trajectories. The data set provided in this embodiment includes most 3D task scenarios, and can achieve more comprehensive training of the network model. For example, it covers four categories: indoor environments, urban landscapes, natural terrains, and stylized art scenes, etc.
[0146] In some alternative embodiments, before processing the incomplete image and the received text description information through the first network model to determine the complete image, it further includes:
[0147] Training the first initial network model with the training data set to obtain the trained first initial network model; the training data set includes 3D images of various scenes, and each 3D image has a known scene category;
[0148] Compressing and accelerating the first initial network model through model distillation technology to obtain the first network model.
[0149] The first network model (FastPaint) proposed in this embodiment does not limit the network structure, and any network model structure for image completion can be adopted. In this embodiment, the model distillation technology is used to reduce the inference steps of the first network model to 2 steps, and the repair ability of the pre-trained model is enhanced through distillation and fine-tuning, making it applicable to interactive 3D generation. Optionally, by synergistically utilizing the advantages of ODE trajectory retention and ODE trajectory reconstruction, knowledge distillation is performed on the first network model, reducing the required inference steps while maintaining the quality of appearance modeling.
[0150] In this embodiment, to address the problem that the repair area in 3D scene generation is different from the fine-tuning stage, a dataset dedicated to training FastPaint (including datasets of various 3D scenes) can be constructed. And the camera pose is simulated and part of the image content is occluded by a mask to simulate the interactive 3D generation process. By obtaining the depth map and the image, and using projection to obtain the mask, it is ensured that the dataset meets the specific requirements of the repair task in this context. The construction of this dataset is similar to the methods of the third network model and the second network model, especially in simulating the camera trajectory, which helps the model to adapt to interactive 3D scene generation.
[0151] In some alternative embodiments, the process of using the second network model to determine the complete depth map may include:
[0152] Based on the incomplete depth map, determine a binary validity mask;
[0153] Among them, the binary validity mask indicates which coordinates in the complete image have depth and which coordinates do not have depth, and is determined based on the incomplete depth map.
[0154] Input the incomplete depth map, the binary validity mask, and the complete image into the second network model.
[0155] Among them, the second network model is obtained through training.
[0156] In this embodiment, the second network model (QuickDepth) is trained based on a lightweight depth estimation model. The network structure of the second network model is not limited in this embodiment, and existing network structures can be adopted. The input of the second network model includes the complete image corresponding to the second pose (output of the first network model), the incomplete depth map, and the binary validity mask. To meet the requirements of interactive 3D scene generation, a dataset containing diverse scenes in indoor, outdoor environments, as well as comics and artworks is created, and multiple camera poses are set to simulate the real 3D scene generation process. In the processing of each frame, the depth map of the previous frame is transformed to the current frame coordinate system through relative pose to generate an incomplete depth map, and the binary validity mask for the area to be filled is marked. During the training process, sometimes the true depth map of the target frame is completely covered, and sometimes the deformed depth map and the corresponding mask are used as inputs. The prediction result is supervised by the L1 loss function to ensure that its prediction result is consistent with the true depth Figure 1 to improve the expressiveness and accuracy of the second network model in complex and variable scenes. This method enables the second network model to effectively fill in depth information and enhance the consistency and integrity of the 3D scene.
[0157] Figure 5 It is a processing schematic diagram of an optional example of the 3D scene generation method provided by another exemplary embodiment of the present disclosure. As Figure 5 shown, the network structure in this embodiment includes a first network model, a second network model, and a third network model. Only the process of the i-th iteration is shown in this embodiment.
[0158] The input data is the second global 3D scene graph output from the (i - 1)-th iteration as the first global 3D scene graph of the current iteration. The rotation and translation matrix of the i-th iteration input by the user (to achieve user interaction) is received to determine the second pose; the first global 3D scene graph is processed according to the second pose, and the obtained incomplete image is input into the first network model, and the obtained incomplete depth map is input into the second network model.
[0159] The first network model receives the incomplete image and the text description information input by the user (to achieve user interaction), completes the incomplete image through the text description information to obtain a complete image, and inputs the complete image into the second network model and the third network model.
[0160] The second network model receives the complete image and the incomplete depth map, determines the binary validity mask based on the incomplete depth map, realizes depth completion of the incomplete depth map to obtain a complete depth map, and inputs the complete depth map into the third network model.
[0161] The third network model receives the complete image, the complete depth map, and the second pose corresponding to the i-th iteration, and determines the second local 3D scene map corresponding to the second pose.
[0162] It may further include an incremental fusion module (the function of this module can be integrated in the third network model or independent of the third network model), which merges the second local 3D scene map with the first global 3D scene map to obtain the second global 3D scene map expanded after the i-th iteration. By continuing the iteration, the global 3D scene map can be finally obtained.
[0163] In order to further enhance depth consistency in this embodiment, a second network model (QuickDepth) with a lightweight depth completion function is integrated to provide a consistent depth prior for the third network model to construct the cost volume. In terms of appearance, a first network model (FastPaint) is proposed for an efficient method of real-time appearance refinement. Compared with traditional diffusion-based inpainting methods sd (these methods require dozens of inference steps to refine appearance modeling), the first network model can achieve a similar effect with only 2 inference steps while maintaining spatial appearance consistency. This method significantly improves the processing speed and makes real-time interactive experience possible.
[0164] Any of the image processing methods provided in the embodiments of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers, etc. Alternatively, any of the image processing methods provided in the embodiments of the present disclosure can be executed by a processor. For example, the processor executes any of the image processing methods mentioned in the embodiments of the present disclosure by calling the corresponding instructions stored in the memory. Details will not be described hereinafter.
[0165] Exemplary Device
[0166] Figure 6 It is a schematic structural diagram of a 3D scene generation device provided by an exemplary embodiment of the present disclosure. As Figure 6 shown, the device provided in this embodiment includes:
[0167] A preprocessing module 61, configured to determine an incomplete image and an incomplete depth map corresponding to the second pose based on the first global 3D scene map corresponding to the first pose.
[0168] Wherein, the second pose is determined based on the first pose after being processed.
[0169] A first processing module 62, configured to process the incomplete image and the received text description information through the first network model to determine the complete image.
[0170] A second processing module 63, configured to process the incomplete depth map and the complete image through the second network model to determine the complete depth map.
[0171] A third processing module 64, configured to process the second pose, the complete depth map, and the complete image through a third network model to determine a second global 3D scene graph.
[0172] The second global 3D scene graph includes at least the first global 3D scene graph.
[0173] The 3D scene generation device provided in the above embodiments of the present disclosure determines an incomplete image and an incomplete depth map corresponding to a second pose based on the first global 3D scene graph corresponding to the first pose; the second pose is determined through processing based on the first pose; processes the incomplete image and the received text description information through a first network model to determine a complete image; processes the incomplete depth map and the complete image through a second network model to determine a complete depth map; processes the second pose, the complete depth map, and the complete image through a third network model to determine a second global 3D scene graph; the second global 3D scene graph includes at least the first global 3D scene graph; in this embodiment, the first network model, the second network model, and the third network model simplify the 3D scene generation process, improve the 3D scene generation speed, and allow interaction with the user during the 3D scene generation process, realizing diverse and coherently connected 3D scene generation.
[0174] In some optional embodiments, the preprocessing module 61 is specifically configured to obtain the first global 3D scene graph corresponding to the first pose; process the first pose based on the received rotation and translation matrix to obtain the second pose; process the first global 3D scene graph according to the rotation and translation matrix to determine the incomplete image and the incomplete depth map corresponding to the second pose.
[0175] Optionally, when obtaining the first global 3D scene graph corresponding to the first pose, the preprocessing module is configured to process the received first pose, the preset image, and the preset depth map through a third network model to determine the first global 3D scene graph; or use the second global 3D scene graph as the first global 3D scene graph.
[0176] In some optional embodiments, the third processing module 64 includes:
[0177] A feature extraction unit, configured to extract features from the complete image through a backbone sub-network in the third network model to obtain matching features and image features;
[0178] A depth optimization unit, configured to perform optimization processing on the complete depth map based on the matching features and the second pose to determine an optimized depth map;
[0179] A local scene generation unit, configured to determine a second local 3D scene graph based on the image features and the optimized depth map;
[0180] An incremental fusion unit for determining a second global 3D scene graph based on a second local 3D scene graph and a first global 3D scene graph.
[0181] Optionally, a depth optimization unit specifically configured to obtain a historical matching feature nearest to the second pose from a feature memory bank based on a rotation and translation matrix corresponding to the second pose; a plurality of historical matching features corresponding to different poses are pre-stored in the feature memory bank; the historical matching feature, the matching feature, and the complete depth map are processed by a cost volume sub-network in the third network model to determine an optimized depth map.
[0182] Optionally, when the depth optimization unit processes the historical matching feature, the matching feature, and the complete depth map through the cost volume sub-network in the third network model to determine the optimized depth map, it is used to generate a plurality of depth candidate values based on the complete depth map; calculate the similarity between the matching feature and the historical matching feature in the depth space corresponding to the plurality of depth candidate values, and determine the probability value of each depth candidate value among the plurality of depth candidate values; perform weighted summation on the plurality of depth candidate values with the probability value as the weight to determine the optimized depth map.
[0183] Optionally, a local scene generation unit specifically configured to process the target depth value and the image feature through a segmentation sub-network in the third network model to obtain a target depth map and a Gaussian model; project the Gaussian model into the 3D space based on the target depth map to obtain a second local 3D scene graph.
[0184] Optionally, the incremental fusion unit specifically configured to project the second local 3D scene graph and the first global 3D scene graph into a pixel coordinate system; filter 3D image information in the second local 3D scene graph and the first global 3D scene graph that does not meet the depth consistency constraint in the pixel coordinate system; perform stitching on the 3D image information in the second local 3D scene graph and the first global 3D scene graph that meets the depth consistency constraint after screening to obtain a second global 3D scene graph.
[0185] In some alternative embodiments, the apparatus provided in this embodiment may further include:
[0186] A model training module for training a first initial network model through a training data set to obtain a trained first initial network model; the training data set includes 3D images of multiple scenes, and each 3D image has a known scene category; the first initial network model is compressed and accelerated through model distillation technology to obtain a first network model.
[0187] In some alternative embodiments, the second processing module 63 is specifically configured to determine a binary validity mask based on an incomplete depth map; input the incomplete depth map, the binary validity mask, and the complete image into a second network model; the second network model is obtained through training.
[0188] Exemplary Electronic Device
[0189] Next, with reference to Figure 7 an electronic device according to an embodiment of the present disclosure will be described. The electronic device may be either or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device may communicate with the first device and the second device to receive the input signals collected therefrom.
[0190] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.
[0191] As Figure 7 shown, the electronic device includes one or more processors and a memory.
[0192] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0193] The memory may store one or more computer program products. The memory may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products may be stored on the computer-readable storage media, and the processor may run the computer program products to implement the 3D scene generation method of the various embodiments of the present disclosure described above and / or other desired functions.
[0194] In one example, the electronic device may further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0195] In addition, the input device may further include, for example, a keyboard, a mouse, and so on.
[0196] The output device may output various information to the outside, including the determined distance information, direction information, etc. The output device may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.
[0197] Of course, for simplicity, Figure 7Only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0198] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the 3D scene generation method according to various embodiments of the present disclosure described in the above part of this specification.
[0199] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0200] Furthermore, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and when the computer program instructions are run by a processor, cause the processor to execute the steps in the 3D scene generation method according to various embodiments of the present disclosure described in the above part of this specification.
[0201] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0202] The basic principles of the present disclosure have been described in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and facilitation of understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0203] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the corresponding description in the method embodiment.
[0204] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the term "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0205] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers the recording medium storing the program for executing the method according to the present disclosure.
[0206] It should also be noted that in the apparatuses, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0207] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0208] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A 3D scene generation method, characterized in that: include: Determine an incomplete image and an incomplete depth map corresponding to the second pose based on a first global 3D scene graph corresponding to the first pose; The second posture is determined based on the first posture after processing; Processing the incomplete image and the received text description information through a first network model to determine a complete image; Processing the incomplete depth map and the complete image through a second network model to determine a complete depth map; The second pose, the complete depth map and the complete image are processed by a third network model to determine a second global 3D scene graph; the second global 3D scene graph at least includes the first global 3D scene graph.
2. The method according to claim 1, characterized in that The determining, based on the first global 3D scene graph corresponding to the first pose, an incomplete image and an incomplete depth map corresponding to the second pose comprises: Obtaining a first global 3D scene graph corresponding to the first posture; Processing the first posture based on the received rotation and translation matrix to obtain the second posture; The first global 3D scene graph is processed according to the rotation and translation matrix to determine the incomplete image and the incomplete depth map corresponding to the second posture.
3. The method according to claim 2, characterized in that The obtaining of a first global 3D scene graph corresponding to the first posture includes: Processing the received first posture, preset image and preset depth map through a third network model to determine the first global 3D scene graph; or, The second global 3D scene graph is used as the first global 3D scene graph.
4. The method according to any one of claims 1 to 3, characterized in that: The step of processing the second pose, the complete depth map, and the complete image by a third network model to determine a second global 3D scene graph includes: Extracting features from the complete image using the backbone sub-network in the third network model to obtain matching features and image features; Optimizing the complete depth map based on the matching features and the second pose to determine an optimized depth map; Determining a second local 3D scene graph based on the image features and the optimized depth map; Based on the second local 3D scene graph and the first global 3D scene graph, the second global 3D scene graph is determined.
5. The method according to claim 4, characterized in that The step of optimizing the complete depth map based on the matching features and the second pose to determine an optimized depth map includes: Based on the rotation and translation matrix corresponding to the second posture, a historical matching feature that is the nearest neighbor to the second posture is obtained from a feature memory library; the feature memory library pre-stores the historical matching features corresponding to a plurality of different postures; The historical matching features, the matching features and the complete depth map are processed by a cost volume subnetwork in the third network model to determine the optimized depth map.
6. The method according to claim 5, characterized in that The step of processing the historical matching features, the matching features, and the complete depth map through the cost volume subnetwork in the third network model to determine the optimized depth map includes: generating a plurality of depth candidate values based on the complete depth map; Calculating the similarity between the matching feature and the historical matching feature in the depth space corresponding to the multiple depth candidate values, and determining a probability value of each of the multiple depth candidate values; The plurality of depth candidate values are weightedly summed using the probability values as weights to determine the optimized depth map.
7. The method according to any one of claims 4 to 6, characterized in that: The determining a second local 3D scene graph based on the image features and the optimized depth map includes: Processing the target depth value and the image feature through the segmentation subnetwork in the third network model to obtain a target depth map and a Gaussian model; The Gaussian model is projected into a 3D space based on the target depth map to obtain the second local 3D scene map.
8. The method according to any one of claims 4 to 7, characterized in that: The determining the second global 3D scene graph based on the second local 3D scene graph and the first global 3D scene graph includes: Projecting the second local 3D scene graph and the first global 3D scene graph into a pixel coordinate system; Screening, in the pixel coordinate system, 3D image information that does not comply with a depth consistency constraint in the second local 3D scene graph and the first global 3D scene graph; The 3D image information that meets the depth consistency constraint after screening in the second local 3D scene graph and the first global 3D scene graph is stitched to obtain the second global 3D scene graph.
9. The method according to any one of claims 1 to 8, characterized in that: Before the incomplete image and the received text description information are processed by the first network model to determine the complete image, the method further includes: Training the first initial network model through a training data set to obtain the trained first initial network model; the training data set includes 3D images of multiple scenes, each of the 3D images has a known scene category; The first initial network model is compressed and accelerated by using a model distillation technique to obtain the first network model.
10. The method according to any one of claims 1 to 9, characterized in that: The step of processing the incomplete depth map and the complete image by using a second network model to determine a complete depth map includes: Based on the incomplete depth map, determining a binary validity mask; The incomplete depth map, the binary validity mask and the complete image are input into the second network model; the second network model is obtained through training.
11. A 3D scene generation device, characterized in that: include: A preprocessing module, configured to determine an incomplete image and an incomplete depth map corresponding to a second pose based on a first global 3D scene graph corresponding to the first pose; The second posture is determined based on the first posture after processing; A first processing module, configured to process the incomplete image and the received text description information through a first network model to determine a complete image; A second processing module, configured to process the incomplete depth map and the complete image through a second network model to determine a complete depth map; A third processing module is used to process the second pose, the complete depth map and the complete image through a third network model to determine a second global 3D scene graph; the second global 3D scene graph at least includes the first global 3D scene graph.
12. An electronic device, characterized in that: include: a memory for storing a computer program product; The processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, the 3D scene generation method described in any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the 3D scene generation method described in any one of claims 1 to 10 is implemented.
14. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the 3D scene generation method described in any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Three-dimensional scene style migration method and device, equipment and storage medium
CN116934936A
Construction method and device of three-dimensional point cloud model, electronic equipment and storage medium
CN118864695A
6D pose estimation method and device and electronic equipment
CN119444843A
System and method for 3D scene reconstruction with dual complementary pattern illumination
US20170352161A1
Cited By
Video generation method and device, electronic equipment, storage medium and program product
CN120434373A